Ep 111 · Aug 27, 2026 · 24 min

Building the Reference Architecture for AI-Enabled Security Operations

Nick Lippis with Peter Campbell, Security Researcher & Agentic AI Working Groups Lead at ONUG

Prefer audio? Listen on
Building the Reference Architecture for AI-Enabled Security Operations — watch on YouTube

About this episode

How should a Security Operations Center work when attacks arrive at machine speed? In this episode, ONUG co-founder Nick Lippis and Peter Campbell, Security Researcher and Agentic AI Working Groups Lead at ONUG, walk through the reference architecture that ONUG's AI-Enabled SOC working group has developed. The group exists because ONUG's board, representing more than 600 of the Global 2000 with roughly 1.5 trillion dollars in combined IT spend, voted the AI-enabled SOC one of its most urgent priorities: alert volume is exploding, non-human identities now outnumber human users by roughly 45 to 1 in this environment, and human-speed response can no longer keep up with AI-driven attacks.

The architecture starts with planning: defining agentic personas that dictate each agent's rights and policies, assigning autonomy levels from L0 (observe only) through L3 (broad authority to act autonomously), choosing between frontier and open-weight models, and scoping the estate under watch, from firewalls to cloud IAM to WAFs. Agents then operate through one of two access patterns, a distributed agent-to-agent model or an orchestrated, federated model, with every interaction tracked for audit and compliance. The SOC version of the blueprint is deliberately more conservative than its network-operations sibling: agents can detect, investigate, and propose fixes, but the decide gate stays with humans, because a false positive can take down legitimate business traffic, and acting too fast can destroy forensic evidence or tip off an adversary still inside the network.

What we cover in this episode

  1. How ONUG working groups are born. Every board member brings one or two use cases to the board meeting, the room votes with ten stickers each, and the winners signal both industry priority and real budget development.
  2. The three active working groups. The agentic control plane drew the most votes and acts as the umbrella, alongside ONUG Connect for WAN provisioning, autonomous infrastructure for the NOC, and the AI-enabled SOC.
  3. Why the SOC must be rethought for AI. Machine-speed attacks, exploding alert volume, and non-human identities outnumbering humans roughly 45 to 1 make the pre-AI security operating model unworkable.
  4. Planning the architecture: personas and autonomy levels. Agent personas dictate rights and policy the way L1, L2, and L3 engineer tiers do today, and each agent gets an autonomy level from L0 (observe) to L3 (act autonomously).
  5. Pattern A vs. Pattern B. Agents reach telemetry, rule sets, sources of truth, and MCP servers either directly in a distributed agent-to-agent pattern or through an orchestrator in a federated pattern, with every interaction audited.
  6. The decide gate: what stays human. Automation belongs in detect, investigate, and propose; the decision to act stays gated, because blanket actions can destroy evidence, break chain of custody, or tip off the adversary.
  7. What to ask security vendors now. Nearly every security product claims agentic AI, so practitioners should press vendors on agent autonomy levels, agent personas, and where agents actually run, including at MSSPs.

The lines worth sharing

“Attacks are happening at machine speed, driven by adversarial use of AI, and we need to be able to respond in a similar manner.”

Peter Campbell

“L3 is God level, meaning that it has broad access and broad visibility to take autonomous actions.”

Nick Lippis

“We never take a blanket action. It would be based on information security and forensic principles: preservation of evidence, chain of custody, being able to have forensic replay.”

Peter Campbell

Common questions from this episode

What is an AI-enabled SOC reference architecture?

It is a blueprint for running a Security Operations Center where AI agents handle much of the operations workload. ONUG's version covers a planning layer (agent personas, autonomy levels, model selection, and the scope of the estate under watch), an access and enforcement layer with full audit trails, and an agentic operations loop that detects, investigates, and proposes responses to security events.

What are AI agent autonomy levels in security operations?

ONUG's working group defines four levels. L0 only observes, L1 observes and recommends actions to a human, L2 asks permission before acting, and L3 has broad access and visibility to take autonomous action. Teams gain trust in an agent at lower levels before promoting it toward autonomy.

Should AI agents be allowed to block threats automatically?

Not by default. The architecture places a decide gate between the agent's proposed response and enforcement, because a false positive can take down legitimate business traffic, and an automatic block can destroy forensic evidence or reveal to an adversary that they have been detected. Decisions like blocking access typically stay with a human.

What should I ask a security vendor about their agentic AI features?

Ask what autonomy level the agent operates at, what persona and rights it is granted, and where it is deployed, whether in your data center or at a managed security service provider running your tier-one SOC. Almost every security product now advertises agentic AI, and these questions separate real capability from marketing.

How do ONUG working groups choose what to work on, and how can I join?

Use cases come from ONUG board members, who present and vote on them at the board meeting before each AI Networking Summit; the top-voted topics become working groups, which signals both priority and buying intent from large enterprises. Practitioners and vendors can request to join at onug.net/collaborative by picking a working group and filling out a short form.

Read the complete conversation

Full episode transcript · 24 minutes

Attacks are happening at machine speed, driven by adversarial use of AI, and we need to be able to respond in a similar manner. Hi, everyone. I'm your host, Nick Lippis, and welcome to the Built for Trust podcast, where you get to hear from all the folks who are building and shaping AI enterprise infrastructure. Now, let's get right into it with our guests.

Hi, everyone. Welcome to the ONU Collaborative Working Group Use Case Updates. So this is really for, we wanted to share with everyone, all of the vendor community is now about to enter into the working groups after myself, Peter, and many others have been working very hard during the summer, doing our summer project around taking the use cases that were developed during the last board meeting, and creating some really useful projects for the industry.

So first, I am Nick Lippis, one of the co-founders and co-chairs at ONU. And Peter, won't you say hi? Hi, Peter Campbell. I'm a security researcher and advisor to the working groups at ONU. Awesome. Peter, you're more than just an advisor.

You're kind of like running a lot of the working groups. And I'm so, so happy to have you, you know, involved. You add such a great dimension, perspective, and to keep everything kind of moving along. No, thank you, Nick. And this has been fun. You talked about a summer project.

This has kind of been like summer camp, right? So we're starting on our journey here. Yeah. And I tell you, I love our working groups. They are, I gain so much insight from what we're doing in the working groups than on almost anything other than the conference itself.

So I agree. You know, what's really unique, and I know we'll get into it a little bit more, Nick, but I mean, it's really got a practitioner focus to it, doesn't it, right? These are people who are actually out there, you know, implementing and feeling some of the pain as well as they're, you know, bringing solutions forward, AI-driven solutions. And we're going to talk a lot about that.

Yeah. I think, you know, I think what you'll see in terms of the audience, you'll see is a real, real life because like all the folks who are participating in the working groups are not folks who are saying, oh, I'm thinking about doing that. I'd like to see this, you know, I, you know, it's like, I think this is how it's going to work.

These are the folks who are actually doing it. So it is so much different when you have real practitioners who are struggling with a lot of these topics. So let's get into it. First, I want to kind of give you a perspective on kind of how the working groups were established. This is ONU.

This is the largest organization of IT executives within the industry that are collaborating together. So if you, if you're not familiar with ONU or so, we have about, it represents about 600 plus of the global 2000 combined spend of IT is about one and a half trillion dollars. So it's a very big buying block in our industry. Most, um, most being over almost nearly 90% have budget, um, uh, purchasing authority and we've been at this for a while.

So we have market permission, uh, to lead, um, you know, so that comes with a gaining trust, uh, within, uh, within the industry. So, uh, as I mentioned, uh, a second ago, just mentioned all of those things, but who is in the working groups right now? It's really, these are mostly ONU board members, um, or folks who are leaders within the industry. So we have representation from a whole group of companies, um, that have had input, uh, into what you're, uh, kind of, uh, about to see.

So, um, like I said, very practitioner, uh, focused, the key thing, uh, about the working groups is that when we start them, the way they come to be, or they get created is that we have a board meeting that, uh, that occurs the day before, um, the AI networking summit. Every board member brings one or two use cases to that board meeting. They present those to the, um, to the rest of the board members. Uh, and then, uh, so that everyone has a good understanding of what they are.

And then, uh, we put them all on kind of like, you know, big sticky easel kind of like the large ones all around the room. There may be some 20 or so use cases and we give basically only 10 little red stickies, you know, and these are the ones that you peel off and then you can actually stick them on a little old school, but it actually, it works because you get a really great visual when it's, when it's done.

So what that voting actually does that it basically certifies that this is what the ONU board feels is important, uh, for not just them because they can't solve these problems on their own. They need the industry to start focusing on that. What it also communicates is that this is propensity to buy. This is budget development for these areas that they are mostly concerned about and that they want the entire community to start focusing, uh, their minds on.

So, um, use cases built by consumers and on the, in the large enterprise, putting their requirements against them, um, and looking to buy. So the work that we're doing in the working groups, take those use cases, which are very, you know, is like on like one kind of bulletin or one kind of like large sticky, right? Not a little sticky, like, you know, a big one, like an easel size. Uh, and then we start to, um, have a lot of discussion, um, about them.

So, um, we developed four working groups and we're really active in three right now. The fourth one, uh, we held off for, uh, for a little bit. We'll probably start that, um, either later in the fall, um, or early winter timeframe. Now, Nick, if I was just going to interject for a second, so as we're going through, as the board meets, I mean, you're looking at a kind of a theme, right? That might get put onto the, you know, onto a whiteboard and people put stickies next to it.

They voted, they've gotten to these, uh, say these three themes here. Um, so, uh, and then the working groups, uh, really get together and refine that further. Correct. Yeah. Yeah, I'm absolutely. Yeah.

We find it, define it, you know, create problem statements and reference. And so architectures, um, and so forth. So, um, yeah, thank you, Peter, uh, for that. And I think this kind of shows how the votes went. So, um, we have, um, we call WG one or working group one, that's the agentic control plan. That was, uh, the most concentrated and highly demanding one.

There's the AI enabled SOC. Um, and then there's also, uh, AI network and autonomous infrastructure. Um, so a couple of words on these before we kind of dive in, right, Peter? So one WG two, uh, we kind of splitting that into two different areas. Uh, there is the ONU connect work. Um, and that has to do a lot around provisioning of wide area networking services.

Um, and then on the autonomous infrastructure, that's a real focus on the knock. And then most of this is really looking at kind of a post mythos, a post, you know, patch storm world where that becomes the new norm, where in essence, operations now needs to be just like in essence, you know, part of a CICD pipeline. Uh, and so how do you get ready for that? How do you deal with that?

And that's really where we're spending a lot of our time and focus. Peter, you have anything you want to add to that? No, cause I, I remember being in the room as we were selecting, as we went through the voting and it's, it'll, I think it'll be interesting for people to see kind of where we've landed and how things have started to, how we've been getting, begun to shape things through the working groups and, and the way that these three working groups sort of tie together.

Yeah. Awesome. Okay. All right. Great. Um, so those are the four, um, so we're going to give kind of updates, um, you know, on, on the working groups, um, WG one, agenda control plane, extremely important, you know, uh, activity.

This is the most, um, highly, um, you know, needed, uh, area. Uh, a lot of this has to do with security, creating an agentic fabric as well. So we'll talk, uh, a lot about that, uh, and then into, um, various different domains. So there's domain one, which is, uh, autonomous networking, which is the knock. And then also, um, the sock Peter, you want to add anything to that? No, I think just going really back to a genetic control plane as being kind of the umbrella, uh, atop the other working groups.

And I know there's been a lot of, uh, panic and concern, maybe, maybe that's too strong of a word around sort of like what people have been seeing in the news recently with open AI and hucking face, but I think this is a really, uh, the time to begin, you know, solving these problems, right? And that's really what the working groups are focused on is, uh, you know, fact-based and approaches like this to solve problems.

Okay. So let's dive into, uh, working group three, and this is an AI enabled sock. So, uh, this got a large number of votes, about nine votes, uh, from our board members. There are like 44 companies that are in scope, uh, that are participating and, uh, in this marketplace, um, you know, the non-human, um, intervention to human ratio is about 45, uh, to one.

So, uh, we needed, um, we need to really start, I think, rethinking, um, how socks are operated in the AI, uh, world. Now, just as a corollary, um, you know, we did another podcast on, uh, autonomous infrastructure, which was focused on the knock. So it's, um, it, they're very similar architectures. Um, if you've seen both of them, you'll see, um, but the entire operations flow, uh, is very different and the whole, you know, um, goal set and requirements and workflow and, uh, playbooks, uh, are different, uh, there as well.

So they really do require, uh, different, um, workflows, but the same architecture, uh, is valid for, for both of them. Uh, we knew this was the problem, um, that the, uh, board, uh, new board, um, specified. Uh, we know alert volume is exploding, hard, uh, for, uh, organization to keep up, um, machine attacks, speed of machines, you know, hard for humans, uh, to keep up. Um, and then also non-human, you know, identities, you know, outnumbering, uh, human users, um, as well.

So the, the whole threat landscape is fundamentally changing, uh, with AI. And so the way that we approach security, uh, in the pre-AI, uh, world is fundamentally being rethought, needs to be rethought and redone. Peter, anything to add to that? No, I would agree. It kind of goes back to what we opened up with around kind of the mythos moment.

And, you know, the patching cycles and, uh, really what we've, uh, you know, attacks are happening at, uh, machine speed driven by adversarial use of AI. And we need to be able to respond in a similar manner. Yeah. Okay. Um, all right.

So let's take a look at the reference architecture for how you AI enabled the SOC and be able to catch up, uh, with machine speed. Okay. Ready drum roll. Sure. Da, da, da, da, da, da, da, da, da.

Here we go. Um, AI enabled SOC. So, um, this is the reference architecture, um, that the working group, um, has developed. So, uh, if you saw, um, if you saw, listen to, uh, working group two, um, everything from this part, the whole planning part to, uh, half of this is the same. And what's really different is the agentic operations loop.

Uh, but we're going to assume that you didn't see, um, that previous, um, working group. And so we're going to talk through the planning part of this. So Peter, I'll start and, you know, and kind of like hand, I hand it off to you. Okay. No, that sounds good, Nick. Okay.

So, um, um, in the planning part, you really need to spend a lot of time thinking through and use AI tools to do this, um, you're the agentic personas, each persona, you're going to have the end different personas, uh, each focused on doing very specific tasks, just like you have L1, L2, L3, uh, engineers, uh, today. And within those engineers, they also are segmented in the kind of the tasks that they do. So you need to think about what kind of persona, um, the agent is going to have, and that's going to dictate their rights, um, that are granted to them.

Um, uh, the policy that's associated with, um, uh, as well. So persona is really important to start to build, uh, that out. Then you need to start thinking about what is the agent autonomy level. So we like to think of these as an L0 through L3. L0 is just observing, taking a look, you know, at things that are going on, but not acting on anything.

Uh, L1 is observing, but asking then now recommending. So going up to a human saying, Hey, I'm seeing this, you may want to take this action. Uh, L2 is that, okay, um, the, I'm seeing these things. Um, I want to act upon this, you know, mother, may I, can I do that? Uh, and then L3 is God level, uh, meaning that it has broad access, broad visibility, um, to take autonomous actions, uh, based upon things that are occurring incidents or things that are occurring within the security posture, uh, of an organization.

Um, so that actually provides a really nice, um, you know, uh, way in which you can segment the various different, uh, agents, um, and then, um, how you can start to gain trust in those agents and start moving them more and more towards, uh, autonomy. Uh, and then those agents are going to rely upon various different models. Um, whether they are, um, kind of frontier models that are in the public domain or whether they are open weight models within your own organization.

So which models, um, are you using, uh, to process a lot of security, uh, alerts, alarms, uh, logs and so forth. Uh, and then the important thing too, is that what is the scope of the estate that is under watch? Um, so firewalls, cloud IAM, you know, um, WAFs, you know, it's like just, what is that scope of, uh, of a state?

Peter. You know, that really, really well said. I would say that, you know, the important thing is that, you know, the blueprint is the same, right? Between working group, uh, two and working group three. If you saw that working group two, that's the network, uh, autonomous operations working group.

Uh, but the SOC version is, uh, deliberately a little bit different, a little bit more conservative as, as Nick said, uh, right. So that, um, as we get into the agent enforcement and audit, uh, there's things that the, uh, agent can do, and there's things that the agent can't do, right? So the, uh, we don't want the agent, uh, rewriting, uh, detection rules or enforcement rules, for example, right?

Um, and I think the critical differentiators in a box four where it says, uh, decide, right? So the decide gate really determines, uh, what types of actions are we going to take? You know, we may not want to block access, right? We may not want to tip our hand to the adversary who may be, uh, within our network, you know, maybe gathering intelligence or, you know, uh, or basic or also, uh, eliminating key evidence that would be needed during the investigation.

So we never take a blanket action, but it would be based on, uh, you know, sort of, uh, information security and forensic, uh, principles around, you know, preservation of evidence, chain of custody, being able to have forensic replay, um, you know, a record that would, uh, that would hold up. So, I mean, there's a number of different decisions and that all goes back to the agent persona and the persona that you have granted.

So I think, uh, the other thing I would add, I think is also within, uh, within cybersecurity, there are a lot of vendors now who are coming into this space with, with solutions that are almost every security solution now has some type of a genetic AI capability built into it. So it's really important as you're meeting with vendors, I know this happens in, uh, you know, the SOCs that I talk to is they're speaking with the, uh, the main players really, really trying to understand, you know, what is level of, uh, autonomy?

What is the persona of the agent, uh, that's being deployed? And in some cases it may be deployed within, uh, your data center. In other cases, a lot of security teams outsource their, uh, tier one SOC out to an MSSP where they may be using agents there. And really it's important to understand the level of autonomy. So this type of, of, um, uh, blueprint, uh, can be deployed, you know, uh, horizontally across, uh, different, uh, verticals or also with partners as well as you're talking about MSSPs, which are very common today with, with SOCs.

Yeah. Awesome. All great points, you know, Peter. Um, I think maybe one thing I want to do is, um, talk a little bit about, um, what we show here is pattern A or pattern B, right? So you did all that planning up above and now those agents, um, are tasked to do various different, uh, actions, uh, within, uh, your SOC and they need access to resources in order to do that.

So, um, everything in the red is all contained, right? Um, and that's all about access enforcement and also audit, which we know is extremely important, uh, within, uh, all of the SOC environment. So we identify a pattern A, which is a more distributed agent to agents, you know, kind of, uh, model or a pattern B, which is more of an orchestrator federated, uh, kind of model. It's almost like the, uh, the agentic control plane, um, you know, that could be plugged into, uh, the SOC.

So on pattern A, uh, the agents may need access to telemetry, alerts, alarms, logs, various different rule sets, uh, for various different pieces of the estate, uh, that it's managing. What are the sources of truth, uh, that the company and the organization has established, uh, and then also access to other agents and MCP, um, you know, servers, uh, for tools that it may need or to be able to spawn off tasks to various different, uh, agents, uh, where on the orchestration side of it, it's at the agents mostly will interact with the agent, with the orchestrator, uh, to request, um, services.

And then the orchestrator will have API calls into these various different pieces on the pattern, uh, a side of things. So, uh, but all of that needs to be documented. All of those interactions need to be, um, you know, tracked, um, observed, um, you know, uh, and monitored, uh, in real time, um, so that audit trails can be done and that when, um, regulators come in compliance, um, reports can be, uh, established, uh, for them.

So, I think with all of those tools available, then it's the agentic, uh, operations loop. And I think Peter, you talked a lot about that, you know, already. I'm not sure if we want to add anything in addition. Um, the only thing I would say is that where we believe like a lot of automation, um, can play, take place and should take place is in one through three. So, it can detect, you know, certain signals in the infrastructure, it can investigate, uh, and then it can propose, uh, fixes, uh, for them.

Um, right. I was going to say the pivot point really is at the decide gate, and that's going to really determine, you know, be based on, uh, the, uh, level of autonomy, uh, for that particular, uh, process within your SOC. And, you know, uh, as far as, uh, do you want to, you know, uh, continue to gather intelligence to understand, you know, what the adversary may be doing, or do you want to, you know, block access now?

And those are some of the things that, um, you know, that, uh, probably Nick, a human will need to decide that. Right. Uh, so I think that that's, that's an area where, um, you know, a false positive could take down legitimate, you know, business traffic. And in some cases it's okay to do that, uh, but at the same time, you know, we don't want to lose forensic evidence.

We don't want to tip off the adversary, uh, cause we're still trying to understand the extent of the event. So these are some of the things that, you know, as we bring practitioners together, uh, that, that, you know, run, you know, large SOCs, uh, they can help us sort of vet and pressure test this, uh, genic operations loop. That's what we're really looking for next.

Yeah. Okay. Um, yeah, I think that's, uh, I think that's basically a, so, um, so again, we have a reference to architecture, uh, for, um, the enabled SOC. We really welcome all the vendors to get involved, um, to kind of, um, share, um, how this architecture can be advanced, uh, within, within the, uh, within the industry and with it for practitioners, uh, to, uh, get involved, to help refine it and, um, add more requirements.

And most importantly, really to be able to kind of lead the industry, um, you know, uh, with this, there'll be a lot of this, um, at the summit, uh, in New York city, October 28th and 29th. Um, so Peter, any, uh, any last words before we go to next steps? No, I think we can, we can move right to the next steps. I think we covered this line really well.

We want you to get involved, you know, as well. Um, and that's really easy to do. Uh, you just come up to the, uh, onoog.net slash collaborative and you'll see this page. Uh, and then on this page, um, you're an IT corporate or collaborative member. Uh, just request to join the collaborative. Just click on that.

Uh, you'll kind of fill out a, you know, a quick form, you know, there, and then, you know, you'll, the last you for like, which working group you want to get involved in. Uh, and that's really it. And so, um, and then we'll get back and, uh, back in touch with you around getting involved. Okay. Peter, any last words?

No, I would say any questions, you know, please reach out. Happy to answer and address any questions that people have. I think the idea is that, you know, if we get more involvement, uh, more practitioners involved specifically, you know, we can really provide something that's going to be really great that you can bring back, uh, to your company. Yeah.

Awesome. Great. Thank you, everyone. Thank you.

The episode summary, topic list, and questions on this page were generated with AI assistance from the episode recording and show notes.