About this episode
How must infrastructure operations change now that frontier AI models can find, and exploit, vulnerabilities in massive, decades old codebases faster than operations teams can respond? In this episode, ONUG co-founder Nick Lippis talks with Cisco's Tom Gillis, SVP and GM of the Infrastructure and Security Group, and Murali Gandluru, VP of Product Management for Data Center Networking, about what they call the summer of turbulence. The stakes are concrete: roughly 80 percent of network outages come from misconfiguration rather than hardware failure, the newest frontier models can comprehend a codebase of a hundred million lines in its entirety and chain together the state interactions that produce exploitable flaws, and an adversary does not even need source code, it can probe a black box firewall until combinations break. Cisco has already shipped critical security updates driven by this shift, and the guests are blunt that more waves are coming.
The episode's core argument is that the traditional operating model, harden a design, then avoid touching it for a year or more, is no longer survivable. Tom Gillis says that model has to change now or organizations will be breached in a significant way.
What we cover in this episode
- From centralized data lakes to federated telemetry. The old model of shipping every alert into one centralized data lake cannot scale, so Cisco is building a decentralized approach with local data repositories, federation built into firewalls and data center switches, and AI agents watching each repository for anomalies.
- Seeing inside containers and VMs from the network. Integrating the Isovalent container networking interface into Nexus Dashboard treats VMs and containers as peers, giving network operators visibility into workloads that used to look like an opaque blob behind a single IP address.
- Frontier models are finding and exploiting real vulnerabilities. New frontier AI models can comprehend a hundred million lines of code in their entirety and chain state interactions across modules, producing a step function increase in discovered vulnerabilities even in hardened devices, and they can exploit a black box system without source code.
- The end of the don't touch it operating model. Instead of qualifying a design and freezing it for a year or more, infrastructure teams should expect significant updates every quarter plus a stream of vulnerabilities in between, handled with lots of small, continuous, easily rolled back changes in the cloud scale out philosophy.
- Live Protect: compensating controls instead of emergency patches. Live Protect is not a patch and does not modify binaries; it applies a laser precise rule, such as blocking a specific process from accessing a specific file, enforced through eBPF with virtually no performance impact and no reboot, starting in monitoring mode before enforcement.
- Organizational change and CI/CD for infrastructure. Repetitive L1 and L2 tasks get automated and roles shift up, from opening and closing tickets to designing the guardrails agents operate within, and one of the world's largest retail banks told Cisco it wants to repave its entire stack every night.
- Smart switches, digital twins, and the path to autonomous operations. The demo shows Cisco Cloud Control and Nexus Dashboard driving inventory, compliance, and shield deployment across thousands of switches, while smart switches with built in DPUs move segmentation in line without firewall hairpinning, and digital twins point toward infrastructure that upgrades itself continuously.
The lines worth sharing
“About 80 percent of network outages come from misconfigurations. The ASICs don't fail. The switches don't fail. The routers don't fail. That's more human error than anything else.”
Nick Lippis
“That model has to change. Nope, that's going to change now, or you're going to get breached in a really significant way.”
Tom Gillis
“Stop the attack, stop the vulnerabilities, don't stop the network.”
Murali Gandluru
Common questions from this episode
What causes most network outages?
Roughly 80 percent of network outages come from misconfiguration and misunderstanding of configurations, not hardware failure: the ASICs, switches, and routers rarely fail on their own. Because misconfiguration is human error, it is a problem well suited to AI, which can drive error rates asymptotically toward zero and keep configurations intact.
How is AI changing vulnerability discovery in network infrastructure?
The newest frontier AI models can ingest and comprehend an entire complex product, on the order of a hundred million lines of code, and string together the chains of state changes across modules that humans miss, so they are finding a step function increase in real, non theoretical vulnerabilities even in hardened devices. They can also exploit vulnerabilities on a black box system without source code by probing until combinations break, so teams should assume adversaries have that capability and expect repeated waves of critical updates rather than a one time patching event.
What is Cisco Live Protect and how is it different from patching?
Live Protect is a compensating control, not a patch: it does not touch or modify binaries. It applies a very precise rule, for example stopping a specific process from accessing a specific file, implemented through eBPF with virtually no performance impact, so shields can be deployed across thousands of running switches in seconds without reboots or downtime. It has a monitoring mode to build trust before enforcement, and it acts as the bridge between more frequent full software updates.
Should infrastructure teams adopt a CI/CD operating model?
The episode argues yes, urgently. The old approach of qualifying a hardened design and then changing it as little as possible for a year or more cannot survive a world of quarterly significant updates plus continuous AI discovered vulnerabilities in between, so infrastructure needs the cloud philosophy of lots of small, continuous, rollback friendly changes. Since what is actually being upgraded is software, infrastructure teams can borrow the lifecycle processes dev teams already use, and one of the largest retail banks has told Cisco it wants to repave its whole stack, switches through web servers, every night.
How do NOC and SOC roles change with agentic operations?
Repetitive L1 and L2 tasks get automated and the roles shift one level up: engineers do less ticket opening and closing and more designing of the guardrails that agents operate within, teaching the system what good looks like instead of firefighting in the middle of the night. The bigger shift is organizational, since siloed platform, fabric, networking, and security teams trading tickets cannot keep pace, and shared visibility across VMs, containers, network, and application context lets multiple teams work from the same view.
Read the complete conversation
Full episode transcript · 54 minutes
that model has to change. Nah, not like, oh, next year we're going to do it differently. Like, nope, that's going to change now, or you're going to get breached in a really significant way. Hi, everyone. I'm your host, Nick Lippis, and welcome to the Built for Trust podcast, where you get to hear from all the folks who are building and shaping AI enterprise infrastructure.
Now, let's get right into it with our guests. Hi, everyone. Welcome. I am, like, just so psyched for this discussion. Before we kind of dive into all the topics, but we're going to really focus on agentic operations and how operations is fundamentally changing, and it has to change.
So, Tom, I know everyone knows you, but just I want you to say a quick hi. Also, Murali, too, you know, after Tom. So, Tom, say hi. Well, thanks, Nick. You know, I didn't know everyone knows me, and if that's actually true, maybe should I be nervous about that?
But in case you don't, Tom Dillis, I'm the general manager for infrastructure and security at Cisco. So, working on all these products and looking at the intersection of those products, right, which is what I think is interesting. Yeah.
We should have that photo of you, like, on stage at Cisco Live, like, you know, a few days ago. It was just, it was just, like, mind-blowing, you know? It's, like, you on stage with, like, 30,000 people looking on. It was just crazy. You know, before we're going to say it again, networking is cool again, right?
And so, this is the biggest Cisco Live we've ever had in history, which is awesome. Yeah. Awesome. Okay. Excellent.
Hi, I'm Murali Ganluru, and I'm in Tom's team. I'm Vice President of Product Management for Data Center Networking. I'm responsible for the data center networking business for Tom. Awesome. Great. Well, okay, let's dive in.
I think I just want to provide a little bit of a reference. I think we all know that our world all around us is changing, and that's really thanks, like, to AI, you know, that things run at a much faster speed. And you can look at lots of different ways in which that manifests itself. I think the one that I look at is just, like, anthropics, like, meteor, meteor, meteor, meteor, I'm totally mispronouncing that.
Like, thank you. You know, that growth, you know, just shows consumption, you know, and that we're still compute-bound, and we're not really demand-bound whatsoever. But that also, you know, says that, you know, especially in large enterprises, like, you know, the dev teams are the ones who are really kind of consuming that and using those tools.
And so code is shipping itself. Workloads are able to kind of spin up and spin down much faster than they ever were before. Things are moving more autonomously. You know, than they ever had. And I think the issue, and those are all really good things because it allows business to move at the speed of the market.
But the problem is around operations. And operations still does not have the tools in order to keep up. And so we're going to really focus in and hit on that, you know, right now. So I think the thesis that we're going to have is that all of us who are infrastructure professionals, this is the biggest shift that I've ever seen in the last 30 years.
But the way that we design, build, and run this infrastructure is going to fundamentally change. And I think probably maybe the best example of that, and Tom, you kind of iterated this, like, when we were doing the prep, And that is, like, about 80% of network outages, you know, come really from misconfigurations and, you know, maybe not understanding the configuration so right.
So that's more human error than anything else. Like, the ASICs doesn't fail. The switches don't fail. The routers don't fail. And that's also a problem that is really well-suited for AI. So we'll use AI, and we'll talk about that here, in order for us to asymptotically start to go to zero in terms of errors and to keep the configurations, you know, intact and also keep us moving more and more towards autonomous.
So the arc that we're going to talk about, we're going to talk about an architectural shift that's happening in operations. We're going to then talk about the operational piece of that. Then Murali's going to do a really good demo for us. So we're going to do a demo and show you how this all kind of comes together, and then we'll talk a little bit about futures.
Sounds good, everyone? Sounds good. Okay, great. So let's first maybe start architecturally. So the way that we've always built infrastructure to monitor it and to run it, this lifecycle management, has been that we have all these emitters all across the infrastructure, and they emit and they, you know, you know, it's basically, it's an alert, an alarm, you know, an event that happened.
And those get distributed up into a centralized data lake, and then something breaks, and then a team is assembled, which takes a long time just to assemble it. And then it's almost like somebody throws a needle into a haystack, and they say, go fetch, go find it.
That's been the model that we've really had for like the last, at least 15 of my 20 years. That model is now fundamentally changing. We're being, we're moving more distributed, moving intelligence further down, not only at the switch level, but at the port level.
So I thought, you know, Tom, that would be a really good place for you to kind of start, you know, you know, just talking about this shift around the operational infrastructure for lifecycle management.
Yeah. So, you know, what you were alluding to there, Nick, you know, the way I'll phrase it is like, how do you know you have a problem?
And the data's there. Again, it's like the perfect analogy, a needle in the haystack, that there is a needle in the haystack. It's just a gigantic haystack, like way, way, way, way too big to shove all of that data into one place to look at it.
So this was the impetus behind the acquisition of Splunk with a combination of Cisco plus Splunk is let's create a decentralized model where we can have lots of local data repositories and then we'll federate across those.
So that's going to provide, and what's cool about this, this isn't like, oh, you know, someday, like we have built that federation into our firewalls, which is like typically for most customers, firewalls, the number one source of ingestion.
But we've also built it into our data center switches. So I'm really into him and put this into the Nexus dashboard, Nexus One. And so, and then we've integrated it into high surveillance.
So we can see right down to the process level, if things are something that's changed, we're going to be able to instrument that, capture that, see that.
Now, we saw the data ingestion problem, but it's still a ton of data. And so, the point of this broadcast is that we couldn't make this workable without AI.
So having little AI agents that are sitting on those local data repositories and being like, hey, wait a minute, this looks a little off to me.
I think something's wrong. You know, that's, you know, it's coming. That's still like kind of just a little bit around the corner.
The federation capabilities here. But that's going to transform how we understand infrastructure. Did something go wrong? Is the application misbehaving? So we're going to have a much finer level of insight there.
And I would argue that's kind of step zero or step one on this journey to bringing CICD or cloud operating principles to all forms of infrastructure.
Yeah. So that's... You know what... Go ahead. I'm sorry. My mind's kind of spinning from what you just said.
You know, one... All right. So great. So we have to, you know, agents like distributed like throughout the infrastructure.
Yeah. Those agents, they tend to be in containers. Right? And so... And ISOVALENT, you know, allows you to see the containers.
Not just like on the infrastructure side, but any kind of containers. And so what I think is really also unique in here is that the network has essentially for the longest time been opaque.
Right? You really can't see into things. Now you can. And I think that's really fundamental, you know, to like this shift.
It's both that we're distributing intelligence, operational intelligence, back into the infrastructure. And now there are tools for us to actually see what's inside the infrastructure.
So... That was something we debuted at Cisco Live just last week. And so we've integrated, as you said, the ISOVALENT container network interface into the Nexus 1 dashboard.
And that gives us complete visibility into a workload that is either a VM or a container. Like, they're treated as peers.
And that, you know, that's not a startlingly new idea. But it's done, right? And it's the platform.
And I think what makes it interesting is that, you know, I'll argue that Nexus is the most widely deployed data center platform in the industry.
and ISOVALENT is the most widely deployed container networking. So there's not a lot of friction to bringing those two things together.
Yeah, but visibility is, like, you know, that's where we start on this journey.
Now, the second thing on this journey that we need to really think about is the operating cadence for the infrastructure underneath the application.
Switches, routers, firewalls, like stuff that, you know, these are generally considered to be high performance, highly engineered, you know, very, very, very tight devices.
And, you know, if I'm being candid, most customers have a view of, like, I qualify a design, I test that design rigorously, deploy that design, and I don't want to change that design.
I don't want to, like, breathe on it. Yeah. Yeah, don't breathe on it, right?
Don't touch it, right? So it's like, you go for a year or 18 months or whatever, you know, and you make maybe little bits and pieces, but you're trying to change as little as possible because you're balancing availability versus, you know, why are you making this change?
You know, security versus availability. And so that's always been a, you know, balance that the folks on this broadcast have to listen to.
And that balance just went like this. I don't know if everyone can see, but it just, it just tipped by 100x, right?
And you know what I'm talking about. Yeah. And if we could just add, we're actually, to your point, Tom, we're actually bridging worlds that used to be separate.
The VM and containers and the network and app context into one framework. That's what you're going to be announced. Yes. Well, there's a meta theme, Murley, that we're going to come through in this podcast, which is a lot of what we're doing is going to require and also is going to precipitate organizational change because the old world where I've got, you know, sort of the platform team and they do the catering networking and I've got the fabric team and they set the, you know, set the fabric up and they, you know, trade tickets to go back and forth.
That has to change. That has to change. Right? And, and, you know, the other thing that we're going to show is, is that the pace at which we update the infrastructure, it has to change.
Right? And Nick, you and I did a podcast on this. No, no, no, the last one we did recently, we were talking about Mythos and the change, the impact that that's having on software, not Cisco products, any form of software.
Any software. Yeah. Any software. Yeah. It's a step function change and, you know, not everyone's got their hands on Mythos yet, but like, you know, enough people are getting on it that, you know, we're all kind of hearing the same cry like, whoa.
Right? This is not just, you know, marketing hype. Like, this is significant.
So, maybe it's probably worth recapping in case anyone didn't listen to the last podcast, but I encourage you to go back and listen because it was good.
It was really great. Yeah. Here's, here's my take on Mythos and you know, you may have a different view, but we've all been using AI coding tools for a long time.
And what I would observe is that for, for new projects, we'd get this 20X, this surge in productivity. The models are incredible.
They could like, understand the whole product and spit out like enormous amounts of functionality really, really quickly.
But for Nexus Switch, Catalyst Switch, Cisco Firewall, they're very complex products that have been developed over a decade and they have all kinds of layers of complexity and, and sophistication and the models couldn't comprehend them in their entirety.
so we weren't getting that 20X lift in productivity. You know, we're getting something quite modest and that's what's changed.
So these new frontier models and it's not just open, it's not just the Anthropic, right? Like we have open AI models as well and we're seeing a similar sort of impact.
They can understand the model in its entirety. A hundred million lines of code and this thing can slurp it all in and understand it and that's what makes them so unbelievably effective at finding these vulnerabilities because it's this kind of weird chain if I change this state in this module and then it talks to this thing which sends a signal over to that thing and it goes to this thing and the computers can string that logic together way, way, way, way better than a human.
Yeah, it is kind of mind-boggling about just all of the interactions that happen in any piece of software, you know, but like in networking in particular, you know, it's like that's, that's what networking is all about is, you know, interactions, you know, between, between so many different points, you know.
Lots and lots of states, right? Yes, exactly right. Lots of states. Yeah. So as a result, these models are finding vulnerabilities that, you know, we have to be careful, we can't tell you the numbers, but it's a step function from what it was before, right?
You know, if you look at the, and these are Cisco devices, which these are very hardened devices. Yeah.
So if you think about, you know, the customer's application logic, that's going to have, like, thousands of vulnerabilities.
So it's creating a bit of havoc in the industry, if I'm being frank, right? This is all anyone wanted to talk about in Cisco Live.
Yeah. Like, what do we do? How bad is bad? You know, what does it look like?
And, you know, I call it the summer of turbulence, where we have rolled out a few security, critical security updates.
And when we say these are critical, we mean they are critical. Like, these vulnerabilities are real. They're not theoretical.
And here's the other part of those mythos things, is that not only will it find the vulnerabilities, it'll exploit the vulnerabilities. Yeah.
And Anthropix has been incredibly responsible about trying to put guardrails, you know, to do the right thing to limit that.
But let's just make the working assumption that the adversary has a capability that it'll exploit these vulnerabilities.
It'll find it to exploit these vulnerabilities on a black box system. It doesn't need source code. It can just poke to your firewall long enough and it's going to find these combinations and make them break.
It's more effective if it knows what it's doing with the source code. So, take out all the other, you know, everyone is in this kind of panic mode where they're going to run around and patch everything.
It's log4j times a thousand. Every vendor, everywhere, is coming in with emergency updated.
So, there's an initial reaction of like, how am I going to update every single system? Every firewall, every low balancer, every switch, if it, you know, go through and make it a prioritized list and we're working with our customers to help make that prioritized list.
this round for a second. Yeah. I'm going to help. You know, what about the next round? Right?
Like, you think that we're going to like, you know, do this once? It's not one and done. There's going to be another wave of this in September and then another wave again and it's just, this is the new world.
Also, one of the things is that the key to this is being able to do this in a repetitive manner at scale and so building an infrastructure that supports the ability to disseminate those shields and do that in a predictable manner and in a way which the operations team provides comfort to them as opposed to alarming them is very critical and that's what essentially stop the attack, stop the vulnerabilities, don't stop the network.
Yeah. So, this is where Light Protect comes in. Yeah. I think Amethyst is going to be kind of a catalyst for lots of different, you know, things.
You know, it's like not just around how do you automate kind of the patch management but also how do you even approach security now, right?
So, it's like, you know, it used to be, you know, modes that we tried to build modes with like firewalls.
Yeah, we went to cloud that kind of really expanded that, you know, that threat vector and now here it's like every piece of software is almost getting rewritten you know, and like from what I learned when we did AI Networking Summit a couple weeks ago and I'm sure you heard it loud and clear as well over at Cisco Live was that there is no working models on how do I deal with this stuff, you know, and how do I deal with it with the velocity that people are imagining that's coming at them.
What are the tools that they can use? How do we organize ourselves in order to deal with this?
So I think it's a big open well, there's just a huge amount of interest in how do I approach this, you know, so, and I know that, you know, Tom, you talked about that.
We own a community, right? Like, everyone needs answers, right? And the truth is, we're figuring a lot of this stuff out on the fly, but we do have a bunch of answers today that are really important we want to cover in this broadcast.
So first is, like, the old model where I harden that, it's particularly true in the data center, but it's true generally of all infrastructure, but like, I harden my data center design and then I try not to touch it once a year, right, for a major upgrade.
Yeah. That model has to change. Nah. Not like, oh, next year we're going to do it differently.
Like, nope, that's going to change now or you're going to get breached in a really significant way.
Like, it's that kind of urgency. Never miss a crisis to drive meaningful change.
So we as your vendor are building tools that are going to allow you to adopt a different operating model.
But where I need ONUG and the leaders at ONUG to lean in is understand that this change is afoot.
And driving change in an organization, especially among networking professionals, is hard because when you're on uptime, you're super risk-averse, right?
Like, that's how our community operates. But we've got to move differently.
You know, it's a good thing to have a bit... I'm sorry, Tom, finish your thought.
Bring in another one. and also, this community, network engineers, they are so risk-averse, which they should be, and they have to be because the network touches everything.
But also, it takes them a long time to absorb new technologies and innovations unless they're made really simple and really easy for them to do.
But I think that model is gone. It's almost like they can't rely upon just waiting a little bit until they can kind of wrap their mind around it.
And so, I think you're in a really uncomfortable place. And it's going to get more and more uncomfortable for them, unfortunately.
Totally agree. Well, Merle's going to show us a demo that hopefully can assuage some of these concerns and that the tools that we're building, I think, are a pretty big step forward and I think they're pretty easy to use.
But I don't want to trivialize because I would say the situation we're in, we together, customers, vendors, everyone, the whole industry, it's going to be rough for the next 12 months.
There's going to be disruptions and there's going to be headlines and it's going to be bad.
So let's try to minimize the problems. And there's again, steps we can take on that front too, right?
How do we know the problems? But let's talk about kind of this new philosophy.
Okay? So the old philosophy was nuclear-hardened it, change it as little as possible.
Operational, yeah. Yeah, right? The new philosophy is, I'm going to argue, is one that's well understood in the industry and it's kind of the scale out versus scale up systems.
It's the view in the public cloud of like, look, you can't achieve the super high availability, so design around it and make lots of little tiny changes on a continuous basis and do it in a way where a change is not going to be massively disruptive and so you can easily roll back and that's the way we're going to manage infrastructure, which I know I'm sure there's people out there that are at heart to skip the beat, like, wait, what?
But we've taken the first step there and we can prove it.
So the way we think about it, and this is a place where I'm really proud of the team at Cisco because we're not on our heels here, we're on our toes.
We've been thinking about this and you and I have been working on this for years to get to this state, right?
So yay. But the way we look at it is we're going to move to a world of more frequent updates.
Again, focusing on the data center where things are the scariest, you know, it's probably going to be quarterly, there's going to be an update from us and from your other vendors every quarter that is significant.
Like you have to adopt it. It can't be like, oh, we'll catch the next one, right?
And so just do the math on that. How many man hours does it take you to process and manage that?
What's worse is even though with these quarterly updates, the number of vulnerabilities which, if you're watching the video, I'm moving my hands.
Again, I don't say numbers, but I'm using the big indicator, right?
Like there's going to be tons and tons of vulnerabilities in between those quarterly updates that have to be dealt with because they will be exploited by these automated attackers.
And that's where we've done something really interesting. So Cisco has introduced LiveProtect, and Nick, you were not faced with this last time, but I want to say it again.
So LiveProtect is a compensated control. It's not a patch.
We're not touching and modifying the binaries, and that's important.
But it's very, very precise. Like laser precise is a very specific control that says if you see this particular process trying to access this particular file, stop it.
Right? And when you could be that precise, there's two big advantages.
The first is you're going to have a really low false positive rate.
It's not zero, but I'm working on IPS systems forever, right?
It's not going to look anything like an IPS. Like it's going to be this really, really, really, like, man, if we get a hit on that rule, we want to have eyes on the glass and see what the heck is happening.
And the second is that we can implement this using EBPF, and I won't go into what that is, but we refer back to the previous broadcast, extended Berkeley packet filter, which is in Linux, and our competitors could do the same thing, and they probably should, we can implement this with virtually no performance impact.
And again, my buddy Murley over here, we've been doing this for, we had it in production since last December.
Back me up. Performance implications of putting a sheet on a running switch.
What happens to the switch? It keeps running as is the same performance.
So that's really key. That is so big because we're used to major patches and upgrades where you actually have to do it on a weekend.
So you're basically talking about being able to do these significant patches without getting a hit, without getting a performance hit, without having downtime.
Without stopping the system. You put it on the system while the system is running.
Like in big, giant monster data centers, we're doing this. We had at Cisco Live, Surface Now, was on stage with Murley and I, and they were saying we have to run it on an entire production data center.
We're putting shields, on these systems without rebooting or modifying the systems.
So the precise controls, right? Yeah. So we think that's a really big step forward and we think the whole industry is going to have to go this way.
Because otherwise, what do you do? You can't up yet every week.
Yeah, really. You know, it's like we're going to shut down the network to upgrade it, you know, and turn it back on.
Yeah, that's a non-starter, right? Yeah. Okay. So I think all of that is really good, but I want to, you know, I do have a little bit of a format, you know, I want to go back to.
So operationally, so, you know, it's like, so normally the way that we would deal with operations, right?
We have networking teams, we have security teams, you know, we have the application teams, we have the development teams, right?
And so in smaller companies, we had DevOps, right? And we're like, you know, you can, you know, where the developers weren't just throwing over something new over to the operations, you know, team, but they were actually not responsible for the operations of it.
That never really took off in the large enterprise. There was still, you know, you got big dev teams and you have big operation teams because the systems are so big and they're so complex.
But that's just not going to scale anymore because like even like, you know, a good example like at JPMC, you know, when they need to make a change to the infrastructure, they need a hundred people to kind of sign off on that.
You know, that won't fly anymore. Networking and security has always been, you know, the technology has been combining those groups and creating more and more overlap around those groups.
We've seen organizational shifts happen before, especially in our industry where we had voice groups and we had data groups and they became one group, you know, as well.
So I think now, in order to be really prepared and to be able to deal with kind of a much faster life cycle management, you know, thanks to like, you know, not only just the LLMs discovering, you know, the need for patching, but also just the breadth and the scale that the dev teams are writing applications, what are you all thinking about in terms of like, how does this, how do we, how should kind of management teams start to think about operations?
Like, we had separate NOCs, separate SOCs, you have L1 and L2 engineers, a lot of that is looking to be automated so they can move up, you know?
Yeah. Have you kind of like, you know, we should provide some talking points for the community to start thinking about how should we really be thinking about organizational-wise the infrastructure teams?
Yeah, it's a great question. You know, and I think the honest answer is we're going to, we together have to figure out some of this stuff out.
Now, the tooling, the procedures, the philosophies that have been in place for decades are changing and they're being driven by this mythos change urgently, not like, you know, over the next few years, over the next few months.
Yeah. And so it's likely that the organizations are going to have to change too, and that of course will be, you know, slower.
Will it look exactly like the DevOps model? I don't know. But I will tell you, I was talking to the infrastructure team for one of the largest retail banks in the world, and they were saying, kind of along the lines, Nick, of we want the CICD model, we want to repave the whole thing every night, the whole stack.
Switches, routers, firewalls, load balancers, you know, database, web server, the whole caboodle.
Like, we want to spit it out, repave the whole thing. And that's technically within reach.
You know, it's not completely out of the question. Now, there might be some nuance on exactly how we do that.
But yeah. it's interesting on the L1, L2 part of the operations, you know, it's the repetitive tasks that get automated, and the roles shifts one level up.
Essentially, you're doing less of ticket opening and closing and more designing of the guardrails that agents operate within.
so, you know, where engineers would firefight past midnight, 2 a.m., 3 a.m. They become the ones teaching the system what good looks like, and that system will learn, and it is able to operate at much wider and bigger scale.
Now, I think the whole day zero, day one, day two has to be really rethought, you know, now, and especially day two, all right, you know, on how much that is going to be now automated and the tools that we're going to get to do that kind of automation.
And I think, you know, it's like, this might be a good segue for you, Merle, to do the demo, you know, but I think maybe one point before we do that, you know, Tom, the more I think about it and the more kind of discussions that we've been having is that, you know, a CICD kind of model for infrastructure makes sense, you know, because, you know, really when it comes down to it, like, we, clearly there's a lot of hardware like in networking infrastructure and security infrastructure, but the main point of that is the software, and that's really what's being upgraded and, you know, and patched, and we have all the processes in the dev teams to do that, and if that's the case, then it would be interesting, okay, do, and I've seen this actually, you know, within the owner community is that the dev teams now are trying to make changes to the infrastructure, and the infrastructure of people are saying, time out, you know, it's like, so there's a rural responsibility thing.
It's a battleground on that, right, but now it's the whole infrastructure stack, but I do agree with Merle, and Nick, you were saying this, is that, you know, it's very clear that we can use agentic applications to either fully automate or highly automate that upgrade process, and I couple that to this live protect thing, so live protect is the bridge between these updates, but the updates are going to be more frequent enough than you can do about it, and we have to together figure out how can customers ingest these updates, you know, much more frequently, and that's the ICD philosophy.
Yeah, well, you know, you find, you know, I actually, Tom, you just kind of tied two, you know, dots to me, you know, that are really important, you know, one is that Mythos is going to kind of drive infrastructure teams into a more CICD kind of model, because the upgrades, it's almost like we're patching, you know, we're going to have a huge amount of patching that's going on, but that's like, it's the same kind of process of like actually committing new software, you know, as well, you know, so it's like, you know, so they're forced now to do this, you know, and so they're going to have to get that, you know, that skill set and that mental model, you know, as well.
So, so, Murali, this might be a good time now. Okay, let's think about, okay, agentic and how agentic is really going to be kind of forcing that change and the kind of tools that we're now have access to.
Absolutely. As you meant to imagine, everything starts from Cisco cloud control.
And Cisco cloud control essentially is the platform that sees across compute, network, security, storage, and then across domains, data center, campus, branch, the backbone, WAN.
And essentially, it acts on what it sees across these domains. That's what agentic ops effectively means in practice.
So when you look at this particular dashboard, you see sites, you see multiple locations, and that was the topology view.
But then if you also look at the inventory view here, you get to see what kind of security exposure do you have, what kind of life cycle horizon, are they any of LDOS environments, are they past LDOS, so the compliance and governance, the compliance posture further around policy and control as well.
So everything is visible here. early in the broadcast we talked about step one, really just showed a step one, which is like, what do I got?
What do I got? So inventory, yeah. Most people don't really know, like, is this out of support, and what's a version of this?
So we've taken the big step to make that, answering that question easier, and you just saw it.
Can I ask a couple of questions on this one? So that's the inventory piece, right?
Okay, so here, so compliance exposure, so I would think that there's a lot of kind of intent-based engine that's kind of going to drive this, so like every industry is going to be a little bit different around their compliance, so there's got to be modules associated with that.
The life cycle horizon, explain that, what do you mean by that? What it means is basically is that you've got a particular set of switches, firewalls, you've got a bunch of servers in your environment, they've got software.
Those software and the hardware associated with it, essentially are they in support?
How many days of support is left? Are they past the last day of support, which is short form LDOS?
Are they within the 90-day window? You get a view of your risk of compliance, our being out of compliance, what's that window look like?
Okay, great. And then security exposure. So is that like systems that have not been patched?
Yes. We have a lot of p-certs, vulnerability, CVEs, all the time showing up, right?
And particularly as Tom just said, in the mythos world, you're going to see a lot more.
And that's just not mythos alone. There's just a lot more other environments.
So this gives you a clear idea about what are your clean assets, what are your high-risk assets, and what's somewhat of a medium-term risk assets.
So I would think in a mythos world, that orange that says high is going to be almost like the entire circle that you're not patching.
Yeah, it's about hardening your environment. And this allows for us to, as Tom said, the philosophies, changing the philosophy around how we think about support, how we think about hardening your environment infrastructure, and what's that cadence look like.
Yeah. Okay, great. Perfect. So when you think about the topology that typical network infrastructure or even cross-domain infrastructure operators look at, you've got a view of the devices, number of van appliances, number of sites, locations, campuses, all those in one place.
And when you provide that kind of a view, you also are able to see what's your health status in all of these, right?
Am I seeing warnings in my data center switches? What kind of unhealthy, poor compliance or other kinds of things exist in that environment?
So let's go into Nexus dashboard. Particularly in this case, Nexus dashboard looks at your end entire data center network asset base.
And typical view of a network topology looks something like this. And remember, we launched from Cisco Cloud Control into this.
Similarly, you could do the same thing in campus. You've got a spine-leaf design, you've got connections out into the internet, but guess what?
You don't have one big viewpoint. And that's somewhat by design, which is what's behind that Kubernetes AI cluster?
So in a world where workloads, I think over 66% of the workloads are AI, of AI centric workloads are on Kubernetes, you've got to have your operations team know how and where issues occur and what kinds of application profiles, transport profiles look like.
And so that's what we've fundamentally changed with the ISOvalent technologies that we acquired.
And now you're able to see inside the application from your network operations window itself.
So you can see it's talking to whom. So this is showing how, so if I look in there, web app database and a namespace, and then here we have agents.
So, okay, great. And that's just by clicking on that one switch. Absolutely.
And essentially you've got configuration that pushes down all the way into those CNIs.
And now you have the ability to not only get visibility, you've got isolation capabilities and control as well.
Tom? We talked about this at the beginning of the broadcast, bridging worlds.
Like today, the DevOps team and the platform team has visibility into the connectivity and the networking for the containers, but the fabric lives in a totally different domain.
And conversely, looking at the fabric, a giant complex Kubernetes cluster, like all the purple stuff, looks like a little tiny blob, just an IP address.
And so we now neutralize that. So at least we know what we're dealing with.
Step zero. Well, yeah, that's true. And the bridge of a gapping between groups is visibility.
inventory and visibility, and then management. So yeah, this is great.
Continuing. The next step from visibility is essentially about policy. Right?
The platform teams, the one that's responsible for the Kubernetes infrastructure, they typically define the policy and now we have a way for that policy to flow through into the network as well.
And now you are able to take a look at how that particular, in this particular case, the AI cluster policy enables a particular traffic flow.
You see the green dots flowing across, you flow across from the app all the way across out into the internet and probably likely ending up in another location or accessing some kind of application or data environment in another location.
That's really cool. So like the box. Yeah. if you've got an app that has containers and VMs in it, so you've got an AI front end but the database is running in a VM, you can now have a PCI requirement for segmenting that app across the VMs and the containers, single view, single policy, single management troubleshoot of it.
Yeah. And if I could hit it late. Yeah. And I guess the only thing I was just saying, just like make sure I got it.
So like the box on the bottom, you know, that's basically just showing the path and the packets within kind of the containers.
And then, so that's a logical view. And then on top of that, the green is the physical view.
Like I'm really how packets are transported. Yes. Exactly.
Exactly. And by the way, all of this hides a key element of enablement and innovation that has happened at the data plane level from the CNI perspective as well as on the switches, which is EVP and VXLAN all the way adjacent to the workloads.
That's now implemented on iSurveillant as well. So it's seamless technology.
That allows you to carry the metadata associated with the policy all the way across as well.
Across one fabric, across sites, locations. Now you can see how it scales.
How policy information also scales across environments. Awesome. Okay.
It's good to you. Yeah. And the third big part is what we spent some time talking about it here is you've got to make sure that your environment especially particularly in the current world we are living in or entering is about making sure your devices first, you have visibility into how out of compliance or what kind of advisories are associated with your networks.
And since we talked about the framework of how you would actually apply policies, look at this.
You can see that you have a whole bunch of switches, all the reds there.
Effectively, you've got all those devices needing to be dealing with security advisories.
So now what we actually can do is we enable LiveProtect.
LiveProtect now is a capability that has two modes. One is monitoring.
We want to earn your trust by first making you get comfortable with the notion of saying, okay, I have visibility into all of these devices, inventories, assets that exist in this fabric that essentially allows me to see which ones are out of compliance or which ones need advisories.
Once you're comfortable with it, now you can actually go into LiveProtect enforcement mode where you can actually use the framework of LiveProtect to go and apply the compensatory controls that Tom talked about earlier.
this allows you to in real time be able to stop the attack, stop the vulnerabilities and in Tom's example it was a granular file all the way down and you can now build a planned approach to patching or a full software upgrade, something that was not possible earlier, right Tom?
Yeah, that really is kind of fascinating. So when I'm looking at this graphic and I'm thinking, okay, when I'm thinking about compliance, I'm thinking, okay, well maybe these have not been upgraded or updated so they're patches that are missing so they're out of compliance, which is huge.
One of the biggest hassles for any kind of networking team is when you're going through you have a regulator and you have to prove that you're in compliance.
The rock fetches are just horrendous. So yeah, this is great probably.
So continue. So we had policy and now we have advisory. Then you apply this across all of these environments and essentially once you do that, it's life protect shield status shows that it's protected and obviously every time there is an approach or an attack of trying to exploit that vulnerability, the hit counters that you see here, they will increase.
Right now there is no hit, so it's zero, but the hit counters will increase.
And Nick, this is what we talked about earlier, which is like it needs to be just brain dead easy and it's that easy.
It takes a few seconds to deploy shields across an entire fleet of thousands of switches.
Boom, push, go. So we want to make it. So we put a human in the loop, as Marilie said, it's got a passive bode, non-disruptive, but once you get comfortable that these things are very, very, very tuned, they're intended to be just automatically deployed.
And that's a big, big step forward. Big step. Yeah, so a shield is an agent.
Is that right? You know? A shield, you could think of a shield as a policy.
So the shield is the policy that says, don't let this process talk to this one file.
It's an expression of what not to do. The agent is already baked in. The thing that does that enforcement is already built into our operating system.
So Marilie's team put it into the NXOS that we were shipping last December.
We announced it on our SE-WAN. We're putting it into our router switches, but it's going to go across the whole Cisco portfolio.
So, yeah. And then... So these are like roles. in place. It's like a wall.
Yes, exactly. Awesome. And the consistency of approach, Nick. The consistency of approach.
Remember, we started at the same Cisco cloud control that the entire portfolio has.
The approach to LifeProtect, again, same across all of the environments. So there's that theme and it's the same product being used across multiple environments as well.
So I want to quickly go and point out a few things and just shift gears.
Another element of secure networking that we're fundamentally disrupting is around how you think about segmentation.
Typically, traditionally, this is how an architecture looks like. You've got spine-leaf designs and you've got firewalls hanging off of a border leaf or a border gateway and essentially all the traffic has to go hairpin to the firewall, third party even.
And it works. It's totally fine. What that results in is essentially a drop in the ability to do higher latency numbers and essentially that means that that is actually sometimes not compliant, particularly for latency-sensitive workloads like AI workloads, that could be a problem.
So what we've introduced is the smart switches and these smart switches essentially are any Nexus 9K today, but again across the whole portfolio, Nexus 9K switches having DPUs built on them.
And you can actually apply and turn on services that enable segmentation that you can now do in-line instead of having to do hairpinning to go to a firewall and have that one and then come back and then go out the same port.
this allows for much better in-line approach to security and that allows for a much, much higher throughput, higher amount of firewalling as we call it, 20% in this case, traffic filtration and a much, much better latency as well.
Yeah. So this kind of like, you know, brings essentially the firewall down to the port level.
right? And so there's no hairpinning, there's no going back to like, you know, some firewall.
It's fully distributed into the fabric. Particularly at the layer 3, layer 4 level where, you know, the network devices can do that at scale, in-line.
That's where essentially this is most effective and that's where you see that the segmentation capabilities that would have sat on the firewall, you can now bring in-house.
So now the firewall is not burdened with, or they can focus more on higher layer, application layer services and stuff like that.
Things like NAT, at scale, all of these things can happen now at the smart switch itself, in-line.
Okay, awesome. What else you got? And so, yeah, and that's effectively it and back to you, Tom.
That's great. So I think like with that demo is that, you know, a lot of, you know, this visibility is afforded by kind of the agents being built within the switch, you know, and then you're kind of building on top of that, you know, is, you know, really, you know, kind of what I picked up from that.
And now, and along with that visibility, then you now have multiple teams that have different roles that can actually be looking at the same thing.
And that's, you know, and that's huge. That's really huge, especially around resolution or being able to make sure like you are kind of, you can move, the infrastructure can move at the pace of the application developers, you know, now, you know, as well.
And the ability to kind of secure and harden the infrastructure with keeping up with patch management, you know, also.
CICD, right? That's the goal, you know, but Nick, as we said at the beginning, the tools that we're building, the way we're building these products are changing dramatically.
But I think the bigger change has to happen is the, you know, organizational change and the operating philosophy of the entire network community.
Never miss a good crisis to drive change. Well, we got one, right?
And, you know, I think together we're going to have to muddle through this pretty quickly.
So if we think about, you know, also like, the progression, right?
Because like, to me, this is like kind of the first step around autonomous infrastructure, right?
So it's like the first step is really having much greater visibility, the integration of the technologies, networking technology and security technology, and also agentic, you know, being built, you know, deep within the infrastructure, the architectural shift that we talked about as well.
There's some amount of automation that that you're building into this, you know, but where do we go from here?
You know, how do we, do we, are we on a path towards autonomous infrastructure, you know, where the infrastructure can run itself, you know, and it can fix itself, and it can patch itself, you know?
Yeah. I think so. I think, I think, you know, within months, you know, and, you know, maybe it's 18 months, but, but, but it's, it's, it's right at hand.
So the other thing we announced that Cisco Live was a digital twinning technology, right?
And if you think about our problem of upgrading network infrastructure, an AI agent that's observing a system, like a switch, running in an environment and understanding that environment, and managing the upgrade process and just say, did something change or did it break?
AI's going to be incredibly effective at that. And so couple of that with what we're doing with LiveProtect, LiveProtect is the bridge in between these upgrades and then we're working on digital twinning and AI agents that are going to drive lots of little changes and it's, it's going to look like painting the Golden Gate bridge, you know, like we could shut the bridge down, paint the whole thing, right?
And the stop of traffic, right? And, or we could have a tiny little crew out there and this is what we did, they do in San Francisco.
It starts on one end of the bridge and they just do, they just kind of make their way all the way across.
They flip to the other side and then they make their way back and they just keep going, right?
Like all the time, there's a paint crew up there. It's amazing. They do it in a non-disruptive way.
Yeah, that's the model and we have a lot of work to do on the tools still to get that to realization, but we're racing the event done and there's no other way.
There's no other answer to this mythos problem. Well, you know also too that dawns on me is that we kind of knew that it was coming but like a massive infrastructure upgrade because like you can't just like go backwards or basically put a lot of these features and these functions onto older equipment.
It's like you need to upgrade all the equipment with this. But the thing that's happening now though is that you upgrade you become AI enabled and you get autonomous operations with it.
So it's the value proposition is really strong. That's right.
When we get through the summer of hell and it may turn into the year of hell I don't know, but when we get through this turbulence that we're heading into the world will emerge to it's going to be way better.
Who likes upgrading firewalls and load balancers and switches is very cumbersome and it's not really valuable at work.
So if we can have it where it's either fully automated or highly automated and the infrastructure underneath will be way more resilient that's a good outcome and I think that's at hand.
I think we can achieve that together with that's why we're here.
Well, that's great. You know, well, the community needs the help, you know, they clearly need to be they're on this path, you know, towards, hey, we've got to like upgrade this infrastructure and we've got to be able to support, you know, the business.
But I think that's a good place for us to like, you know, stop, you know, yeah, wrap up, you know, I thought this was like really, really great.
So, you know, I think it was the first really kind of like hardball around agentic ops and, you know, especially, you know, you know, Tom Amorelli and the work that you all have been doing over the last couple of years and how that's coming to fruition.
And I think the trend lines that we're now just going to start to see and really, I think my big takeaway from this was this whole patch, this summer of hell that we're going to be going through is also going to give teams more AI tools and skill sets but also it's going to force an organizational change you know, around software lifecycle management for infrastructure.
So, anyway, this was great. Thank you guys. Always a pleasure. Thank you.
Thanks Tom. Thanks Morelli. Thanks everyone for plugging in.
The episode summary, topic list, and questions on this page were generated with AI assistance from the episode recording and show notes.