About this episode
Why has container networking become the foundation of AI infrastructure? In this episode, ONUG co-founder Nick Lippis puts that question to the people who built it: Thomas Graf, CTO of Cisco Security and co-founder of Isovalent, the company behind Cilium, and Craig Connors, CTO of Cisco's Infrastructure and Security Group, who was instrumental in Cisco's acquisition of Isovalent. The framing comes from the ONUG community itself, which voted that the future of enterprise AI is hybrid, spanning open and closed models across public cloud and private data centers. The logical layer of that multi-cloud, hybrid world lives in containers, which makes the ability to network and secure containers across multiple trust domains a core requirement of AI infrastructure.
The episode's central argument is that identity, not network addressing, has to be the basis of security in this world. Cilium was designed around that principle: policy binds to a workload's identity rather than its IP address, because containers, and now AI agents, are volatile and come and go constantly. That design, powered by eBPF in the Linux kernel, was adopted by the major cloud providers' managed Kubernetes services and became the default plumbing for most Kubernetes deployments, and it now underpins extreme workloads such as OpenAI's massive learning clusters. The stakes rise with agentic AI, because as Craig Connors warns, an agent given a goal will actively circumvent security policies, changing IPs, protocols, and ports to get its task done.
What we cover in this episode
- Why Cisco acquired Isovalent. Craig Connors saw two shifts converging: enterprises moving to hybrid cloud and Kubernetes becoming the substrate for AI networking, and Isovalent led both container networking with Cilium and eBPF-based runtime security with Tetragon.
- How Cilium became the default Kubernetes networking layer. Thomas Graf's team built Cilium on eBPF around scale and identity-based security, and once Google baked it into its managed Kubernetes engine, with the other major cloud providers following, it shipped by default with most Kubernetes distributions.
- Rethinking the CMDB for containers and agents. No single database can track infrastructure that spins up and disappears constantly, so the CMDB becomes an AI-augmented data layer composed from multiple sources with an identity layer on top, extending to non-human identities for agents and endpoints.
- OpenAI as a proving ground for container networking at scale. OpenAI's learning clusters combine massive Kubernetes scale, server links in the hundreds of gigabits, long-running batch jobs where an interruption is insanely costly, and the need to let data engineers pull data on demand while locking down model theft.
- Agentic traffic and why agents circumvent security policy. An agent that needs internet access to achieve its goal will change its IP, change its protocol, and spoof a port, so enterprises need to discover the agents they have, be able to cut them off, and grant access by injecting credentials from outside the workload so the agent never holds them.
- Identity in the fabric and Cisco's converged security strategy. Policy binds to identity carried as a network-level tag, maturing toward mutual TLS with cryptographically signed identities and post-quantum safe ciphers, while Cisco turns campus, data center, and Kubernetes environments into uniform observation and enforcement points.
- AI is pulling workloads back into the data center. Running containers on premises is no longer hard, and the data volumes AI pushes onto the network make public cloud markedly more expensive, so even cloud-first enterprises are expanding or building data centers again.
The lines worth sharing
“You can lock down the behavior of an agent the way you used to lock down the behavior of any other container, but that agent is going to start circumventing your security policies in order to achieve its goal.”
Craig Connors
“We want to grant access without giving the actual credentials to the agent, because the agent will abuse them if we do so.”
Thomas Graf
“Everything touches the network. No matter what an agent is doing, it's not doing anything without crossing the network somewhere.”
Craig Connors
Common questions from this episode
Why is container networking the foundation of AI infrastructure?
Because the future of enterprise AI is hybrid, spanning public cloud and private data centers with both open and closed models, and the logical layer of that multi-cloud world lives in containers. Kubernetes has become the substrate for AI workloads, and most AI agents will run in containers, so networking and securing containers across multiple trust domains is what makes hybrid AI infrastructure workable.
What is Cilium and why is it the default networking layer for Kubernetes?
Cilium is a container networking interface built by Isovalent on eBPF, designed from the start for scale and identity-based security rather than repurposed network virtualization. Google adopted it and baked it into its managed Kubernetes engine, the other major cloud providers followed, and it became the default plumbing that ships with most Kubernetes distributions, meaning most enterprises are already running it.
Why does security for containers and AI agents need to be identity-based instead of IP-based?
Containers and agents are volatile: they come and go, and a service may scale across many replicas at once. Binding policy to a workload's identity, carried as a tag in the network packet rather than inferred from an IP address, means policy holds as instances appear and disappear, and it matures toward mutual TLS with cryptographically signed identities, with post-quantum safe ciphers now treated as a must.
How do you stop an AI agent from misusing credentials?
Do not give the agent the credentials at all. The pattern discussed in the episode, originally built for untrusted batch workloads, injects the certificate needed to reach a service from outside the workload, so the agent is granted access to the service without ever holding the token or certificate itself. That matters because an agent pursuing a goal will actively circumvent policy, changing IPs, protocols, and ports to get its task done.
Why is AI pulling workloads back from the public cloud into the data center?
Two reasons from the episode: running containers on premises is no longer hard, so public cloud is not required for container platforms, and the volume of data AI forces onto the network is making public cloud a level more expensive than before. Large cloud users were already repatriating workloads for cost reasons, and AI is accelerating that, with even cloud-first enterprises expanding or building data centers again.
Read the complete conversation
Full episode transcript · 46 minutes
You can lock down the behavior of an agent the way you used to lock down the behavior of any other container, but that agent is going to start circumventing your security policies in order to achieve its goals. Hi everyone, I'm your host Nick Lippis and welcome to the Built for Trust podcast where you get to hear from all the folks who are building and shaping AI enterprise infrastructure. Let's get right into it with our guests. Hi everyone, welcome to the Built for Trust podcast. I am really just so psyched for this podcast. We're all about containers and we're going to really talk about how fundamental containers are as almost like a stability point or a major building block within not only just AI infrastructure, but clearly around kind of cloud infrastructure, you know, as well. And those two, how they are inextricably linked together. And so with that, I have the folks who really have developed container networking on the podcast. So actually Thomas and Craig, let me, I don't want to do injustice to your background. So let me ask you guys to like, say hi and just do a quick intro.
Thomas, why don't you go first? Sure. Hey, I'm Thomas. I'm the co-founder of iSurveillance and our CTO of Cisco security. iSurveillance joined Cisco two years ago. Before that, we ran iSurveillance for eight years and developed iSurveillance, Cilium, the standard for container networking. And before that, I was a curl developer at Red Hat for 10 years doing low-level networking development. That's my, that's my background.
Awesome. Great. Thanks, Thomas. Welcome. So psyched to have you. Craig? Hey, Craig Connors. I'm the CTO for Infrastructure and Security Group, which covers data center networking, computing, compute, security at Cisco. I've been at Cisco for a little over three years. Most of my background prior to this in SDN and SASE space, I had companies like Tulare and Bellopod. Nice to be with you. Awesome. Yeah. And those, those are great companies too, you know, as well. So Craig, thanks. Welcome. Really just so psyched to have both of you. So we're going to really kind of dive into this.
I want to provide a little bit of context. I think this is really coming from like within the ONU community to kind of frame like our, our discussion. So we all know that enterprises are really kind of facing a large amount of uncertainty today because we have a token economy that we're all working within that really hasn't been budgeted for, for AI. There is now an AI FinOps kind of program within most corporations that are starting to like emerge to try to manage that token economy. But you know, an open question is this, you know, is that these frontier model providers haven't really earned enterprise trust. And we clearly know that from the last couple of weeks, as we've been seeing jail breaks from, you know, various different frontier models. So there's an earned enterprise trust, you know, on security deficit that they are now really kind of working towards. But what's interesting too, is that the cloud providers have that trust. And they built that trust over decades. It wasn't too long ago that we all were just saying, "Oh, we don't want to put our workloads into the clouds because it can't be trusted." And they work hard on that to really build the trust for the enterprise marketplace and they have. Now economics is starting to shift a little bit and there's repatriation kind of going, has been going on for some time. So there's this kind of movement and kind of going back and forth around workloads. But the community within ONU voted at the last summit that really our future is all around hybrid, it's around hybrid AI infrastructure, both open and closed models.
And the cloud providers play a big role within that. So the logical layer associated with that, that multi-cloud hybrid world, it lives within containers. And what makes container networking core to AI infrastructure is that movement towards hybrid. So I want to kind of give you that perspective on the importance of containers, the importance of the ability to be able to network containers across multiple different trust domains in particular. So that's the context. Let's kind of dive into this first. And so I thought it'd be good for us to lead maybe with Craig. Before we got into the technology, I want to really kind of have this original, what kind of motivated Cisco to acquire isovalent a couple of years ago. So Thomas built the technology, but Craig was instrumental in the acquisition. So Craig, what did you see in isovalent that kind of drove you to kind of spearhead that acquisition. Well, you kind of made the case for me already in your intro. Two big things that we saw happening. One, the shift towards hybrid cloud. And two, the fact that Kubernetes was becoming the substrate for AI networking made it imperative for Cisco and as a major networking player to have a bigger foothold in the cloud and Kubernetes networking space. So that was one side. How do we extend this market leadership we have in enterprise on-prem networking to cloud and Kubernetes networking? But the other was the security angle, right? Isovalent built Tetragon. They were using EDPF for runtime security. And really, as Cisco's on this mission to merge security into the fabric of the network and converge networking and security and expand the cloud and Kubernetes and Isovalent's there, they're the leader there, and they're doing converge security at the same time. It was a perfect fit.
Yeah. Awesome. Yeah. Well, in 2020 hindsight, it makes sense, but maybe it took vision to do that. When was Isovalent acquired in 2024? Was that kind of around the time frame? Yes. Yeah. Yeah. So it took vision to kind of see that, and not to take anything away from that, but it took even more vision for Thomas for you to develop the technology because you had to be you had to be doing this probably in 2020, 2021, you know, time frame. So, you know, maybe, you know, kind of the question here is that what's kind of interesting, you know, about Isovalent, you know, is that most enterprises, you know, are already running your technology. They didn't have to buy it. It was part of the platform, right? It came bundled within. It was almost like the early days of like, you know, Intel inside, you know, when you basically went, you know, K8, you also, you know, got, you know, Isovalent, you know, and you got, you know, the core, I think, technology that came with it was Selenium. And, you know, and that's essentially the default plumbing for most, if not all Kubernetes deployments. So, so maybe, you know, if you can kind of talk a little bit about, you know, about in particular Selenium and, and I think is it Tetragon, you know, as well as the core kind of networking components and technologies there. Sure. Yeah. So it was, I think, around 2016, when we got started, it was around when, when a company called Docker introduced containers for the first time. Right. And, uh, at the time we were working on virtualization. So I was working on OpenV switch, which was, uh, kind of the de facto, uh, network virtualization platform or open source project. Uh, it was developed by a company called NYSERA that got acquired by VMware. This became VMware, NSX, T in the end. Um, but when containers came around, it was becoming clear that like the, the, the level of networking that we have been developing to that point would not actually fit containers, even though the initial wave of solutions that hit the market were all repurposed, like network virtualization solutions repurposed for containers. And we really saw the need for scale, but also identity-based security as the core pillar that we had to, uh, base this on.
Uh, we had this amazing technology called EBPF playing around, which was kind of a different, different companies used it for different purposes. Um, it was ideal. It was ideal for software-defined networking, but we also saw that it was ideal for application profiling and monitoring and networks and runtime security. So it's, it's a super, super powerful technology where you could actually see fit for, for multiple, uh, solutions. We ended up building Cilium first, which is to contain a networking interface originally built for Docker. Then Google brought Kubernetes onto the market.
Uh, and that quickly became the standard for container, uh, um, orchestration. And that was all the opportunity for us to really tailor to Kubernetes and step-by-step become the standard. And I think initially it was a dozen of users, a couple of hundreds of users, and then eventually, uh, the first cloud provider in the form of Google adopted us and baked Cilium straight into their Google Kubernetes engine. And voila, when you, when you created a managed Kubernetes cluster, you had Cilium on the hood automatically. AWS followed eventually Microsoft Azure with, uh, with, uh, with, uh, AKS, Azure Kubernetes service follows. And step-by-step, we kind of got baked into, um, the Kubernetes distributions one by one, which was ideal for us as a company because our technology was already embedded in, in kind of the default Kubernetes distributions. And all we have to do as a company is to convince our customers that the enterprise distribution that we have that adds a bit of more value on top of the open source version is definitely worth it. And that's how we turned this into a successful business, business model as well. So started with, with, with Cilium on the networking side, and eventually evolved this into Tetragon, which is the runtime security component. We always saw that from a security perspective, you need both, you need the runtime, the process context, and you want the network context as well.
And for true security, you need to eventually combine the two together. And that's what we're doing at Cisco now. Yeah, that's great. So we, all right. So great. So obviously Selenium kind of, uh, networking of containers, um, and, um, uh, Tetra. Oh my God, I forgot the class. Yeah. I kind of, kind of went through a kind of, uh, a mind block there for a second. Um, security. So actually, you know, what's really kind of fascinating, you know, is that there's a couple of things that you mentioned there. You know, one is those days back around, um, Nasira, um, with, you know, Martin Quesada, and there was the whole focus around kind of, um, you know, open flow, um, and, um, kind of, um, you know, software defined networking focused on the data center, but it was focused in the wrong place. It was focused around these white boxes that really never, you know, took off or took a long, long time, you know, for them to actually start to get a little bit of traction, which they're, which they're getting now, uh, with Sonic. Um, but it said, you guys really, you just really honed in and focused on containers. Um, it's so funny. I remember the first time when Docker came into like, you know, the, um, um, the vernacular within the industry, I was having this conversation with someone and they were like, Oh, you mean the people who make pants?
Like, no, that's a different kind of Docker, you know, but, uh, but I think, you know, there was a, there was a whole flurry of activity that was going on there that I think was mostly in the domain of the hyperscalers. Um, and then, you know, and then it really started to branch out into like, you know, all parts of like enterprise compute. Um, so, uh, really insightful for you guys, you know, to be in that space at that time, um, and to be, uh, kind of integrated, you know, as kind of that default kind of networking in the container, uh, in the container world, um, led by really the hyperscaler companies. I was actually super, super funny anecdote because you mentioned, um, um, Martina Casada. Martina was our first investor. So he was like, he later said this was the easiest investment decision ever because you guys are completing the vision that I started.
I really had this, this vision and this, this, this thought a little bit early for containers. This makes even more sense. So for him, I think it took like minutes to convince him and he wrote our series a check. Yeah. You know that, well, that, that, that's a great backstory. Um, I remember like, you know, I was invited, um, you know, Martin and Steve Mullaney, uh, who was running the Sierra at the time, invited me to come in and to kind of just give me an overview about what they were doing. And they talked, the only thing they talked about was open flow at that time. There's, you know, um, and I'm not sure if it was like, kind of like mostly kind of like, uh, um, kind of an open version to be, well, VX land is kind of open, you know, um, uh, open switch, you know, kind of a open switch, you know, um, um, um, I forget the, is it kind of, yeah, maybe it's just like, uh, open V switch or something like that. Open V switch was the project. Yeah. Yeah. Right. Um, but I thought like, I'm not sure if it was just like, okay, it's almost like a lot of times, like, you know, when you're kind of in that early, uh, early stage, a lot of companies will kind of talk about something over here to distract you from what they're really doing over here. And I'm not sure if that was it, or if there really was like, you know, kind of a belief in open flow and, um, you know, in, in the company, but clearly they did really, really well, uh, with the acquisition, um, on the Sierra buying by VMware. And, um, and I think that really kind of like, um, helped establish, you know, Martine, you know, as a, you know, a major figure like within, uh, within our industry.
So anyway, um, where my mind was kind of going, as you were talking, Thomas, is that one of the, one of the disruptions that, that Kubernetes or containers did was with CMDB. Right. Um, and so when the world was physical, um, we could really have stability around, you know, what is the database around all of the devices, uh, within an infrastructure. Now that's like nearly impossible, you know, with kind of containers coming up and going down and, you know, and just that level of dynacism. Is there anything that you, that you're kind of working on to kind of help, you know, kind of close that gap, um, so that you can't do this with people.
That's got to really be kind of an AI dimension, you know, to kind of keep track of, you know, the, the state of, uh, of an infrastructure at any one particular point in time. But you guys have got to be at the center of that. Yeah, absolutely. I think that actually goes back to like the fundamental design consideration of Cilium and it, we haven't even invented it. I think we, we, we kind of looked at Cisco ACI, which did it before us and, and, um, just kind of to, to decouple the networking and the security concept. And very similar to like, if you look at something like SSL or TLS, that does not depend on how your network plumbing is being done. And while for, let's say for virtual machines, but I think in kind of a virtualization age, it's, it's, it was still possible to look at like network topology and also use network topology as a way to do segmentation, things like that. For the age of containers, or if I'll, if we look now at AI agents, that's impossible. These things are volatile, they come and go. You really need to look at the connectivity layer. And then on top of that, completely decouple to security identity based. So if I'm not depending on network addressing to understand who somebody is, do an enforced network security. And that, that was conceptually what we introduced with Cilium, which matched exactly what Kubernetes designed or how Kubernetes saw the networking as well. And it's actually the same way as Cisco ACI is doing this as well. Yeah. Awesome. Okay. Yeah. So we had that separation. And then what would feed, is the concept of CMDB outdated then? I think it's no longer kind of the, the CMDB may not see everything anymore, right? You may have a partial picture. You may have a partial CMDB in the cloud, like you or different CSP. Kubernetes is kind of a mini CMDB. For some customers, VMware is still part of the CMDB as well. We now need to compose this together from multiple sources, right? And over that CMDB be all an identity layer where we can do security. And also for observability, use that identity layer to identify who are you talking with. And what's interesting is it's no longer just workloads. It has to span to the endpoint. And then on endpoints, on your laptops, you will, you will, on agents, they will also need identities. They will need non-human identities. So it really, I think that the, the traditional sense of a CMDB needs to be reshaped, right? Because it's just too, too, too static of a concept. But that is, that was the core value of CMDB, right? So is CMDB, CMDB debt as a concept? No, it's more important than ever. But the way we implement CMDBs has to be this AI augmented data layer. It has to be completely different.
Yeah. Yeah, I totally agree with that. It's, you know, the world was a lot easier before, you know, but we have a lot more flexibility now. And we, we definitely need a new approach, you know, for both kind of physical and kind of virtual, you know, worlds to understand state at any one particular point in time. The identity piece, I never really kind of like associated the identity part of, of, of what you all have built, you know, with that. But clearly, I see it as a component of agent identity. And that, that is key and fundamental towards, you know, the agentic world that we're all all moving into right now. So, um, maybe we can kind of circle back to that. But I think what I wanted to do is, um, you know, make all of this real, you know, because like, it's, it's almost like, um, maybe not the best analogy, but it's like, when you're kind of like, you know, when you're studying math, you know, you kind of do algebra, you know, and you don't really understand algebra until you do calculus. You don't understand calculus until you do differential equations. Uh, and you have all this level of kind of, um, virtualization, you know, that happens. And so, um, and I think this has a little bit of that. And so I think as you go higher in the virtualization piece of it, um, really good use cases come into focus. So, um, and it kind of really explains everything that's going on underneath.
And so I think if we talk about open AI, that might help, you know, people kind of understand because that's open AI, you know, kind of like runs, you know, on isovalent, you know, um, and containers. So why don't we talk a little bit about, um, about, uh, about open AI as a, as a use case, a real good concrete use case. Yeah. I think open AI is, I think fascinating. We have, I mean, I think lots of fascinating customer reference stories that we can talk about. Open AI is super interesting because it combines multiple aspects. I think that are really interesting for, let's say, core networking people, their scale, these learning clusters are massive, right? These are massive Kubernetes clusters that data engineers are using to build models. And then there's also open AI applied where they run ancient forms. Um, so if you run, uh, open AI codecs in the cloud, actually your call, your task runs on a Kubernetes cluster. And your infrastructure has put entirely new levels of networking demands on servers. Like we're talking 400, 800 gig server links to the switches. These are like performance numbers that for traditional software networking were not approachable before.
So it kind of shows that EBPF or the use of EBPF by Cilium really allows us to be high-performing as well. And the last factor is these are kind of even more extreme than containers. It is as volatile, but it, these are long running batch jobs as well. Like these learning jobs, they run for days and guess what? If they get interrupted and you have to restart, it's insanely costly. So the reliability angle is uh, probably telco level plus two, some, something like that. I think it actually combines multiple levels of data standard requirements, scale, high throughput performance, um, reliability, and then also modern security. I think we've been very challenged by, um, some of the demands and some of the questions that came up. Well, in a learning cluster, you need access to various data sources because you pull in data from here, data from here, and these data engineers, they just want to pull, pull data from somewhere to, to pull into the, into the models. At the same time, you want to prevent that the model needs can get stolen. So you're in this position where you don't have a static kind of connectivity demands. It's very, very on demand. At the same time, you want to protect the models as they get developed. So we had to come up with completely new network security models that allowed us to lock down, um, model theft. So that was interesting. It was a very challenging, I would say, um, um, onboarding, but it was super interesting and really brought the product and, uh, technology forward.
Yeah, I, I bet, you know, the, the level of scale and dynacism that you're seeing in there has got to be kind of unparalleled, you know, what, what a great proving ground for scale. Um, yeah, just, um, yeah, just amazing. You know, Craig, do you want to add anything to like, um, the open AI, you know, uh, use case or, or any other use case? Yeah. I mean, I think Thomas says, has covered it, the, it all ties back to the intro, right? Uh, what did open AI need? They needed the ability to run across multiple different environments.
I think the one thing Thomas didn't mention is what happens if you need to extend more compute and you're, you're the company buying all the compute in the world. Well, you need the flexibility to adopt any type of substrate, uh, at any time. And so I think that's the one of the thing I would call out that like it not only gave them the flexibility to do this in a high performance way with security, but also across whatever infrastructure essentially they could find. Yeah. That's actually a really good point. Just the kind of, um, um, interoperability or the mobility, um, mobility, you know, of containers kind of affords that, you know, flexibility around compute.
Um, and you have consistency around networking and you have consistency around security. Yeah, that's good. Um, yeah, great, uh, great, great point. Let's, let's dive a little bit into, um, then kind of agents. And, um, I'd like to maybe give you an example that I think most in the owner community are struggling with, and you might already have a lot of data, you know, on this and it's basically this is that I think as we kind of move into an agentic kind of world, um, what most, you know, and most haven't, you know, they have really haven't kind of have this, you know, large scale deployment of agents where you have an agent that will, uh, is given a task, um, with an intense, you know, it now has, uh, has rights and access that are granted to it. It has a persona associated with it. It has an identity associated with it. Um, and then maybe it accesses an MCP, um, a server for access, various different tools could access, you know, data, you know, um, kind of a data, um, repo, um, to help in that task. Uh, it may realize, hey, I don't have all of the skills, uh, in order to do this, uh, task. So I'm going to go to an agentic directory and find an agent that actually could help in this. It pawns or distributes part of the task there. And then all of a sudden you get this distribution of traffic flows, uh, in order to, um, survive and to deliver on that original task with that, you know, with that basically, and we hear that mostly from like the hyperscalers who are kind of dealing with this stuff like in scale.
But what I think in a large enterprise, they don't understand is what does this do to balance requirements? What is it due to latency requirements? What are the traffic patterns, uh, that are about to, um, emerge? Um, you know, how does it work, you know, and does it work on existing enterprise infrastructure? So, um, are we looking at in essence, like appearing into, um, a world where, you know, we're going to need a kind of a wholesale upgrade around the infrastructure that we built? And maybe we contain this right now to the data centers, um, because the wide area is a totally different, you know, a beast, you know, in this as well. So, you know, maybe kind of building upon all your expert expertise and all these different use cases, you know, are enterprise networks ready for this kind of a traffic flow? I mean, traditional enterprise security is not, that's for sure.
Uh, you know, one of the, I think the most challenging thing, and this was hammered home this week with the open AI and entropic jailbreaks is that you can lock down the behavior of an agent the way you used to lock down the behavior of a, any other container, but that agent is going to start circumventing your security policies in order to achieve its goal. So if it, if it needs to access the internet to achieve its goal, it will change its IP. It will change its protocol. It will spoof a port. It will do things that are, uh, never before seen. Right. Um, so it's, it's, it's, I used to say agents behave like humans because they were unpredictable, but now it's more than that. It's like, uh, humans that have the intent to circumvent your security. Yeah. With the best intentions of accomplishing it. Yeah. And the skills that are beyond like so many, you know, to like, you know, understand the infrastructure probe and infrastructure, you know, and then find exits, you know, that maybe others couldn't find. Yeah.
Yeah. Well, I think there's so many different angles. So it's just like, if you look at the connectivity itself, additional, a traditional data center may not even have the bandwidth with the, like the, the rise in East to West communication that we'll see from enchanted workflows. Then your typical active visibility, I think the first thing you ever need, if you have enchanted workflows, is to be, to be able to identify and discover agents, because you may not even know that you have them. So you need to be able to like identify them. Then you need a way to cut them off because as Craig is saying, they may be doing all sorts of things, uh, not intentionally bad, but like, maybe like an eight year old, uh, trying to find creative ways around rules. So I think even those basics are not always found in traditional build outs. I think definitely time to look at network infrastructure level up. Uh, some may actually come with, um, I would say modern ancient platforms.
If you look at something like Nvidia open shell that is based on Kubernetes. So it's a Kubernetes platform and that has some capabilities in, and because it is Kubernetes space, again, we can bring in eye surveillance. Then we get quite a bit of the capabilities that we have built to lock down agents, to get visibility, to, for example, grant, um, tokens to agents without actually giving it to the agent itself. It's very interesting. It's not a new use case to us. Uh, one of, uh, one of our customers Palantir had this a while ago, they are running lots of batch jobs of untrusted workloads. So they got from their customers to get workloads that they run and they need to, they need to address services. And because these were untrusted workloads, they did not want to give them the credentials directly. So we've built essentially what is called TLS origination, which means that the workload can access the service, but from the outside, we inject the certificate required to access this. So you can grant the workload access, but you're not giving the token, the SSL certificate to the workload itself. And that's exactly what we need for agents as well, even on a new level. But that same concept is needed. We want to grant access without giving the actual credentials to the agent because the agent will abuse them if we, if we do so.
Yeah. So I think, you know, so I think what I'm hearing clearly on the security side, um, there is a, you know, wholesale upgrade, refresh, new approach, um, around security infrastructure. Um, that's, uh, that's needed. Um, I, and on the network infrastructure, uh, the physical, um, the physical plumbing may be there. Um, you know, and upgrading that, you know, obviously you're upgrading, you know, switches, you know, and, and moving depending upon like, you know, your requirements is there are plenty of options around increasing the bandwidth and, you know, uh, the two tier architecture that we've built in the data centers, uh, should be able to scale because those are not even sure they will be able to scale. That's what, you know, all of the, uh, hyperscalers use and, uh, and also all of the kind of, um, frontier models uses a kind of a two tier, um, class, you know, non-blocking matrix, you know, to, to connect, um, you know, uh, servers and, uh, and switches together. So I think we're there that that could be scaled up and we've kind of proved that can be scaled up to like, you know, hundreds of thousands of, of GPUs, but it's so, but it's really kind of the virtual networking, you know, piece, um, that, that needs to be kind of thought through carefully, uh, to start thinking about, um, you know, agents and agents at scale.
Does that kind of capture it? You know, um, kind of networking is more on the virtual side of it. security is almost like a, a total rethink. I think many, many network services that were previously maybe in hardware are moving more to the virtualization side because it's often hard. Like say we can take firewalling as an example, like scaling the, just the plumbing itself doable, right? We know how to do that. Scaling a centralized firewall and trying to apply to all these two S traffic that's becoming very hard, very expensive. So we're seeing definitely more distributed models of services, firewalling, low balancing visibility. Uh, and one way to do that is through software, right? To actually use either server capacity or a DPU on the server or a DP on the switch. Makes sense. Okay. All right. Great. So I think, all right. So everyone get ready, you know, you know, if, uh, if you're an operations, you know, person, you're a CISO, you know, CTO, um, it's time to really kind of start thinking about your story to the board for, you know, refresh, um, and increasing, um, expenditure in your infrastructure to support the AI era that we're, that we're all in. Um, but let's kind of guide that a little bit. Um, and I think probably the two most important, I think topics, you know, um, around, you know, the scaling out.
And I think, you know, Craig, you talked about it around security and that's really identity. Um, and security and how those are becoming kind of, um, you know, kind of converge together. So I think the previous discussion that we just had really everything up to this point is that we know that all the CISOs like within our commuter asking, you know, what is that Asian, you know, what is it allowed to touch? Uh, who assigns it a persona or what are the rights that it is granted? How are those rights assured? How do we monitor what that agent, um, is doing? How do we we make sure it doesn't go rogue, um, you know, within, uh, within our environment? So there's, there's an awful lot of discussion, um, and concern around, uh, around agents and, uh, and their identity and how do we put in the right controls and the right guardrails, um, to govern, um, this agentic world that, that we're all in. So, you know, Thomas, maybe if we can start with how identity is built into the fabric, I think you already started talking about that, but maybe if we can kind of isolate that around, around identity, um, and then, you know, Craig, if you can, um, kind of take that up a notch, you know, around, you know, um, Cisco and you and Tom Gillis have done just really a fabulous job around, um, kind of reinventing security, um, within the Cisco product set. And so when we talk a little bit about the Cisco security strategy, uh, after Thomas, uh, talks about, um, how identity is kind of baked within the fabric. Yes. I think we really see identity and I mean, overall security should be a property of networking in general. And in order to achieve that identity is the core pillar. And for containers, that's non-negotiable, uh, containers come and go.
And in particular with containers, we often see that multiple instances of a container implement the service. So you can scale up the number of, of containers to have more capacity to deliver a particular service, right? Which is different from, let's say a traditional database where you need to build, put a bigger box. If containers, you can simply scale up the number of containers and you get more processing power. But that also means from a security perspective, if you don't want to then constantly change your policy, like if a new container comes up, you don't want to change the policy. So you bind policy to identity and identity describes all the replicas, all the instances of a container. So that's baked into our understanding. Um, and we're doing this very similar to how Cisco ACI is doing as well, which means the network packet carries on the network level attack that describes the identity. So we're not looking at the IP address to figure out who this is coming from. We're looking at attack. So that's kind of the baseline. And then we also have more modern ways of, uh, carrying this identity and doing enforcement all the way to TLS and mutual, uh, mutual authentication, which would be in modern container roles is the ultimate form of zero trust networking, where you're not just doing identity based, but you're actually doing, um, TLS with cryptographically signed identities and certificates that you can mutually authenticate sender and receiver. That's kind of the, the end form of this. Uh, but typically this is kind of a journey from identity description of policy to network tag based identity to end the last step being full TLS enabled authentication.
Yeah. Is there like, um, does kind of like quantum cryptography play a role in here? Absolutely. Absolutely. Yeah. I mean, I think traditionally, I think, uh, maybe five, 10 years ago, FIPS was the primary kind of complication around, um, encryption and authentication. Now, um, I think post quantum safe ciphers are kind of a must. So that is obviously what we're building in as well. Yeah. Yeah. Especially with the recent, you know, executive orders that were signed and, and, uh, and, and actually what the prelude to that was, um, was kind of a scare, uh, within the, you know, intelligence community that really just kind of drove a lot of activity, uh, in quantum. And now quantum is, I know, for example, in our fall program, we have a dedicated track, you know, just around, you know, how to get ready, you know, for quantum. So that's great. You know, you guys are already baking this in because it's now it's like, you know, containers are fundamental, you know, uh, to the world that, uh, that we're living in. So, um, uh, Craig Cisco brought our security strategy. Yeah. So, uh, Thomas and I said something that might've seemed disconnected. So I want to pull it all back together to explain the story, because I talk about using security in the fabric of the network. And Thomas mentioned decoupling networking security. Yeah. That's when I caught that you're at odds, but they actually, they actually mean the same thing in a, in a weird way. Um, so obviously our strategy has been to fuse security into the fabric of the network. Why? Because everything touches the network, no matter, no matter what an agent is doing, it's not doing anything without crossing the network somewhere. Right. Um, and it provides a uniform point for us to apply enforcement and get observations. The shortcoming is that the network's not enough. Thomas mentioned identity. Um, we talked about agent behaviors and that's things like semantic inspection of prompts and stuff like that. All of that is contextual metadata on top of what we have from the network. So the question is how do I build a network that is uniform and providing the observation points, uniform and providing the enforcement points, but isn't the only thing involved in doing that observation and enforcement. And that's the two pieces that we have managed to put together by bringing the Cisco security portfolio on top of networking. So now when I've got agents, they're not only running in the data center in a trusted environment. We're going to have agents everywhere in the world. So if I can build common visibility across wireless campus switching data center, switching firewalling my Kubernetes runtime environment, if I can use all of those places as potential places to proxy traffic to an MCP gateway or an ADA protocol gateway, if I can use all of those places as places to insert enforcement. Now I have turned my network into a living, breathing security fabric. And that's really the vision of Cisco security. And that's why ISO valent and Cisco made such a perfect sense. And you already see this in like the Nexus and ISO valent coming together with HyperShield as the bridge. You know, you see this in practice, how well this works.
Now I have my Kubernetes environment, my switches and these firewalls running inside the switches, doing layer four to layer seven security consistently across the infrastructure. Yeah. It was interesting, you know, when you were kind of describing just outside of the data center too. But I think that's within a trust demand. But I think also like, you know, your architecture also scales across multiple trust demands, right? Because we'll have the agents in kind of in the frontier models. We'll have agents within the cloud providers. And I think, you know, like Cisco's kind of security solution is the way I kind of view it, you know, correct me if I'm wrong, is that it's both a multi-agent, multi-trust domain kind of architecture, you know, to, you know, to provide both secure, you know, agent creation and destruction, monitoring across these various different domains.
Correct. Yep. And yeah. And whether it's you know, how do I detect an agent talking to an on-prem model, Cisco is always there in the path to inspect that and secure it. If an agent's using an API to talk to a frontier model, again, Cisco's always in the path to inspect that and secure it. So having that ubiquitous footprint is really what allows us to deliver that. Yeah. Okay. I think this is kind of like, maybe just kind of the last, you know, last point. You know, when we were doing the prep for the podcast, you know, Thomas kind of mentioned something that I've actually been thinking about. And, and he basically, what you said, you know, Thomas was that like, that I would be surprised about how many people are now kind of like pulling workloads back into the data center from cloud because of AI. So it's kind of AI is pulling workloads back, you know, into the data center. So, so even kind of cloud first, you know, enterprises are, you know, either expanding their data centers or, you know, or, you know, building new data centers again.
So can you go into a little bit of that? Because I think what that to me says is that we're not just talking about AI workloads here. We're talking about all kinds of workloads that are now being hosted in containers that are portable and mobile. And AI is suited for that because most agents are just going to be in a container. All right. So I guess the bottom line is that, you know, make that make sense. Why is AI pulling that in? I think it's too fast. One factor is that I think it used to be hard to run containers in your own premises and not in the cloud. That's no longer true. There are plenty of simple ways to run containers in your data center that you do not need public cloud for this. And the second aspect is AI and agentic workflows and essentially the maturity of applications will become agentic in the future. So yeah, and the change of AI in terms of how much data it will force onto the network, I think will make public cloud or is already making public cloud, I think a level more expensive that we have not seen before. And even before AI, we have seen large users that were in the clouds move back because of cost reasons. AI is accelerating that. So definitely we're seeing a container is no longer a barrier and AI is driving workloads back into a data center.
Okay, well, mission accomplished. You made it made sense. Good job. All right, excellent. Well, I think that's probably a good place to stop because I think what it really hits on is that we are kind of in this uncertain phase around building hybrid infrastructure. Part is, you know, in a public demand and part is within, you know, a private domain, your own data center. And we don't quite know where the weights are going to shift. You know, they will shift over time based upon economics and new technologies that come into the marketplace, new kinds of workloads and applications that get created. But I think for infrastructure professionals, what this fundamental and what is important is that you design and architect infrastructure that enables you to put your corporation in the best possible position to be able to move when the market moves or when technologies move or when the economics shift so that your corporation can be competitive within its particular respective market. And I think hopefully what we did in this podcast is show how Kubernetes and also the isovalent, you know, is fundamental to that new infrastructure that we're all building and how it supports both existing workloads as we've been doing for the last, you know, decade. And also how now AI workloads, especially agentic, are building upon that really through that kind of building block around containers.
So I think that's, I think that's all we got. Any final words, Thomas or Craig? Oh, it's been amazing. Thanks for having or giving us this opportunity. I know Craig and I will love to talk about about security. And I think in terms of Kubernetes, we're just getting started. I feel like, like a couple of years ago, it felt like we have been peaking. But if I now look into kind of enterprises and customers, a lot of them, many of them are just learning about Kubernetes, whether it is for containers or now for agentic farms, I think we'll have a lot more fun with customers around Kubernetes in the future.
Awesome. Yeah. Craig, any last words? Yeah. Thanks for having me. I'm excited about the things we can't talk about yet as in addition to the things that we are talking about. So a lot more to come in the future on this. Awesome. Great. Thomas, Craig, thanks so much. And everybody out there, thank you so much for plugging in. Take care, everyone.
The episode summary, topic list, and questions on this page were generated with AI assistance from the episode recording and show notes.