Ep 94 · Apr 16, 2026 · 46 min

Programmable Silicon to Drive Cisco AI Networking Innovations

Nick Lippis with Tom Gillis, SVP & GM, Infrastructure & Security Group at Cisco, and Nick Kucharewski, SVP & GM, Cisco Silicon One

Prefer audio? Listen on
Programmable Silicon to Drive Cisco AI Networking Innovations — watch on YouTube

About this episode

How should large enterprises replatform their networks for AI when they cannot build hyperscaler-style infrastructure? In this episode of the Cisco sponsored series, ONUG co-founder Nick Lippis talks with Cisco leaders Tom Gillis, who runs Cisco's infrastructure and security products, and Nick Kucharewski, who leads Cisco Silicon One, about the build versus buy tension gripping the ONUG community: enterprises want private AI infrastructure for governance, security, and IP control, but GPU obsolescence, power constraints, colocation scarcity, and skills gaps keep pushing them toward public AI services instead.

The central argument is that AI workloads turn the network into the backplane of a composable system: the fabric plays the role the PCI bus played inside a single server, stitching together clumps of CPUs, GPUs, memory, and storage at runtime. Ethernet wins that job because it has no intrinsic scaling limits and removes physical reach constraints, and programmable silicon lets the same hardware adapt to workloads nobody can predict, which matters when customers expect five to seven year device lifecycles. Security and observability move into the network itself, with every switch port becoming an intelligent enforcement point and local reasoning watching east-west traffic that humans can never process.

What we cover in this episode

  1. The build versus buy quagmire in enterprise AI. Enterprises want private AI infrastructure for governance, security, and IP leakage control, but GPU obsolescence, data center power and colocation availability, and missing skill sets keep many consuming public AI services instead.
  2. The network as the backplane of composable systems. In a composable architecture that assembles clumps of CPUs, GPUs, memory, and storage at runtime, the network plays the role the PCI bus played in a server, brokering GPU-to-GPU communication at extremely high performance with very little latency tolerance.
  3. Ethernet frees where AI compute can live. Unlike proprietary interconnects, Ethernet has no intrinsic scaling limits and runs across a rack, a row, a building, or multiple buildings, so enterprises gain freedom in physically partitioning a cluster that still behaves like a single parallel machine.
  4. Programmable silicon as the answer to unknown workloads. Because nobody knows what AI applications will look like a couple of years out, Cisco builds its networking silicon around a programmable packet processor so deployed hardware can gain new packet processing capabilities through software across a five to seven year lifecycle.
  5. Fusing security and observability into the network. A smart switch architecture uses silicon programmability to redirect suspicious traffic to sidecar processors, making every switch port a layer seven enforcement point, recording process-level detail on every east-west flow, and adding local reasoning so agents can check whether traffic makes sense.
  6. Tokenomics and the case for hybrid AI. The three reasons for premise-based AI infrastructure are security, sovereignty, and price per token, and with premise-based open models roughly an order of magnitude cheaper than cloud frontier models for many tasks, enterprises will route each job to the right model.
  7. AI-found vulnerabilities and compensating controls. New frontier models can find decade-old vulnerabilities no human ever caught, so CVE lists will grow by an order of magnitude or more, and since enterprises cannot patch that fast, the network must apply compensating controls, including eBPF in the kernel of switches and routers, without touching the application.

The lines worth sharing

“In this world of composability, the network becomes the backplane. The network plays the role that the PCI bus used to play.”

Tom Gillis

“It almost seems a natural conclusion that networking for AI has to have an element of programmable packet processing, because you really don't know what you're going to need two years from now.”

Nick Kucharewski

“Every switch port becomes a layer seven intelligent security agent.”

Tom Gillis

Common questions from this episode

Why is Ethernet winning over proprietary interconnects for enterprise AI networking?

Ethernet has no intrinsic scaling limits and runs over many physical topologies, from within a rack to across multiple buildings, so an Ethernet fabric lets distributed compute behave like a single parallel machine without the reach constraints of a proprietary interconnect. Enterprises also already know how to run Ethernet: large organizations that followed the hyperscalers into specialized interconnects like InfiniBand have come to regret it because they lack the skill sets to manage a second networking stack.

What is programmable networking silicon and why does it matter for AI?

Instead of hardwiring packet processing so it can never change, a programmable packet processor lets software choose how to inspect and distribute traffic after the hardware is deployed. That matters for AI because workloads change dramatically from year to year, and customers expect five to seven year hardware lifecycles, so the silicon must gain new capabilities through software rather than replacement.

Why would an enterprise run AI infrastructure on premises instead of using cloud AI services?

Three reasons: security, sovereignty, and price per token. For many tasks a premise-based open source model is an order of magnitude cheaper than a cloud-based frontier model, a much bigger gap than the private versus public cloud debate ever produced, so enterprises are learning to route simple jobs to cheaper models and reserve frontier models for complex work. Power availability, not cost, is often the real limiting factor on premises.

How does AI change network security architecture?

Alert and log volume is growing by orders of magnitude, so shipping everything to a central data lake stops working. The alternative is a distributed architecture where every switch port acts as a layer seven enforcement point, a smart switch redirects the small fraction of suspicious traffic to sidecar processors for deeper inspection, and east-west flows are recorded locally at process-level detail with local reasoning agents checking whether the traffic makes sense.

Will AI agents overwhelm enterprise network bandwidth?

Not in the near term: agent-to-model communication is relatively compact, so despite enormous flow counts the traffic is not high volume, though that could change as new applications emerge. The immediate problem is access control, because an agent can carry a person's broad, long-lived credentials while showing very little judgment, so networks need session-based, ephemeral, task-oriented access controls now.

Read the complete conversation

Full episode transcript · 46 minutes

Hi, everyone. Nick Lippis here. I want to welcome you to a six-part series of Build for Trust brought to you in collaboration with Cisco Systems. At ONU, we focus on what matters most to enterprise infrastructure leaders, and this series is exactly that. Together with Cisco, we're diving deep into the technologies and architectures that are redefining networking for the AI era that we're all living in right now. This is a sponsored series, and we want to recognize Cisco for stepping forward, not just as a sponsor, but as a leader, helping to shape and advance the conversation around AI-driven networking. They're bringing forward real innovation, real perspectives, and real solutions that enterprise teams can act on today. Across the series, we're going to explore a lot, everything from silicon innovation and AI data center networking to security for the agentic AI, agentic operations, container networking, and the future of open networking operating systems, and also common fabrics. Fabric is an important word. We'll also get a close look at what's coming next in optics and infrastructure performance.

Each episode is designed to cut through the noise, bring you practical insights, real world experience, and forward-looking thinking from leaders who are building the AI native enterprise fabric. So whether you're preparing your infrastructure for AI, modernizing your network architecture, or trying to understand what agentic really means in practice, this series is for you. Without any further ado, let's get into it. Hi, everyone. I'm your host, Nick Lippus, and welcome to the Built for Trust podcast, where you get to hear from all the folks who are building and shaping AI enterprise infrastructure.

Now, let's get right into it with our guests. Hey, don't forget, there's an AI networking summit coming up, and you're invited. So go to the ONUG, O-N-U-G, dot net website to get your tickets now. So, and also, if you're on YouTube or Spotify or any of the places where this podcast is being broadcast or streamed, make sure that you hit like or subscribe to it. Thanks.

Hi, Tom. Hi, Nick. Welcome to the Built for Trust podcast. So excited about this podcast and the discussion that we're about to happen. It's about to happen. So I want each of you just to say a quick, you know, hi and, you know, an intro. Everybody knows you all, you know, but it's always good just to, like, say a quick hi.

So, Tom, would you start? Yeah, I'll start. So, Tom Gillis, I run the infrastructure and security products at Cisco. Awesome. Well, that was quick and sweet. Yeah.

I get it in there. Name, rank, and serial number. That's how we roll. Yeah. My name is Nick Kuchereski. I'm responsible for Cisco Silicon One, our in-house Silicon development team.

Awesome. Great. Well, I think we're going to have a discussion about, I think, some of the fundamental, almost like replatforming that's starting to happen within the industry. Let me just give us a little bit of kind of stage setting, and then we'll kind of dive right into it.

I think what we're finding within kind of the own new community is that there is a massive, you know, quagmire or tension around build versus buy around AI infrastructure. You know, it's like everyone wants to build their own, but there are so many headwinds that are stopping them from doing that, right? There is the obsolescence of the GPUs, you know, that are just advancing so quick. So, they're afraid of, like, if they invest now, you know, how long will they last?

Power requirements in data centers, co-location availability, data center just, you know, availability. Do they have the right skill sets? So, there's all of those headwinds that are just really just, like, making them pause. So, many of them right now are essentially, you know, they're utilizing public services. And I think you can look at Anthropics, you know, month of February of, like, $8 billion, where a huge amount of that, like, more than half of that is really coming from sandboxes within large enterprises or from the enterprise marketplace and not just from the consumer space.

So, there is a lot of kind of public AI services that are being consumed. And I think, I know that what we want to get in the ONU community is that they want to do more private because they're concerned about governance, they're concerned about security, they're concerned about IP leakage, and they want just control is really it. But they can't build the kind of infrastructure that the industry has set up for the hyperscalers. So, this really, the conversation we're about to have is really all about what does the enterprise marketplace do and preparing for their deployments and their workloads around AI because it is fundamentally changing the behavior of networking.

So, I think that's kind of the preamble, right? So, Tom, why don't we start with you, you know? So, this is also creating a big shift in architecture around data centers. And so, one thing that we've also been seeing is that networking has evolved from connecting devices into really being a fabric where machines connect into it and resources are being allocated. So, anyway, let me stop talking and let you talk, you know?

But I think that's really kind of what we're seeing. Go ahead. Yeah, totally. And Nick, we're seeing the same thing, which is why we're on this podcast. So, first of all, let's start with the kind of building blocks, and then we'll think about the business case for, like, why would we run this stuff on-prem and how will we run this stuff on-prem?

But the fundamental components of what makes up a server that runs an AI workload are going to look very, very different than, it's not so much the components, the architecture of the machines that power these AI workloads are going to look very different than the architecture of the machines, the powered Kubernetes, and VM-based workloads. And I'm a computer science geek.

Like, going all the way back to the 80s, I went to engineering school in Boston, and, you know, thinking machines, and Danny Hillis, and there were these notions of these massively parallel arrays and composable systems, right? The idea behind a composable system is that you have a big array of resources. So, a big clump of CPUs, another clump of GPUs, another clump of memory, another clump of storage.

And then at runtime, you assemble those resources in a manner that you can have a, you know, very large-scale system, coherent, meaning like it's keeping track of, you know, one memory space across multiple machines. So, that's been the dream in computer science for, like, you know, as long as I've been doing this. It's finally here.

It's finally here, right? And I think, you know, the nature of these AI applications is they are inherently very, very parallel. And so, even a modestly sized AI application is going to transcend a single box. So, in this world of composability, which is a very exciting and a very efficient world, the network becomes the backplane.

So, the network plays the role that the PCI bus used to play in a Kubernetes or VM-based server architecture. It's the fabric that we use, just as you said, it's the fabric that we use to stitch together the workload. Now, that's a different kind of network. Yeah.

That's what you think of as a data center network, right? If you're going to be a fabric that's going to be brokering GPU-to-GPU communication, it is extremely high performance, extremely latency intolerant. And there's, it puts a bunch of challenges on the networking community. Yeah, that's for sure. And yeah, I think we have to start thinking a little bit differently about what a network actually does, you know, in this new world.

So, there is kind of enterprise kind of architectures in the hyperscaler, you know, kind of reality. And we're going to really focus on the enterprise stuff. So, Nick, you know, why don't you spend a little bit of time maybe talking about how networking really enables this kind of more geographically distributed kind of environment to take place where you get a greater range of options and choices of where workloads reside, you know, and how they're, you know, kind of invoked and enabled.

Yes, absolutely. And I think it might be helpful here also to talk about a little bit of historical context. When we, many years ago, when we used to talk about highly parallel compute, we were usually talking about proprietary links, which were developed specifically to enable massively parallel computing. And then you also had traditional computing where the entire workload would sit within the confines of a single server.

And then the networks were enabling more and more jobs to be running on different computers. But now when you look at modern cloud architectures, what you see is a large amount of compute. That could be traditional CPU compute, or it could be accelerated GPU or XPU compute. But the interconnect fabric, in many cases, is based on Ethernet, right? So you're a massively parallel computer that is using Ethernet as the interconnect fabric that allows all of those elements to behave like a single parallel machine.

Now, that's very exciting. One, because you can build up these very, very large clusters with the scalability of Ethernet that doesn't have any intrinsic scaling limits. But it also gives you the option of having freedom of physically placing those devices. Because like with Ethernet, you can run that over a lot of different physical topologies. You can be within a rack, within a row, across a building, across multiple buildings.

Now you have this situation where you can have your compute resources connected by an Ethernet fabric that allows it to behave like a single parallel machine. But you now have a lot of freedom in terms of how you're actually partitioning that system. And there's no longer a physical reach constraint that would be driven by a proprietary interconnect. And so that's very exciting when we talk about AI compute in the context of the enterprise environments, because here you can think about it in terms of where are my compute resources, where are they located, how can I connect them together, and then how can I provide network overlays and network management overlays to allow you to then pull those resources together to complete any specific workload.

So there's a level of flexibility that we've never had in the past in terms of high-performance compute. Yeah. You know, what's also amazing, too, is the speed in which your networking is moving right now. You know, it used to be like years to go from like 10 gig to 100, you know, or 10 meg to 100 meg to a gig. You know, we went from 800, you know, gig to like, what, 1.6, you know, T, and that's going to like 3.2 T, like, you know, within what, like 18 months or something like that. So the acceleration of the speed is happening, and I think it's because it's not only just the industry understands, you know, kind of Ethernet, but like the consumers do, too.

So, you know, maybe, Tom, if you can talk about kind of open standards here, like, you know, one thing that I think, you know, the networking industry really deserves a huge pat on the back. You know, it's like we really fostered open systems, you know, into the environment, like with TCPIP and with Ethernet in particular, and a whole range of other kinds of protocols. And I think right now, you know, it's, we should kind of like recapture that kind of open standard, you know, mantra within the industry that has served us all so very well. Yeah, and if you just scroll back in time, open Ethernet has always won, right? There's been all the protocol wars, if you will, right? We have these specialized protocols that maybe solved a particular job, but at the end of the day, what customers want is they just want networking.

They don't want like these weird proprietary things. And so, you know, one of the advantages that we have at Cisco is the silicon that Nick K and his team are building has programmability, P4 programmability at the core. And so that allows us to be able to say whether we're running in, we call it a front-end network mode, which is traditional network, or in this back-end network mode, we will change the networking algorithms, but it's all Ethernet. It's all standards-based. If a customer doesn't like us, they could rip us out and put in, you know, one of our competitors. And that's the world that we're living in.

Now, Nick, you pointed this out. The majority of this AI infrastructure today is going into these massive, massive, massive scale deployments. At that massive, massive scale, they can do different networking protocols and, like, it's okay. But for the enterprise customer, you really want one set of networking tools. You don't want to have to manage two. People are struggling just to manage one, and so, you know, let alone another one.

And this is really the heart of our partnership with NVIDIA is that, you know, Cisco and NVIDIA together are very motivated to make the infrastructure consumable for an enterprise customer. And so, you know, having that Cisco switch that you've trusted to run your data center, be able to string together the GPUs, and even handle data center to data center communication, that's very, very appealing. And that takes a lot of the friction out. Not all, but a lot of the friction out that you were referring to in your opening commentary. Yeah, yeah, it's like, even, you know, on this, so obviously we have InfiniBand in the marketplace, and that really came in, you know, for GPU communications. Those in, and I know this firsthand, you know, I can't mention their names, you know, but those in a large enterprise that have kind of followed the hyperscalers, you know, and consumption of that, really regret it.

Because they don't, they realize they don't have the skill sets to kind of manage that infrastructure, you know, and so they understand Ethernet like really, really well. And so building upon that, so I love the, and I'm sure the industry is going to love the kind of versatility that I think, Nick, that you've been working on within Cisco around programmability. So I want to talk a little bit about kind of that as a differentiator, like for Cisco. So it's like one architecture, you know, that can be scaled across many different environments. So why don't you talk about like kind of the new silicon chips or that are coming out? Yeah, absolutely.

And I think, I think Tom's comments highlighted really well the value of having open standards for the interconnect, but then a specific implementation that delivers features that allow customers to do different things with the network they're trying to build. And that has always been, right, part of the, part of the market for networking equipment and that's the environment that Cisco has really driven where you have that open interoperability, but then differentiated features and you can do specific things. And as you apply it to AI within the enterprise, what we see is different workloads have different attributes in terms of how you want to distribute and interconnect that data across multiple compute elements. And at the end of the day, doing that very efficiently really comes down to optimized hardware that is responsible for directing that data and doing it in a very efficient and low latency way. And what we found is it's extremely difficult to build a single piece of silicon that's going to be deployed for a long period of time and anticipate all of those different customer use cases. The answer is having programmability built into the silicon.

Now, some who come from the compute world would say, well, of course it has to be programmable, aren't all chips programmable. But the reality is when you talk about networking silicon, there are really two schools of thought or two camps. One is to do it all in what's called a hardwired way, where all of that network packet processing is done in a very specific way and it cannot be changed after the fact. There's another approach where you actually have a programmable network packet processing engine so that you can look at different fields of the packet. You can choose how to handle that traffic and how to distribute it. But what that means is, is that the software can then adapt to customer needs.

And so Cisco networking silicon is all based on a programmable packet processor and that gives us the ability to deploy equipment and then, after the fact, to drive new software that would enable new packet processing capabilities, which means we can adapt to new emerging applications. And that is so critical based on what we're seeing in AI, like the applications we're going to have two years from now. It's going to be extremely different from what we have this year, because that's what we saw the previous two years. So it almost seems to, it almost seems a natural conclusion that networking for AI has to have an element of programmable packet processing, because you really don't know what you're going to need two years from now. So you need to be able to deploy the hardware and then have software to upgrade the functionality in time. You know, that's really interesting because like then I'm kind of like my mind is spinning around the various different products, you know, that will, you know, in manifest itself, you know, with more programmable, you know, silicon where, you know, you can have products that are very either geared towards a particular place in the network, like whether that's on the edge or in the data center or in campus areas, or it could be very workload, you know, focused, you know.

So it could be around agentic, you know, and the inference load that is going to be placed upon, you know, networking and being able to handle various different, like the amount of traffic that these agents are going to generate, it's just going to be huge. And the latency requirements are going to be so stringent, you know, as well. So you can understand that things are going to change pretty significantly. And so having that flexibility in the programmability of these products, I think are going to be really key. Tom, you have, you know, maybe some thoughts around this programmability and how maybe, you know, consumers might be able to kind of use them in different ways. So like what was your thinking?

You know, it's like, this is your group, you know, so you were, you know, really kind of driving this. Yeah. Just to be clear, like the programmability is something that, you know, we're talking about like P4 microcode, right? So Cisco will provide that programmability, but the way it's going to manifest itself to our customers is flexibility. You'll get a bunch of new features and capabilities that Nick was referring to without having to replace the hardware, right? Many of our customers are looking for, you know, five and seven year lifecycle on these devices.

And so the programmability we think is going to be really important. Now we talked about like performance and sort of networking, but where I think the biggest opportunity for us actually is in security. And so, you know, something you've been hearing from Cisco is we're really taking network security and fusing it into the network itself. And so that means a distributed architecture where you're going to have hundreds of thousands, possibly millions of little tiny enforcement points that are just baked into the network. Every switch port becomes a layer seven intelligent security agent. So in this world, the packet forwarder is like the traffic cop, right?

And it sees everything. Now we want to be able to apply intelligence at that layer to say this 98%, you know, we're looking for needle in a haystack, right? So, so we want to keep 98% of that haystack and send it on its way. But the 2% that looks like it needs further investigation to be able to seamlessly and intelligently redirect it to a coprocessor that can do that processing. And we've already introduced this architecture and you'll see a pattern in this whole podcast, Nick, is technologies we developed for the hyperscalers are making their way down into the enterprise.

But we built an architecture for our hyperscale customers we call a smart switch, where we've got DPUs that are running in the sidecars. And then leveraging that programmability of the Silicon One, we can redirect traffic into those DPUs for layer four or layer seven inspection as required. That general architecture is going to expand and continue to develop where you're going to have, it doesn't have to be a DPU, right? An array of security processing, a DPU, a GPU, right? We can do some very, very cool stuff in the network that is kind of co-resident with the network packet forwarder, but independent of the packet forwarder. So this is an important, a subtle, but important distinction, which is while we believe in the fusing of network and security together, you still have two personas.

You have the network person. Yeah. That's to update their rules all the time. And then you have a network person. How often does network person want to update the software in a network device? And so we can make a device that is actually lifecycled independently and can be managed, you know, it's like two personalities in a single product.

Yeah. That's a smart switch. And that's a class of device, right? That's not just a feature. And I think that's going to be more and more important as we move into this agentic world, because you're going to need to put security controls in front of devices that you can't put anything on the endpoint. Yeah.

Right? Yeah. Right? And so, you know, we just make the assumption that the endpoint is a wild, wild west, and it's going to be hard or impossible to control. So in that world, the network is the thing you can trust. Hmm.

You know, it's interesting because, like, I just, you know, kind of flash back to, like, you know, a presentation that you gave at the summit. I think it was in New York when you were talking about the amount of, like, you know, logs, alerts, alarms that are coming out of the infrastructure, you know, projecting to grow by orders of magnitude. It's, you know, 100x, 1,000x. Yes. And so that, you know, you can't do the processing that the, you know, like, you know, basically send everything up into a data lake and then kind of, like, tries to find that needle in a haystack. That just doesn't work, you know, you need to have this really distributed processing so that you can kind of understand signatures really deep into the network, you know, before, you know, it gets sent up for, to some incident response center, you know, that now, like, you're basically telling everybody, go fetch, you know, go find that needle there, you know?

Yeah, totally, totally. And so, you know, in this kind of post-AI world where agents are talking to agents, the amount of east-west traffic is just going to burgeon. And what's particularly challenging is that these agents, incredibly powerful, oftentimes they have really broad access. Yeah. But they have the common sense of a printer, meaning they'll do just crazy things, right? And so they have no judgment at all.

And so east-west inspection. Yeah, right? Like a toaster town has a car. Like, what the hell? And so east-west inspection, I think, is going to become more important ever. Now, we are living in a, you know, again, where I think anyone who's listening to an Ono podcast with Nick, Nick, and Tom is going to self-select as being, like, one of my crowd, right?

Like, pretty techie, pretty nerdy. So we're talking to techie folks here, but what a time we're living in. What an absolutely astounding time. Like, the tools that we have at our disposal make stuff that we've been talking about for 30, 40 years suddenly eminently achievable. And here's one of the things that's achievable, Nick, and I talked about this in New York, yeah, two years ago, but now it's real, which is every signal east-west will be observable. We see every single little, you know, hay stock in that gigantic haystack.

We're Cisco. We're moving the packets from A to B. So now we're going to have the ability to record in very great detail every single flow and the process that initiated the flow and the process that terminated the flow. Process level visibility. We can record that. Now, that is a thousand times more data, three orders of magnitude more than what you're ingesting out of your, like, firewall logs into a SAM or an observability platform.

So you're not going to ingest that thing, but we can record it, we can store it locally, and then we federate across these local data repositories. And what I think is really interesting is we are actively working on putting reasoning into those local repositories. So we're putting GPU resources there so that you not only store it locally and federate it, but you can actually run local agents to say, does this make sense? Yeah. And so the robots are going to be watching the robots. Like, that's the world we're going in, but it's the only answer because otherwise nobody's watching the robots, right?

Well, yeah. Yeah. Humans can't process, you know, the amount of, like, you know, data that's coming out of them, you know? So, no, we need the tools, you know? Yeah. I don't know, Nick, do you want to add anything to kind of, like, a distributed, you know, why this, you know, I think, Tom, you made the point.

Like, you know, this whole distributed infrastructure is, like, you know, inevitable, you know? It's like, this is, you know, the way you're putting it soon. These ideas have been around for decades. It was too hard to implement. Now, all of a sudden, we can do all this stuff. And it's the silicon underneath it that is the revolution, right?

Like, whether it's the ability to process multi-terabits of network traffic or, you know, these massively parallel GPUs or both. You can't have one without the other. That's what makes all this stuff suddenly, like, magically come to life. I'm like, whoa, it's moving fast. Yeah. When Nick K talks about the silicon aspect of it, I do want to circle back and talk about some of the security aspects.

Because there's some kind of heartbreaking news there that, again, comes back to, like, network, save the day, right? But, yeah, we'll circle back on that one. All right, Nick, take it away, Nick. I think it points to a really new and very exciting view of networking equipment and networking software running on that equipment, which is viewing the network as a platform that delivers connectivity, security, visibility, and the ability to layer on more and more intelligence to accomplish those goals. It's a very different way of thinking of networking.

If you would, like, think of things in the past, networking hardware could be thought a little bit like appliances. I have my edge device. I have a firewall. It has to do that because I know the nature of traffic passing that boundary. You might have a Wi-Fi router. You'll have a wiring closet switch.

And you think about these as having certain functions that were based on an assumption of what's the nature of the traffic passing at every level of the network. But now, as you move to an AI-enabled enterprise where you have far more traffic being generated by machines than by humans, it's necessary to look at the entire network as a platform providing security, observability, and then enforcement points that you can then layer on top of that intelligence for observability. And that's very exciting from a silicon standpoint because then what do you do for the silicon to actually enable that platform to really drive deeper and deeper intelligence that can be applied at all points within the network? Moving away from this idea of a specific appliance or a specific box to do a specific function to an entire unified platform that you layer this intelligence on top of. Yeah.

Yeah. Yeah. I love it. And, you know, a teaser for our listeners, in this world of composable systems, just exactly what Nick was says, like the box becomes irrelevant. The sheet metal doesn't matter. One of the things that I think is really, really interesting is the possibility of actually having memory resident in the network.

So putting in shelves of memory under a top of rack switch so that can provide caching for legacy storage systems. It can provide memory for processing and it can streamline this composability and make these systems continue to optimize. That's not science fiction. Like we're really working on that. And I think it's really, really cool. It's far from a plan, but like, yeah, that's the world we're heading into.

Well, you know, it's like really what I'm picking up, you know, from this conversation and we knew it had to happen that the products that are out there in the marketplace today just didn't enable the large enterprise to kind of really build their own AI infrastructure. And I think really what we're talking about is that what you both are doing and what Cisco was really doing is now you have the building blocks to enable a whole new generation of products that allow us to replatform, you know, the network as this AI fabric, you know. And so, you know, building upon what the hyperscalers have built or what you have supplied to the hyperscalers and now to bring that to the large enterprise. And a shout out to a very, very close and cooperative relationship with NVIDIA, right? Like, you know, we are both heavily, heavily motivated to make this stuff easy for the enterprise. You don't have to set up these complicated, you know, you don't have to be a hyperscaler to run AI infrastructure in your environment.

And we want to do it in a manner that we squeeze every drop of efficiency out of the available power. Like I actually think, and I spend a lot of time talking to these enterprise customers. It's power that is the limiting factor, not even cost. Cost is a huge one, but we just don't have the watts available. So there, we got to do what we got to do, meaning like, you know, you have a certain amount of power available. We're working heavily with NVIDIA to make sure that, you know, the tokens we can generate per watt are going to be optimized.

And having a systems level approach to this, I think, is really important, right? We want to keep those GPUs spinning full time all the time when they're engaged and make sure that we're feeding them properly, that, you know, the network plays such a key role. But there's a whole bunch of stuff we can do with the orchestration layer as well to optimize. Well, I think, you know, what kind of really excites me about this whole conversation is that when we start thinking about, like right now, we're in a very token consumption, you know, environment. And I think what happens or what gets unlocked here is that now we can kind of segment, you know, this market based upon various different use cases. So there may be open source models that are really good for like low level kinds of applications.

There are like you might want to just use the, you know, the public AI models, you know, or LLMs for other kinds of functions. And so now we can, you know, actually start to think more intelligently about how we want to apply, you know, where we want to consume tokens and where we want to kind of grow our own. And I think it moves us, you know, I think the options and the choices that you're enabling now for the large enterprise is to focus on transformation of their business. And not just do that on token consumption, because lots of companies have realized that they are spending so much on tokens and they're like, okay, well, where's the return on that token, you know, investment. And I think you're now giving them a wide range to deal with, you know, so. Right on.

So when we talk about premise-based infrastructure, there's three reasons you would do it. Security is the first one. Sovereignty is the second one. But the third and one that I think is going to be a much bigger deal than it was in past revolutions is going to be tokenomics or the price per token. So, you know, at Cisco, in my organization alone, we have 12,000 software developers. 12,000 software developers, right?

A lot, yeah. We sent a message down from on high, like, use these tools. You have unlimited tokens. And then I got the bill and I was like, hey, whoa, hey. Okay, slow down. So what we're finding is the philosophy around using these tools is that you want to treat them like a digital team member, right?

Literally, you converse with these tools as if they were humans. And just like with human team members, not all team members are created equal. And so what we're finding is that there are certain tasks that you want to put toward what we call the frontier models, right? Like the absolute best model that can understand complex code base, 20 million lines of code. But then you want to implement dark mode. You don't need the frontier model, right?

You can use one of these open source models, which are usually just like one click behind, you know, the frontier models. But we're seeing economics that these are an order of magnitude. Now, it's not a fair comparison. I want to be clear, right? Like a premise-based open source model versus a cloud-based frontier model. But it's like an order of magnitude difference.

It's not like it's 2x. It's 10x, maybe even 20x difference. And the tools are designed in a way that, you know, your development environment can use multiple different models. And so starting to understand what model to use for what tasks so that we can optimize when there's that much of a disparity is a good thing. NVIDIA is actually working on a model router where, like, you put a job in and it'll figure out which model to do. So we are going to move into a world where, you know, use the right tool for the right job.

And so there's not going to be just one model that does everything, right? You're going to have these optimized models and that'll deliver way more efficiency. Yeah. Like I said, order of magnitude. If you remember the whole debate private versus public cloud? Yeah.

It was within 2x, meaning, like, is the public cloud twice as expensive as the private cloud or maybe, you know, half as expensive? But there was never an order of magnitude difference. Now there's an order of magnitude difference, right? And so I think we're going to see a lot more interest in premise-based stuff because the consumption of the tokens is, like, directly impacted to the pulse of your business, you know? And so, like, it goes through the roof. Yeah.

It's, yeah. I think everything that we're saying is almost like the Appian way leads to all roads to Rome. All roads here lead to hybrid, you know, like a hybrid model is really going to be the way in which we're going to, like, you know, build. So, Nick, I'm sorry, do you want to add? I want to make sure you, you know, get your time in. Hey, there's one more topic I want to throw out.

This is topical. It's news. So, by the time the podcast goes to air, it'll be public information. The next generation of models that we're seeing from these frontier providers are yet another leap forward in terms of their ability to understand software. An interesting artifact of this capability is that these models have an absolutely uncanny ability to find vulnerabilities in the software, like nothing we've ever seen before. And so, and you can't stop progress, right?

So, these models are coming out. And what it means is, you know, we're seeing in, like, open source software programs, these models can find vulnerabilities that have been in place for like a decade. And no human has been able to find these vulnerabilities. Wow. So, this is a problem. And what it's going to manifest is, is that as these models hit the street in the next, you know, few weeks, all of a sudden, the list of CVEs that, you know, an enterprise customer has to deal with is going to grow by, like, one order of magnitude, you know, maybe even two orders of magnitude.

Can you imagine? Wow. Like, you have thousands of patches that you need to apply. And so, it is simply, and this is, we think this is not a moment in time. Like, this is going to just continue to go on where, like, the models are going to be poking holes and finding vulnerabilities. And, you know, every enterprise customer is going to be adopting an AI-powered, you know, red teaming capability.

And so, it's going to be the white hat versus the black hat, robot versus robot. In this world, you have to be able to apply a compensating control. You can't patch this fast. Right? It's just not practical. Yeah.

And that compensating control needs to be applied to the enterprise application, but also to the infrastructure itself. Yeah. And this is something that we've been working on at Cisco, is that because all of our products have a Linux kernel, we have the ability to apply compensating controls using eBPF in the kernel itself to a switcher router without removing that switcher router. Yeah. So, that's going to be a big, a big body of work. But for the enterprise customer that's thinking about their applications, I've got a mortgage payment application.

All of a sudden, it's going to have 500 new vulnerabilities that have to be serviced. Having the ability to apply a compensating control from the network without touching or modifying the application is super important. And so, happy ONUG, right? Like, the network is the thing that ties together the engine that powers these AI applications. But the network is also the thing that we're going to rely on to protect our own applications from the dangers of AI. Yeah, the value prop for, like, the network fabric is just really, is getting just so sky high, both in terms of, it's a brand new operational model around lifecycle management.

You know, ingrating, like, you know, patch management and being able to kind of take advantage of vulnerability identification. And then, okay, how do we kind of patch, you know, for that is huge. The observability going down to kind of port level, looking at every particular, you know, packet, you know, in the infrastructure, the programmability enabling us to, you know, have products that come to marketplace that are kind of purpose fit for, you know, a range of different, you know, use cases is huge. And then, especially as this shift, this kind of build versus buy is going to move back and forth based upon, you know, economics at the time and kind of new breakthroughs that occur. So, you need that flexibility in the infrastructure also so that you have, as a consumer, you have the widest range of choices and options, you know, there, you know, as well. So, I think I want to maybe just hit one last point, you know, before we go, and that is, what do we think will happen with traffic patterns?

Because I think, you know, we've been, there are some that I've seen, you know, traffic, you know, and the whole concept of network being inverted. And we've been talking about that around the need for a fabric, you know, and how, you know, the network is the PCI bus. So, now, you know, it's really, it's kind of accessed into that memory and, you know, connecting all these different devices together. But do we see just exponential growth, like in traffic patterns? Do we just see, you know, latency at levels that, you know, well below kind of like the kind of 50 milliseconds, 20 milliseconds, 30 milliseconds levels, you know? So, yeah.

So, a great question, the one we're asking ourselves a lot, right? So, in the data center, this GPU to GPU communication thing, I don't see any end to that, right? So, it's not like we get to 3.2 tier, we're like, we're good, don't need any more. And we're going to hit some like physical limitations in and around that 3.2 terabit per second per port timeframe, where you simply can't move electrons over a copper wire that fast. And so, we're going to need to go direct to optics. So, there's going to be a whole wave of innovation that's going to reshape the industry around co-packaged optics, you know, sort of direct-to-optic type of solutions.

So, that's, but that's going to be the data center. I think the question you're posing is, you know, will agents talking to agents swamp branch, you know, and wide area networks? Yeah. It's an appealing idea as a networking vendor, you know, like, okay, cool. And Silicon One can handle, you know, like, put it anywhere you need. We have the capacity.

Right now, I don't see, there's a ton of traffic. It's just not super high volume traffic. You know, like, when we analyze our own software development practices, the communication of the agent back to the model is relatively compact, you know? So, while there's tons and tons and tons of flows, it's not, like, going to overwhelm the bandwidth and the throughput. But I think that's a jump ball. I don't think that it's too early to call and say, people coming up a crazy new application.

These agents are watching YouTube videos, you know, to be able to answer questions. And so, you know, interesting times. Yeah. Yeah. That's a really fair answer, you know, to that. I think, you know, I think that's the right metaphor.

It's a jump ball. You know, we just don't know. But I think one thing that we do know is with the introduction of OpenClaw and the exponential growth in Sandboxes and now with NemoClaw and others, you know, coming out, we know that there is a massive experimentation process that's underway right now. You know, maybe one good proxy for that would be the sales of, like, Apple Studios or Apple Minis, you know, where everyone's kind of, like, reusing those to kind of run agents on, you know? No, it's the introduction of OpenClaw is sort of the chat GPT moment for agents, meaning now anyone can make an agent and people are realizing these things are actually pretty useful, right? Yeah.

So, but it, again, is causing networking to reinvent itself because, in my opinion, networks do two things. They move packets from A to B as efficiently as possibly, and then they move packets from A to B but not to C, meaning they do access control, right? And so, so in this world of agents, access control has to fundamentally change because we've had access control for humans, least privileged access. We've had that for decades, right? Everyone knows how that works. Salespeople can go to the sales apps.

IT goes to IT, but you don't want sales in IT. Okay, fine. That's still pretty broad access and it's still pretty long-lived. Then we have machine access. So, like the printer can connect to the print manager and only the print manager. And there, we actually go to the lengths where we tag the packets and we say, wherever these packets go, if it's coming from a printer, it only connects to the print manager.

So, very, very tight, narrow, you know, sort of restricted access. All of a sudden with OpenClaw, you can write an agent that has your credential, has your broad access, but it has the common sense of that printer. And so, we have to have session-based, ephemeral, task-oriented access controls in the network. Yeah. Have to. Yeah.

So, it's not a bandwidth question that the agents are going to blow up your networks, but in terms of access control, it is absolutely blowing up your networks and that's happening now. Not theoretical, not like, oh, next year that's going to be a problem. Like, dude, right now, someone in your organization is writing, you know, an OpenClaw agent, you know, to process their email, but it could be sending emails to, you know, the CEO, you know, and the board of directors, right? Like, these agents, you know, have judges. Yeah. Yeah, for sure.

Like, governance, control, trust, you know, are really kind of top of mind. And I think you started to address that, you know, earlier on, both you, Tom, as well as Nick. So, I think, I think that's probably a good place to end it, you know, unless there's, like, anything else we want to kind of bring up, you know, in, you know, in this discussion. So, you know, exciting times, you know, lots more to talk about. I think we could do a whole separate session on the intersection between silicon and optics and how that's going to impact networking. So, yeah.

Well, you know, it's like they had their, what was the trade show on that? OHC. Yeah, massive. It was, like, massive, you know. It's like, yeah, the whole optics space is just exploding now. And I think, you know, obviously, you nailed it, Tom, and you kind of know this, but you live in it.

But, okay, well, I think this was in, this has been, like, really great. I learned a ton from both of you, and I wanted to thank you, you know, for sharing, like, all your insights, you know, with the community. But I think we are in a, I think one of the key messages, and we're going to continue the conversation at the summit, you know, in May 13th, you know, to the, and the 14th over in Frisco and Dallas. I will be there live for any of our listeners. You want to come get an autograph? I'll be signing autographs.

Well, I'm not really signing autographs. Well, and also photos. Photos. How's that? Yeah, that works. All the favorite drinks.

Yeah, drinks and photos, autographs are all good. But I think, you know, we'll continue the conversation there, because we just really scratched the surface of all the various different ways in which AI is impacting our infrastructure, and how it is now real-time. Everyone has got to be planning for how you're going to replatform your infrastructure for AI. Okay. Great. Thanks, Tom.

Thanks, Nick. All right. Thanks so much. Have a good day. You too. Thanks, everyone, for listening.

The episode summary, topic list, and questions on this page were generated with AI assistance from the episode recording and show notes.