Ep 80 · Dec 18, 2025 · 16 min

AI Remembers Everything: The Sovereignty Dilemma

Nick Lippis with Tom Gillis from Cisco

Prefer audio? Listen on
AI Remembers Everything: The Sovereignty Dilemma — watch on YouTube

About this episode

Can an enterprise trust AI infrastructure with its most sensitive data when a model that learns your secrets never forgets them? That is the question at the heart of this episode, recorded with Tom Gillis of Cisco just after his ONUG keynote. Unlike a field in a database, knowledge absorbed during model training gets infused into the model itself and cannot be deleted or undone, which makes intellectual property leakage a harder problem than anything the industry faced in the move to cloud. That tension is pushing enterprises back toward air gaps, premise-based infrastructure, and sovereign deployments, and it is forcing a rethink of what the data center itself looks like.

Gillis, whose mandate at Cisco is to figure out the data center of the future, argues that AI infrastructure is fundamentally different from the networks built for VM and Kubernetes workloads. GPU clusters need a back-end network for direct GPU-to-GPU traffic, with programmable silicon that applies different load balancing and buffering than the front-end network, and the industry is heading toward co-packaged optics because copper cannot carry electrons fast enough at the speeds coming next.

What we cover in this episode

  1. Front-end vs. back-end networks for GPU clusters. GPU-to-GPU communication is a different class of networking from the traditional front-end network, with scale-up interconnects from GPU vendors and Ethernet-based scale-out fabrics where programmable Cisco silicon detects back-end traffic and applies different load balancing and buffer handling.
  2. The road to co-packaged optics. Somewhere beyond today's 800 gig speeds the industry cannot move electrons fast enough over copper, so optics will sit directly next to the silicon on a substrate, a shift the industry is still working out how to deliver at scale.
  3. Making AI infrastructure consumable for the enterprise. Most enterprises run hundreds or thousands of GPUs rather than giant AI factories, so Cisco's partnership with NVIDIA focuses on bringing lessons from the cutting edge into deployments enterprises can run in their own data centers, close to their data.
  4. The sovereignty dilemma: models never forget. Because knowledge learned by a model is infused into it rather than stored like a database field, IP leakage cannot be undone, which is driving air gaps, premise-based infrastructure, and a swing from cloud-first back to on-prem for sovereignty and trust reasons.
  5. Multi-cloud and the case for heterogeneity. A major cloud outage just before the conversation reinforced a long-running ONUG theme: diverse, multi-cloud infrastructure gives a better outcome than putting all your eggs in one basket.
  6. Federated analytics: from data lakes to data ponds. AI systems generate two to three orders of magnitude more operational data, so instead of ingesting everything into one costly data lake, Cisco and Splunk propose moving analytics to where data is generated and federating search across local repositories with no ingest.
  7. AI breaking out of the data center. Gillis expects AI to finally drive widespread edge computing through robotics and everyday AI systems, with the network as the common thread, and points to AI coding tools making developers several times more productive as the business force making infrastructure catch up fast.

The lines worth sharing

“When a model learns your secrets, it never forgets. You can't delete it, can't undo that, because it's not like a field in a database. It gets infused into the model, and that's it.”

Tom Gillis

“Instead of moving all that data to the analytics, let's flip it around: let's move the analytics closer to where the data is generated, and then federate search capability across it.”

Tom Gillis

“We're going to use these tools regardless of how ready the infrastructure is, so the infrastructure needs to catch up, and it needs to catch up fast.”

Tom Gillis

Common questions from this episode

Why can't data be deleted from an AI model after it has been trained on it?

Because a model does not store information the way a database stores a field. What the model learns gets infused into the model itself, so once sensitive data or intellectual property has been absorbed in training, it cannot be deleted or undone. That is why safeguards around what data reaches a model matter far more than after-the-fact remedies.

What is the difference between front-end and back-end networks in AI infrastructure?

The front-end network is the traditional enterprise network connecting servers, storage, firewalls, and load balancers to incoming traffic. The back-end network carries direct GPU-to-GPU communication, which behaves more like an interconnect than a general purpose network, with scale-up links from the GPU vendors for small clusters and Ethernet-based scale-out switching, using different load balancing such as packet spraying and different buffer sizes, for applications that need a thousand GPUs.

Why are enterprises moving AI workloads back on-prem instead of the cloud?

Three reasons come up in the episode: the data an AI system needs already lives in the enterprise data center, security and IP leakage concerns are sharper with AI because a model cannot unlearn what it absorbs, and there are economic and speed of light advantages to keeping compute next to the data. Sovereignty pressures add to this, with organizations in some regions requiring AI to run on their own infrastructure rather than on a hyperscaler.

What are co-packaged optics and why do AI data centers need them?

At the speeds AI networking is heading toward, beyond today's 800 gig links, electrons cannot move fast enough over copper wire. Co-packaged optics puts the optical components on a substrate directly next to the switch silicon, so the connection is optical rather than copper. It removes the parasitic effects of copper but creates new operational challenges, and the industry is still working out how to do it at scale.

How should enterprises handle the explosion of operational data from AI systems?

AI infrastructure generates two to three orders of magnitude more system data, such as traces, logs, alerts, and alarms, than today's model of ingesting everything into one central data lake can economically absorb. The approach described in the episode flips the model: keep data in local repositories where it is generated, whether a cloud storage bucket, a data warehouse, or a firewall log manager, move analytics closer to that data, and federate search across it without ingesting it at all.

Read the complete conversation

Full episode transcript · 16 minutes

when a model learns your secrets yeah it never forgets yeah you can't delete it can't undo that because it's it's not like a field in a database it gets infused into the the model and that's it it does it's you can't go after it hi everyone i'm your host nick lippis and welcome to the built for trust podcast where you get to hear from all the folks who are building and shaping ai enterprise infrastructure now let's get right into it with our guests tom thanks so much for being you know back on the built for trust podcast just finished your keynote you know up on the main stage yeah good fun really really great i tell you i love the house too congratulations to you all for drawing a crowd oh thank you yeah it's like actually our biggest crowd standing room only for you and well there's a reason yeah right this is this ai stuff is driving change that everybody needs to understand right and and onug is the place to learn in my opinion so oh thank god thank you for that and we try you know to do that i think you know it's like this is becoming uh over time i think it is now it's just only going to grow bigger it's really the epicenter of like where do you figure out how to build an enterprise ai yes you know company and i think that's really what's happening and your your keynote fed right into that because really you know it's like when i hear you know you speak and now cisco really speak you're really kind of rethinking and reimagining what infrastructure means and what it does that's my job at cisco you know that right and so uh and that's relatively new it's like in the last year but they've asked me to go figure out what does the data center of the future look like and it's going to look a lot different than the data center we know today you know yay you know let's get on it you know well you know we we did that three tier and now two tier you know and now there's like all the gpus scale out and it was like you know sort of bringing the cloud scale and then the the the infrastructure for for large-scale ai compute is dramatically different yeah than the infrastructure we've had for kubernetes-based apps or for vm-based applications yeah i think one thing that i get from the from the whole gpu experience around networking is how fundamentally different the protocols have got to be you know obviously ethernet had to be fundamentally rethought and um you know and redesigned uh for that and so how but we also know like that these systems you know these ai infrastructure systems the gpu clusters yeah um in our community you know the large enterprise there might be they may have a thousand gpus maybe some might have 10 000 gpus they're not of the scale of that you're not here most most enterprises are just starting on this journey and there's a lot of friction around it so one of the advantages we have at cisco is we have a multi-billion dollar business selling components to these monster ai factories you know like the big giant model trainers and we've learned a bunch of stuff about how to make this work at scale so so let's start with some of the basics yeah when we talk about a network the network that we know and love in the enterprise today that connects the servers to the storage the firewalls and load balancers we call that the front-end network it's traffic coming in but as you start to deploy gpus the gpus need to talk directly to each other so gpu to gpu communication is a really different class of networking and we refer to that as back-end network there's a scale up and a scale out component to the back-end network so scale up is going to be something that's really going to come from the gpu vendors yeah where they build interconnects that allow these gpus to talk directly to gpus at blistering speeds but that's going to go to you know 78 gpus something like that 100 gpus yeah but you're going to have applications that need a thousand gpus right and when you need a thousand gpus and you use called a scale out back-end network those that today we are doing with ethernet-based switches and so but the the the mode of operation the way we do load balancing uh you know we do packet spraying across the network the buffer sizes when you're running gpu to gpu it's more like a like an interconnect it's almost like a patch panel for a computer as opposed to a general purpose front-end network so we have a different mode of operating our silicon is programmable and so it knows oh i'm doing back-end network and then it'll apply different algorithms for doing traffic handling on the back end yeah then for the front end but the point is it's it's all the same cisco switch wow i didn't know like so the silicon actually changes wow it detects and we work with nvidia on that where they we have a handshake back and forth like okay this is this is back-end networking and we've got to run you know these different load balancing algorithms um uh and then the the different um uh buffer sizes to handle that traffic yeah and it's crazy fast traffic 800 gigs yeah yeah it's yeah we're kind of like you know 800 gig you know kind of uh speeds right now and it's like it's and that's not enough yeah and we're kind of like we're kind of in star track land yes with optics like in ethernet speeds so it's well so we could almost have a separate podcast on this which is at the cutting edge the top of the market to go from 800 gigs somewhere between 800 gigs 1.6 t and 3.2 t somewhere in that zone we're going to reach a point where we can't move electrons fast enough over copper wire and so we're going to have to go directly to optics yeah that's a whole nother realm of complexity so that would be kind of optic to optic gpu to gpu connection so the whole co-packaged meaning there's there's the silicon yeah then there's a optic sitting right next to it on a substrate so it's not running over copper wire and then what you plug into is an optical connector there is no silicon i mean there is no copper yeah the optical connects directly to the silicon into the silicon so like the a the um a to o conversion happens actually on chip on chip yes exactly but that's just going to create an operational and logistical challenge and that that now your silicon vendor your optics vendor has to be the same yeah right and your cabling is you know all optical it you know it's it's a big big big step and we you know we're still trying to solve a bunch of the technical challenges of how to make this work that's interesting i wonder so will now this kind of the substrate be kind of like um kind of optical modes you know that go right into the chip you know said and so you have kind of like both lasers and detectors like you know right there and it's an electrical signal that goes into the optic but it's on a laminate substrate so it so it doesn't have the parasitic effects of a copper wire so it's literally you put the the two two uh silicon die right next to each other yeah and that's how you can achieve the really high speed wow it's super complicated and you know sort of in full disclosure nobody has worked this out yet like we we the industry are still working on how to do co-packaged optics at scale so kind of good news is we're working on it you know the bad news is is that that doing ai networking can be very complicated yeah right now we're talking about hardware and we're talking about stuff that's three years but you know as cisco we have the advantage of sitting we kind of have a front row seat to like what the cutting edge looks like yeah we learn from that like the programmability we talked about in the switch and then we bring that into the enterprise to make it easy for the enterprise yeah and this is this is my sort of mission and my passion is how do we take this crazy stuff we're doing at the high end yeah where there's liquid cooling and these guys literally buy nuclear power plants right yeah yeah it's crazy crazy are you gonna buy nuclear power plants around your data center like i hope not right um uh and they have fixed space too yes yeah exactly right so so there's a lot of work that the industry needs to do to make ai infrastructure consumable for the enterprise right i can tell you we're on it yeah and this is the heart of our partnership with with nvidia is we're trying to make this stuff consumable yeah for the enterprise customer because enterprises are going to want to deploy this stuff in their data center because it's next to the data that they have security reasons economic reasons and speed of light issues where you want this stuff in your own data center yeah i think you know it's it's almost like we've seen a little bit of this story kind of play out in the cloud where there's a lot of kind of um hesitation there still is hesitation around like having workloads in the cloud from a security point of view ip leakage point of view and so forth and that's going to play itself out also in the ai space and as well it's it's the parallels are really similar um um however i think that the data leakage issue is even more challenging because there's an inherent tension with ai and that when a model learns your secrets yeah it never forgets yeah you can't delete it can't undo that because it's it's not like a field in a database it gets infused into the the model and that's it it does it's you can't go after it and so so protecting intellectual property in a world where where you know the models can absorb this data and kind of maybe they're not going to give your source code away but if it was going to be giving away some of the ideas that's actually the hard part of this right the intellectual property how do i ensure safeguards around that it's pretty muddy yeah right that's that's not assured yeah and so we do think that that you know having air gaps having premise-based infrastructure whether for sovereign reasons or intellectual property protection we think that's going to become increasingly important yeah and i think the whole data solving the data stewardship uh is really key and the ip leakage is like really one of the critical you know you know what's funny is like we've been doing this a long time right it used to be everything was in the cloud it was like tom how fast can you get your stuff in the cloud i'm racing to get all the cloud in the last two years it's gone the other way no no we want it on-prem yeah for sovereignty reasons for you know there's just a lack of trust unfortunately across international boundaries right so in europe they're saying hey great it's running on amazon but you know we need this running on our own infrastructure well even that like you know okay so this is um wednesday the 23rd on monday the 21st of october aws went down you know and it's like i am personally so happy that we don't want we don't run our registration system on aws if that would have went down like we got like almost like 200 registrations yesterday yeah you know it's like business critical yeah for us that was like business critical so yeah yeah and i know it's like you know we had this whole discussion we had a board uh conversation yesterday about this and like it was back in 2014 yeah 2014 that we were starting to talk about multi-cloud yeah and like a lot of people were like okay well why multi-cloud really everyone all you really need is that do aws and so it's like these little reminders you know uh really are kind of helpful um you know um you know to remind us that this is really a multi-cloud kind of world heterogeneity is always good right you know and you know sort of a diverse infrastructure is going to give you a better outcome than all your eggs in one basket common sense there yeah yeah so um one other thing that you talked about that i wanted to kind of hit on so repackaging or not repackaging but like packaging products for the enterprise marketplace for ai infrastructure yes that's that's on its way the other is the operational model you you gave a data point that i was really blown away with basically you're saying like around system data these are trace logs alerts alarms so forth like two to three multiple yeah orders of magnitude orders of magnitude like you know that's you know that's a thousand times you know bigger you know tend to a thousand times more data right which is just incredible it's incredible it's you know that's huge and so the way that we've always dealt with you know operational data is the haystack model right you know so it's like you put all this potentially somebody throws a needle in that haystack and then you have all the engineers yeah yeah so but the haystack lives in one place on a data lake yeah that's just not going to work yeah first of all it's it's it's almost breaking today because we you know we own splunk right and so when i talk to splunk customers they're all like oh my god it's really expensive well yeah it is expensive even if i give you the software for free yeah ingest is expensive it burns cpu and it burns storage yeah and so so if you take that model and you throw a thousand times more data at it like it's not going to work and so this this is sort of the big deal between cisco and spunk combining is that our view is instead of moving all that data to the analytics let's flip it around let's move the analytics closer to where the data is generated yeah and then federate search capability across it so you instead of a data lake you've got a series of data pawns yeah local data repositories that could be an s3 bucket right out on amazon it could be data living in snowflake but it could be your cisco firewall log manager is gathering that data locally we can search it we can analyze it with no ingest yeah if you're a splunk customer you hear this like like not in just i'm not saying the ingest is free i'm saying you don't ingest it at all yeah that's a big deal and i think that's the only way to scale to this ai world we're moving to and you know we tend to focus on ai applications in the data center yeah but what i think is really cool and again we could do another podcast on is ai is going to really break out of the data center and i think it's going to be the thing that drives edge computing yeah you know like i feel like the industry was all waiting for this revolution in edge computing and it kind of happened in pockets retail right you know it's doing it so hotels and iot a little bit iot a little bit you know factories manufacturing but it wasn't widespread i think ai is going to be the thing that makes it widespread because we're going to find robotics and ai systems that are in our everyday life the robotic coffee maker the robotic bartender you know these are real things yeah thank god robotic bartender right like actually hey i'm a big gardener yeah i want a robot help me garden definitely yeah it'll go out and it can tell the weeds from the you know the plants yeah it's like you know just prepare the beds you know yes you know all that it's gonna happen and so data is going to be all around us and we just have to have a different architecture for that what's the common thread to all this stuff it's the network yeah right which is why i think onug is coming back to what we talked about in the beginning like onug is is the place to learn about these massive changes that are coming at us not in years you know this is happening in months yeah right it's um it's it's scary fast and i don't think we don't i don't think the industry's really comprehended it you know it's like when we did the polling you know and the polling data that i showed like in the opening it's like there's like people haven't wrapped their mind around okay this that the infrastructure needs to change they haven't put budget for it but they believe it's all going to happen they just don't have the time scale right you know and so well you know what's going to drive that is the business yeah so case in point i think one of the killer apps for ai is ai coding yeah you know we have 12 000 developers in my organization 12 000 developers and when i look at these tools windsurf codex from open ai like they're unbelievable like like this is like my developers can be not not 20 percent more effective 200 more effective 300 400 like this is this surge and productivity unbelievable it's unbelievable so we're going to use these tools regardless of how ready the infrastructure is so the infrastructure needs to catch up and it needs to catch up fast yeah interesting times we live in right for sure you know but okay i want before we close i want to basically kind of recap so you're thinking uh operational data is going to be more distributed and federated yes um and that um doing the analytics on all that data will be centralized but it'll be kind of utilizing agents to interact with multiple agents that are distributed all across the the analytics will be somewhat centralized it's it's very difficult to decentralize you know complex analytics because you're looking across multiple things but you can process some of it locally yeah and then some of it is is done centrally but in a very efficient fashion right so you just can't have the current model where throw it all in one big bucket we'll go look for it it just doesn't work yeah yeah i'm like it's really it's ingenious and it makes so much sense and especially with you know a hundred to a thousand times like more yes you know data coming in how do you sift through that yeah yeah let's get ready yeah okay thanks tom thanks all right this was really great take care okay thanks everyone thanks everyone you

The episode summary, topic list, and questions on this page were generated with AI assistance from the episode recording and show notes.