Ain’t livestreaming grand? We had a few rough-around-the-edges moments in today’s episode of Ziti TV (I worked on the demo a *little* too late last night), but the important thing is: THE DEMO WORKED.
The demo built on the previous episode’s work where we took the OpenZiti Appetizer (a service that displayed whatever text you feed it on a public web page via a secure zero-trust OpenZiti overlay) and added a RoBERTa classifier to it to prevent users from entering offensive messages for the Appetizer to display. There are lots of people out there who seem to not have aged past 12, and that’s what this is supposed to handle.
RoBERTa classifiers are good at handling profanity and direct insults like “You are too stupid to understand this”. It *will* miss more subtle jabs like “Nobody wants you here and everyone knows it”, but try saying that to a cop who’s writing you a speeding ticket, and they will classify it very differently.
So in this week’s, the classifier keeps its job, but when it isn’t sure, it asks for a second opinion. LLM Gateway, the OpenZiti-based open source OpenAI-compatible gateway, decides which LLM gives that opinion: a small model on your own machine for everyday messages, and Claude for the subtle ones. You won’t change a line of the Appetizer’s code.
I did the demo without much chance for a rehearsal, and I’ll post a GitHUb repo with all the steps and material you need to try it out for yourself. But in the meantime, catch Clint and I work our way through the demo!
Pictured above: Another one of my totally organic, AI-free, hand-drawn “back of a literal envelope” diagrams showing the “limit the blast radius” benefits of network segmentation and microsegmentation.
Jack Poller, NetFoundry’s VP Product Marketing, kicked off a six-part microsegmentation series with the gap that nobody in the space likes to say out loud.
An Omdia survey of 352 security decision-makers found that:
99% of organizations are implementing or planning microsegmentation, and
only 9% had protected more than 80% of their critical systems.
Half of them had been through a lateral-movement attack in the previous year, which is the exact thing microsegmentation is supposed to stop.
Jack argues that the conventional playbook is the problem: Discover every flow, model the policy for the whole environment, simulate, then enforce. Each phase gates the next, and phase one never finishes because reality collides with your best laid plans:
Applications talk to more things than anyone documented and
dependencies shift while you’re charting them.
In the end, your initiative languishes in discovery for quarters and the crown jewels stay exactly as exposed as they were on day one.
The proposed fix? Turn the order upside down! Protect the single most critical asset first, one workload at a time, starting now.
Of course, this approach works only if enforcement stops depending on network location. As long as policy is based on IP ranges and zones, you’re back to needing the full topology before you can trust a rule. If you base policy on cryptographic identity, you can wrap one asset in tight policy today without having mapped what’s around it.
I’m just sayin’: if you’re trying to promote AI as a beneficial thing, maybe the whole Terminator 2 motif is something you might want to avoid. Those AIs, Arnie’s T-800 excluded, were not our friends!
In the meantime, let’s enjoy the classic line from that film:
If there’s one thing that Donald Trump and his sycophants love to do, it’s name or rebrand things to name them after hinself, make them sound more “extra”, or both. Consider:
John McCarthy, coiner of the term “Artificial Intelligence”, creator of the Lisp programming language and garbage collection, influencer on ALGOL, which influences most programming languages.
The executive order, titled INAUGURATING THE ERA OF SUPER INTELLIGENCE, basically declares that “AI”, a term that we’ve been using since 1955 when John McCarthy coined it, should now be referred to as “Super Intelligence” or “SI” in executive-branch communications.
It says that agencies of the federal executive branch must use the new terms in official correspondence, public communications, websites, reports, policy documents and other non-statutory documents. Thankfully (because it would mean a lot of pointless work otherwise), existing regulations, presidential actions, contracts, grants and historical documents don’t have to be changed.
It also means that any content I make aimed at a U.S. government buyers or techies (especially tutorials, presentations, and documentation) will need to include the “SI” terms for both practical (for example, search terms) and political (“We call it ‘SI’ here, son.”). I’m already not looking forward to this.
For now, “SI” means exactly what “artificial intelligence” already means under 15 U.S.C. § 9401(3).; only the name has changed.
Sometime in the next 60 days, Michael Kratsios, the Director of the White House Office of Science and Technology Policy (OSTP), will have to propose legislative language for a federal definition of SI. It’ll have to say whether that definition should modify or replace the statutory definition of AI, and include any conforming amendments and follow-on executive actions. If such changes are made, expect them to be dumb.
Unsurprisingly, it didn’t take long for the first kiss-ass to use the term in Trump’s presence:
Super Cyber Friday is about cybersecurity practitioners and vendors, and every episode has a title that follows a “Hacking [Topic]” format. This particular episode’s title is Hacking Microsegmentation for Machine Workloads, and features:
Galeal Zino, NetFoundry CEO (and my boss’ boss, and therefore the best damned skip-level in the world)
Howard Holton, Founder and Principal Counsel at Phronia Counsel, an independent technology analyst and advisory firm
If you’re wondering what microsegmentation is, here’s the short version:
Microsegmentation is the practice of carving your network into tiny zones so that when something gets compromised (and something will), the damage stays in one small room instead of spreading through the whole house.
Everyone agrees it’s a good idea, almost nobody enjoys doing it, and few actually do it. It’s the brushing and flossing of cybersecurity.
Before the notes, some disclosure
You should know:
I do developer advocacy work at NetFoundry in exchange for money, health insurance, and to convince people I’m more than just an accordion-playing reprobate.
Once again, Galeal is my boss’s boss. And he’s a great boss’ boss!
NetFoundry sponsored this episode.
Feel free to apply the appropriate amount of skepticism to this writeup, but keep in mind that this episode features both a CEO who builds security stuff with an analyst who’s had to live with the results. The episode’s a little more blunt than your typical PR piece.
The tl;dr
C’mon, this article isn’t that long! But still, if you want a very quick summary, here it is…
The old way
What machine workloads need
What you segment by
IP addresses, subnets, CIDR blocks, VLANs
A cryptographic identity for every machine, plus attestation
Where enforcement happens
Firewalls and middleboxes at the edge
Inside the application, via an identity-driven overlay
How AI agents get treated
“It’s basically an employee” or “It’s basically an app”
As a non-human actor with explicitly bounded access
What failure looks like
Firewall surgery, configuration bloat, shelfware
A contained blast radius
What the board says
“We better never get hacked.”
“How bad will it be when we get hacked?”
“Early microsegmentation sucked so bad…”
Howard set the tone in the first minute, when David asked for his microsegmentation pet peeve (0:08):
“Early microsegmentation sucked so bad that it’s hard to get people to listen when you talk about microsegmentation.”
For the record, Howard is a big microsegmentation fan. What drives him (and us at NetFoundry too) up the wall is the feedback he hears most often: “Nope, we did that before, we’re not doing it again.” People tell him it was too complicated and a waste of money. He hates that feedback because, in his words, “it’s just wrong.”
Later in the show (7:25), he explained how early rollouts earned that reputation. Teams would get frustrated during the discovery stage, put some segments in place, and then cause an outage by blocking something that only ran once every 30 days and that nobody had planned for. After that happened a few times, they’d shelf whatever microsegmentation tool they were using. His summary: “Complexity always bites you in the ass.”
If you’ve ever been anywhere near a rollout like that, you probably just winced. The problem wasn’t the idea. Least privilege and blast-radius reduction are good ideas! The problem was that vendors tried to implement them with the tools that just happened to be conveniently lying around: IP addresses, subnets, VLANs, and firewalls.
When an audience member asked whether application-level or network-level microsegmentation is more effective, Galeal didn’t mince words (9:05):
“Network-level is DOA: dead on arrival. It’s an oxymoron to begin with… It’s spilled milk, and you can’t put it back in the bottle. If you’re going to try and use IP addresses, VLANs, and firewalls to do ‘microseg,’ good luck to you. It’s not going to happen.”
That’s coming from someone who’s been doing network engineering for 30 years. Here’s the core problem, especially for machines: an IP address tells you where something is, not what it is. In a world of containers, autoscaling, and serverless functions, the IP address your billing service has this morning might belong to something completely different by the time lunch rolls around. Writing security policy against IP addresses is like keeping track of your friends by remembering which seats they sat in the last time you all went to the movies. Or, as Galeal put it later in the show (35:10), “IP addresses are not identities.”
Howard then described what happens when you try to brute-force it anyway (10:03):
“Microsegmentation and segmentation are not the same thing… If you’re like, ‘Well, I’m going to create 437 subnets in my network and I’m going to create 300 VLANs and I’m going to turn on host-based firewalls,’ you’ve just created a nightmare that no one is ever going to be able to manage after you. And you have failed job number one, which is make sure that the next person can be as successful as you are.”
“Make sure that the next person can be as successful as you are.” I want that on a poster, and not just for network engineers. It applies to code, documentation, and pretty much anything you build that someone else will inherit.
And if you try to do it the classic way regardless, making your firewalls enforce all that traffic between your services? According to Galeal (29:48), that’s the first lesson you’ll learn: “You will melt your firewalls.”
Credentials aren’t identity
My first bearer tokens.
My favorite moment in the episode came from an audience question. Sierra Montgomery asked (11:05):
“If a compromised workload acquires valid credentials and begins behaving like a legitimate service, what signal does your microsegmentation architecture use to distinguish legitimate machine-to-machine communication from lateral movement?”
Galeal’s answer was one word: Attestation. His reasoning: a credential doesn’t prove whether Howard, David, or an AI agent should have it. Credentials, he said, are “necessary but not sufficient.”
Explaining bearer tokens is something that goes back to my first article for Auth0, which was also my “take-home assignment” in the job interview process. The concept always needed the most careful explaining, even though the name told you everything: whoever bears the token gets the access. The token doesn’t know or care who’s holding it. It’s more like cash than a credit card.
That’s fine as long as you know where the token is. The trouble starts when you don’t. A credential proves that something possesses a secret. It doesn’t prove that the thing should have that secret, and it says nothing about whether the machine presenting it is in the state it’s supposed to be in. Attestation fills that gap: it’s verifiable evidence that the workload is what it claims to be, and that it’s running where and how it’s supposed to.
Take the pwning like a champ
At a DEF CON party a long time ago, a few friends and I came up with a joke talk title for the following year’s edition of the conference: “Security is for lightweights. Take the pwning like a champ.”
But behind the joke was an important idea: every boxer gets hit. What makes a champ isn’t never getting punched; it’s being able to take the punch and stay on your feet. Security works the same way. Sooner or later, one of your credentials is going to end up in the wrong hands, whether it’s through a phished password, a leaked API key, or a token that got copied out of a log file. Back then, planning on getting pwned was a punchline. These days, it’s a design principle, and it’s exactly where Galeal starts.
The DEF CON joke came to my mind when an audience member asked how a platform can alert on a compromised token moving laterally (16:06). Galeal’s answer was in the spirit of “Take the pwning like a champ”: Assume that tokens will be compromised, and design your system so that a token by itself doesn’t grant access.
If your architecture opens the door and grants network reachability before it verifies identity, an ill-gotten token can be a starting point to explore your network. The “Take the pwning like a champ” approach that NetFoundry takes verifies identity first and uses policies to spells out which identities can talk to which services. A stolen token doesn’t open any new doors, because the path the attacker wants isn’t accessible to them.
AI agents are chaos monkeys nobody scheduled
History time! If you were a developer around the time the iPad came out, you might remember Chaos Monkey, Netflix’s tool designed to randomly shut down servers in production. The idea behind it could be summarized as “The random shutdowns will continue until resiliency improves”. Netflix’s dev teams were forced to build in such a way that their systems would survive failure. It’s a brilliant (if sadistic) idea, and it works because Netflix chose to unleash it deliberately, with rules.
This isn’t all too different from a pattern Galeal described (28:56): taking an AI agent and saying, “Hey, cool. Here’s some API keys. Here’s the internet. Here’s some enterprise resources. Go do something useful.”
Howard’s response: “I see that like 40 times a week.” That’s a chaos monkey too, except nobody scheduled it, nobody wrote the rules, and it will explain its reasoning very confidently afterward.
My last job prior to my current Developer Advocate gig at NetFoundry was optimizing an MCP server, so this isn’t a hand-wavey “some customer has this issue” thing to me. Every tool or function you expose through an MCP server is something an agent can decide to call, whenever it wants, for reasons you didn’t anticipate. Its network traffic doesn’t follow a script.
When an audience member asked how to segment an AI agent whose needed connections change with each task, Howard’s advice was (27:27):
“If you try to design something that is entirely flexible, what you’re going to end up with is opening yourself up for AI chaos in your network… If you do it the other way around and hope that your microsegmentation tool set is going to keep up with your AI, you’re basically saying, ‘I want microseg to follow my chaos monkey.’ Do it the other way around. Use microsegmentation to restrict the chaos monkey.”
Galeal’s answer to the same question (26:43) supplies the other half: The old bag of tricks isn’t going to work on agentic flows. Treat agents as identities, define a graph of what each one can talk to and under what conditions, and have the visibility and enforcement to make sure they stay on that graph.
That’s the right mental model. An AI agent isn’t a human employee sitting behind your SSO portal, and it isn’t a cron job that does the same thing at the same time every night. It’s a whole new thing.
As Howard put it later in the show, it isn’t another application, and it doesn’t act like a human either. It needs boundaries defined up front, including things like:
A specific identity,
ashort list of things it’s allowed to reach, and i
solation that travels with it whether it’s running in AWS, on bare metal, or in an OT facility.
What this looks like if you write code
Time for me to put my work hat on for a minute. The table above mentions “an identity-driven overlay” enforced “inside the application,” which is a mouthful, so here’s what it means in practice.
OpenZiti is the open source zero trust networking platform that NetFoundry builds, and one of the things it lets you do is embed zero trust directly into your application with an SDK. Your app gets its own cryptographic identity, and it can only reach the services that policy says it’s allowed to reach. On the other side, services don’t need open inbound ports at all, so there’s nothing sitting on the network for an attacker to scan, probe, or point a stolen token at.
Galeal’s “kill switch” (more on that below) gets pretty concrete here too: if an identity is compromised, you revoke it or change the policy, and its paths go away. No firewall surgery required.
If you want to try it yourself, the docs and quickstarts are at openziti.io.
What do you tell the board?
Near the end, David read an audience question: Which metric best shows leadership that microsegmentation is working? Galeal boiled his answer down to three questions (39:11):
The graph: Can you answer the question “What identity can talk to what identity, according to what policy?” (Note that he said identity, not IP address.)
The kill switch: When something unexpected happens, where’s the kill switch that lets you deal with it without compromising uptime, human safety, or business continuity?
Centralized governance: How much of all this can you see and govern from one place, instead of going to a bunch of different environments and firewalls?
Then Howard put on his gloves (41:27): “I couldn’t disagree more.” Speaking as a CISO, he said:
“What my board said 10, 15 years ago was, ‘We better never get hacked.’… Today they say, ‘How bad will it be?’ That is their question to me. That is the question they want me to answer in every board meeting they invite me to.”
My answer would be “Take the pwning like a champ”.
David pointed out that the two answers aren’t really in conflict: Galeal was describing what you measure for yourself and your security team, and Howard was describing what you tell the board. Galeal agreed. Howard then explained why the board version has to be so compressed (43:29):
“I get one slide as a CISO… I get like five minutes… I make that one slide tell them how I’m spending money to make it less bad than the last time they gave me money.”
Ah, the dreaded “You get one, and only one, slide” directive for board meeting presentations. Replace “board” with “VP of Engineering” or “whoever approves your budget,” and it’s still excellent advice.
My takeaway
If I had to boil the whole episode down to one idea, it wouldn’t be about AI, firewalls, or attestation. It would be Howard’s “job number one”: make sure the next person can be as successful as you are.
Galeal landed in the same place from a different direction. When David asked why he started NetFoundry (24:00), he said that after years of beating his head against walls like microsegmentation, he wanted to tilt the playing field so that the next person who has to solve these problems doesn’t have to.
That’s the real case against building microsegmentation out of subnets and VLANs. Even if you get it working, you’ve built something nobody else can understand, let alone maintain.
A policy that says “the billing service can talk to the orders database, and the AI agent can talk to these three APIs and nothing else” is something the next person can read, reason about, and change without breaking everything. And when some of the things doing the talking are AI agents that nobody can fully predict, that kind of clarity is what keeps the chaos monkey in its cage.
Go watch the episode, and if you’re a developer who wants to see what identity-first networking looks like from the inside, take OpenZiti for a test drive!
I was working on one of my “back of the envelope” diagrams for a post on microsegmentation and decided to draw the one pictured above as a warm-up exercise. It turned out well enough that I thought it was worth posting on its own. You might find it useful with clients when they look at you with puzzlement when you talk about “north-south” versus “east-west” in network security.
You can also use the text below when explaining the concepts; feel free to “copy, paste, and adjust to taste”:
North-south crosses the boundary of your environment. A browser hitting your web tier, an API call from outside, your service calling a third party. Up and down the diagram, in and out.
East-west stays inside. Web tier to app tier, app tier to database, service to service, pod to pod. Sideways across the diagram, and it typically never touches your perimeter controls at all.
The entire traditional security stack is pointed at north-south. Firewall, WAF, load balancer, the DMZ, the VPN concentrator, most of the alerting.
East-west traditionally get an implicit pass, because it’s inside, and inside was assumed to be fine. If this were true, we wouldn’t have expressions like “It was an inside job!”
An attacker who phishes a laptop or pops a container has already cleared every north-south control. From that point on they’re moving east-west, which is the direction that typically has the least in the way. The initial breach is usually small. The blast radius is what turns it into an incident, and the blast radius is all east-west.
That’s the gap microsegmentation exists to close: putting authorization on the connections between your own workloads, not just on the ones crossing your perimeter.
If you’ve spent time with theOpenZiti Appetizer, you may have noticed this response:
you sent a message, but it can’t be qualified at this time for offensiveness
There’s a hate-speech classifier on the path every message takes, and it’s been waiting for a classifier to talk to. On theSeptember 18 Ziti TV episode, I got curious about it and went down the “How do I get this thing to work?” rabbit hole. It turned out to be a nice tour of how OpenZiti authorization actually works!
Here’s the fun part. The first thing you notice is that the classifier URL has no instance prefix, while every other service name in the demo is scoped. That looked like the answer, but it wasn’t.Prepare tags the Appetizer’s own identity with an unprefixed classifier-clients attribute, deliberately, in a file where everything else goes through scopedName. The dial side was built with one shared classifier in mind rather than one per instance. The URL is right as written.
What’s left to build is the network side: the service and a dial policy. The application was always correct, and it’s been accurately reporting that it couldn’t reach something that wasn’t there yet.
In the end, the fix was a service and two policies, and (yay!) no changes to the code at all.
In the exercise, you add a real classifier (a RoBERTa model fine-tuned on OffensEval tweets) and host it as a dark service. By the end you have two processes talking by service name, with ss -ltn showing nothing listening inside the classifier container while the controller shows it actively hosting. Then you take access away with one policy change and watch the caller get “not found” for a service that’s running perfectly well three feet away.
The exercise shouldn’t take longer than a half hour, and it all runs in Docker containers, which means you don’t need Python, Go, or even the ziti CLI installed locally. I wrote it for people new to OpenZiti, with the diagnosis reasoning spelled out rather than just the commands, since that part transfers to debugging your own dials.
A few things I’d have liked to know going in, which I haven’t seen written down elsewhere:
Flask’s built-in dev server can’t host a ziti service. The bind succeeds and a terminator registers, but it never accepts. werkzeug polls the listening fd through selectors, and a ziti fd isn’t pollable that way. Callers see EOF and the Python side logs nothing. waitress works, and it’s what the SDK’s own Flask sample uses.
The SDK enumerates services at startup and doesn’t re-check per dial, so a service created after your process starts stays invisible until it refreshes.
Thanks to Clint for the Appetizer, which is a genuinely good teaching demo, and for co-hosting the episode this came from.
This is the first in a series! I’m turning Ziti TV episodes into standalone repos people can work through on their own. Episodes:https://www.youtube.com/@OpenZiti/streams
Issues and PRs welcome, especially if you hit something the troubleshooting section misses. Help make it better!
Once again, the repo is here: https://github.com/AccordionGuy/openziti-appetizer-classifier