Ain’t livestreaming grand? We had a few rough-around-the-edges moments in today’s episode of Ziti TV (I worked on the demo a *little* too late last night), but the important thing is: THE DEMO WORKED.
The demo built on the previous episode’s work where we took the OpenZiti Appetizer (a service that displayed whatever text you feed it on a public web page via a secure zero-trust OpenZiti overlay) and added a RoBERTa classifier to it to prevent users from entering offensive messages for the Appetizer to display. There are lots of people out there who seem to not have aged past 12, and that’s what this is supposed to handle.
RoBERTa classifiers are good at handling profanity and direct insults like “You are too stupid to understand this”. It *will* miss more subtle jabs like “Nobody wants you here and everyone knows it”, but try saying that to a cop who’s writing you a speeding ticket, and they will classify it very differently.
So in this week’s, the classifier keeps its job, but when it isn’t sure, it asks for a second opinion. LLM Gateway, the OpenZiti-based open source OpenAI-compatible gateway, decides which LLM gives that opinion: a small model on your own machine for everyday messages, and Claude for the subtle ones. You won’t change a line of the Appetizer’s code.
I did the demo without much chance for a rehearsal, and I’ll post a GitHUb repo with all the steps and material you need to try it out for yourself. But in the meantime, catch Clint and I work our way through the demo!
I’m just sayin’: if you’re trying to promote AI as a beneficial thing, maybe the whole Terminator 2 motif is something you might want to avoid. Those AIs, Arnie’s T-800 excluded, were not our friends!
In the meantime, let’s enjoy the classic line from that film:
If there’s one thing that Donald Trump and his sycophants love to do, it’s name or rebrand things to name them after hinself, make them sound more “extra”, or both. Consider:
John McCarthy, coiner of the term “Artificial Intelligence”, creator of the Lisp programming language and garbage collection, influencer on ALGOL, which influences most programming languages.
The executive order, titled INAUGURATING THE ERA OF SUPER INTELLIGENCE, basically declares that “AI”, a term that we’ve been using since 1955 when John McCarthy coined it, should now be referred to as “Super Intelligence” or “SI” in executive-branch communications.
It says that agencies of the federal executive branch must use the new terms in official correspondence, public communications, websites, reports, policy documents and other non-statutory documents. Thankfully (because it would mean a lot of pointless work otherwise), existing regulations, presidential actions, contracts, grants and historical documents don’t have to be changed.
It also means that any content I make aimed at a U.S. government buyers or techies (especially tutorials, presentations, and documentation) will need to include the “SI” terms for both practical (for example, search terms) and political (“We call it ‘SI’ here, son.”). I’m already not looking forward to this.
For now, “SI” means exactly what “artificial intelligence” already means under 15 U.S.C. § 9401(3).; only the name has changed.
Sometime in the next 60 days, Michael Kratsios, the Director of the White House Office of Science and Technology Policy (OSTP), will have to propose legislative language for a federal definition of SI. It’ll have to say whether that definition should modify or replace the statutory definition of AI, and include any conforming amendments and follow-on executive actions. If such changes are made, expect them to be dumb.
Unsurprisingly, it didn’t take long for the first kiss-ass to use the term in Trump’s presence:
Super Cyber Friday is about cybersecurity practitioners and vendors, and every episode has a title that follows a “Hacking [Topic]” format. This particular episode’s title is Hacking Microsegmentation for Machine Workloads, and features:
Galeal Zino, NetFoundry CEO (and my boss’ boss, and therefore the best damned skip-level in the world)
Howard Holton, Founder and Principal Counsel at Phronia Counsel, an independent technology analyst and advisory firm
If you’re wondering what microsegmentation is, here’s the short version:
Microsegmentation is the practice of carving your network into tiny zones so that when something gets compromised (and something will), the damage stays in one small room instead of spreading through the whole house.
Everyone agrees it’s a good idea, almost nobody enjoys doing it, and few actually do it. It’s the brushing and flossing of cybersecurity.
Before the notes, some disclosure
You should know:
I do developer advocacy work at NetFoundry in exchange for money, health insurance, and to convince people I’m more than just an accordion-playing reprobate.
Once again, Galeal is my boss’s boss. And he’s a great boss’ boss!
NetFoundry sponsored this episode.
Feel free to apply the appropriate amount of skepticism to this writeup, but keep in mind that this episode features both a CEO who builds security stuff with an analyst who’s had to live with the results. The episode’s a little more blunt than your typical PR piece.
The tl;dr
C’mon, this article isn’t that long! But still, if you want a very quick summary, here it is…
The old way
What machine workloads need
What you segment by
IP addresses, subnets, CIDR blocks, VLANs
A cryptographic identity for every machine, plus attestation
Where enforcement happens
Firewalls and middleboxes at the edge
Inside the application, via an identity-driven overlay
How AI agents get treated
“It’s basically an employee” or “It’s basically an app”
As a non-human actor with explicitly bounded access
What failure looks like
Firewall surgery, configuration bloat, shelfware
A contained blast radius
What the board says
“We better never get hacked.”
“How bad will it be when we get hacked?”
“Early microsegmentation sucked so bad…”
Howard set the tone in the first minute, when David asked for his microsegmentation pet peeve (0:08):
“Early microsegmentation sucked so bad that it’s hard to get people to listen when you talk about microsegmentation.”
For the record, Howard is a big microsegmentation fan. What drives him (and us at NetFoundry too) up the wall is the feedback he hears most often: “Nope, we did that before, we’re not doing it again.” People tell him it was too complicated and a waste of money. He hates that feedback because, in his words, “it’s just wrong.”
Later in the show (7:25), he explained how early rollouts earned that reputation. Teams would get frustrated during the discovery stage, put some segments in place, and then cause an outage by blocking something that only ran once every 30 days and that nobody had planned for. After that happened a few times, they’d shelf whatever microsegmentation tool they were using. His summary: “Complexity always bites you in the ass.”
If you’ve ever been anywhere near a rollout like that, you probably just winced. The problem wasn’t the idea. Least privilege and blast-radius reduction are good ideas! The problem was that vendors tried to implement them with the tools that just happened to be conveniently lying around: IP addresses, subnets, VLANs, and firewalls.
When an audience member asked whether application-level or network-level microsegmentation is more effective, Galeal didn’t mince words (9:05):
“Network-level is DOA: dead on arrival. It’s an oxymoron to begin with… It’s spilled milk, and you can’t put it back in the bottle. If you’re going to try and use IP addresses, VLANs, and firewalls to do ‘microseg,’ good luck to you. It’s not going to happen.”
That’s coming from someone who’s been doing network engineering for 30 years. Here’s the core problem, especially for machines: an IP address tells you where something is, not what it is. In a world of containers, autoscaling, and serverless functions, the IP address your billing service has this morning might belong to something completely different by the time lunch rolls around. Writing security policy against IP addresses is like keeping track of your friends by remembering which seats they sat in the last time you all went to the movies. Or, as Galeal put it later in the show (35:10), “IP addresses are not identities.”
Howard then described what happens when you try to brute-force it anyway (10:03):
“Microsegmentation and segmentation are not the same thing… If you’re like, ‘Well, I’m going to create 437 subnets in my network and I’m going to create 300 VLANs and I’m going to turn on host-based firewalls,’ you’ve just created a nightmare that no one is ever going to be able to manage after you. And you have failed job number one, which is make sure that the next person can be as successful as you are.”
“Make sure that the next person can be as successful as you are.” I want that on a poster, and not just for network engineers. It applies to code, documentation, and pretty much anything you build that someone else will inherit.
And if you try to do it the classic way regardless, making your firewalls enforce all that traffic between your services? According to Galeal (29:48), that’s the first lesson you’ll learn: “You will melt your firewalls.”
Credentials aren’t identity
My first bearer tokens.
My favorite moment in the episode came from an audience question. Sierra Montgomery asked (11:05):
“If a compromised workload acquires valid credentials and begins behaving like a legitimate service, what signal does your microsegmentation architecture use to distinguish legitimate machine-to-machine communication from lateral movement?”
Galeal’s answer was one word: Attestation. His reasoning: a credential doesn’t prove whether Howard, David, or an AI agent should have it. Credentials, he said, are “necessary but not sufficient.”
Explaining bearer tokens is something that goes back to my first article for Auth0, which was also my “take-home assignment” in the job interview process. The concept always needed the most careful explaining, even though the name told you everything: whoever bears the token gets the access. The token doesn’t know or care who’s holding it. It’s more like cash than a credit card.
That’s fine as long as you know where the token is. The trouble starts when you don’t. A credential proves that something possesses a secret. It doesn’t prove that the thing should have that secret, and it says nothing about whether the machine presenting it is in the state it’s supposed to be in. Attestation fills that gap: it’s verifiable evidence that the workload is what it claims to be, and that it’s running where and how it’s supposed to.
Take the pwning like a champ
At a DEF CON party a long time ago, a few friends and I came up with a joke talk title for the following year’s edition of the conference: “Security is for lightweights. Take the pwning like a champ.”
But behind the joke was an important idea: every boxer gets hit. What makes a champ isn’t never getting punched; it’s being able to take the punch and stay on your feet. Security works the same way. Sooner or later, one of your credentials is going to end up in the wrong hands, whether it’s through a phished password, a leaked API key, or a token that got copied out of a log file. Back then, planning on getting pwned was a punchline. These days, it’s a design principle, and it’s exactly where Galeal starts.
The DEF CON joke came to my mind when an audience member asked how a platform can alert on a compromised token moving laterally (16:06). Galeal’s answer was in the spirit of “Take the pwning like a champ”: Assume that tokens will be compromised, and design your system so that a token by itself doesn’t grant access.
If your architecture opens the door and grants network reachability before it verifies identity, an ill-gotten token can be a starting point to explore your network. The “Take the pwning like a champ” approach that NetFoundry takes verifies identity first and uses policies to spells out which identities can talk to which services. A stolen token doesn’t open any new doors, because the path the attacker wants isn’t accessible to them.
AI agents are chaos monkeys nobody scheduled
History time! If you were a developer around the time the iPad came out, you might remember Chaos Monkey, Netflix’s tool designed to randomly shut down servers in production. The idea behind it could be summarized as “The random shutdowns will continue until resiliency improves”. Netflix’s dev teams were forced to build in such a way that their systems would survive failure. It’s a brilliant (if sadistic) idea, and it works because Netflix chose to unleash it deliberately, with rules.
This isn’t all too different from a pattern Galeal described (28:56): taking an AI agent and saying, “Hey, cool. Here’s some API keys. Here’s the internet. Here’s some enterprise resources. Go do something useful.”
Howard’s response: “I see that like 40 times a week.” That’s a chaos monkey too, except nobody scheduled it, nobody wrote the rules, and it will explain its reasoning very confidently afterward.
My last job prior to my current Developer Advocate gig at NetFoundry was optimizing an MCP server, so this isn’t a hand-wavey “some customer has this issue” thing to me. Every tool or function you expose through an MCP server is something an agent can decide to call, whenever it wants, for reasons you didn’t anticipate. Its network traffic doesn’t follow a script.
When an audience member asked how to segment an AI agent whose needed connections change with each task, Howard’s advice was (27:27):
“If you try to design something that is entirely flexible, what you’re going to end up with is opening yourself up for AI chaos in your network… If you do it the other way around and hope that your microsegmentation tool set is going to keep up with your AI, you’re basically saying, ‘I want microseg to follow my chaos monkey.’ Do it the other way around. Use microsegmentation to restrict the chaos monkey.”
Galeal’s answer to the same question (26:43) supplies the other half: The old bag of tricks isn’t going to work on agentic flows. Treat agents as identities, define a graph of what each one can talk to and under what conditions, and have the visibility and enforcement to make sure they stay on that graph.
That’s the right mental model. An AI agent isn’t a human employee sitting behind your SSO portal, and it isn’t a cron job that does the same thing at the same time every night. It’s a whole new thing.
As Howard put it later in the show, it isn’t another application, and it doesn’t act like a human either. It needs boundaries defined up front, including things like:
A specific identity,
ashort list of things it’s allowed to reach, and i
solation that travels with it whether it’s running in AWS, on bare metal, or in an OT facility.
What this looks like if you write code
Time for me to put my work hat on for a minute. The table above mentions “an identity-driven overlay” enforced “inside the application,” which is a mouthful, so here’s what it means in practice.
OpenZiti is the open source zero trust networking platform that NetFoundry builds, and one of the things it lets you do is embed zero trust directly into your application with an SDK. Your app gets its own cryptographic identity, and it can only reach the services that policy says it’s allowed to reach. On the other side, services don’t need open inbound ports at all, so there’s nothing sitting on the network for an attacker to scan, probe, or point a stolen token at.
Galeal’s “kill switch” (more on that below) gets pretty concrete here too: if an identity is compromised, you revoke it or change the policy, and its paths go away. No firewall surgery required.
If you want to try it yourself, the docs and quickstarts are at openziti.io.
What do you tell the board?
Near the end, David read an audience question: Which metric best shows leadership that microsegmentation is working? Galeal boiled his answer down to three questions (39:11):
The graph: Can you answer the question “What identity can talk to what identity, according to what policy?” (Note that he said identity, not IP address.)
The kill switch: When something unexpected happens, where’s the kill switch that lets you deal with it without compromising uptime, human safety, or business continuity?
Centralized governance: How much of all this can you see and govern from one place, instead of going to a bunch of different environments and firewalls?
Then Howard put on his gloves (41:27): “I couldn’t disagree more.” Speaking as a CISO, he said:
“What my board said 10, 15 years ago was, ‘We better never get hacked.’… Today they say, ‘How bad will it be?’ That is their question to me. That is the question they want me to answer in every board meeting they invite me to.”
My answer would be “Take the pwning like a champ”.
David pointed out that the two answers aren’t really in conflict: Galeal was describing what you measure for yourself and your security team, and Howard was describing what you tell the board. Galeal agreed. Howard then explained why the board version has to be so compressed (43:29):
“I get one slide as a CISO… I get like five minutes… I make that one slide tell them how I’m spending money to make it less bad than the last time they gave me money.”
Ah, the dreaded “You get one, and only one, slide” directive for board meeting presentations. Replace “board” with “VP of Engineering” or “whoever approves your budget,” and it’s still excellent advice.
My takeaway
If I had to boil the whole episode down to one idea, it wouldn’t be about AI, firewalls, or attestation. It would be Howard’s “job number one”: make sure the next person can be as successful as you are.
Galeal landed in the same place from a different direction. When David asked why he started NetFoundry (24:00), he said that after years of beating his head against walls like microsegmentation, he wanted to tilt the playing field so that the next person who has to solve these problems doesn’t have to.
That’s the real case against building microsegmentation out of subnets and VLANs. Even if you get it working, you’ve built something nobody else can understand, let alone maintain.
A policy that says “the billing service can talk to the orders database, and the AI agent can talk to these three APIs and nothing else” is something the next person can read, reason about, and change without breaking everything. And when some of the things doing the talking are AI agents that nobody can fully predict, that kind of clarity is what keeps the chaos monkey in its cage.
Go watch the episode, and if you’re a developer who wants to see what identity-first networking looks like from the inside, take OpenZiti for a test drive!
(On the bright side, I had time to draw a “back of the envelope” comic that covers the big CVE!)
The CVE numbers
Here are the numbers:
1,248 new network-exploitable CVEs
117 of these CVEs are rated at 8.6 or higher
12 of these CVEs are a perfect 10.0:
9 of those twelve in vm2 alone
Worth reading
1. The Langflow flaw
CVE-2026-85025 (9.8) is unauthenticated RCE in IBM Langflow OSS 1.0.0–1.11.5. It’s reachable through publicly shared MCP project endpoints, with read and write access to chat sessions on top. Whatever context, credentials, or proprietary data flowed through those sessions comes along with it.
The mechanism is the interesting part. Langflow lets you share a flow via an MCP project endpoint so other tools and agents can call it programmatically. The code that’s supposed to enforce “this one flow is public, everything else isn’t” doesn’t hold that boundary. So “publicly shared” quietly generalizes from one flow to the host. This happens without a privilege escalation chain, user interaction, or even any waiting for someone to click anything.
All in all: 8 CVEs against Langflow in a single seven-day window, including the one above: code injection, OS command injection, path traversal, an incomplete scanner denylist, and this one.
2. The difference between CVE vs. KEV
CVE: Common Vulnerabilities and Exposures. These are publicly-disclosed security flaws, and may or may not have been used in an attack. These are weaknesses that have been announced.
Mark writes that he’d bet money that the Langflow flaw above becomes a KEV. Another Langflow CVE has already done that: CVE-2025-3248, a missing-authentication RCE, which went onto CISA’s KEV catalog on May 5, 2025, with a three-week federal remediation deadline and confirmed ransomware use. Sysdig later documented that same flaw on a server nobody had patched in over a year as the entry point for what they assessed as the first fully autonomous agentic ransomware operation.
The traits that move a CVE onto KEV are the traits this one has: no auth, trivially scannable, and sitting on infrastructure that isn’t in anyone’s asset inventory. Langflow has supplied both ends of that pipeline before.
3. Why AI agent platforms keep showing up inReachability Watch
Because these tools are three things all at once:
New
Fast-moving
Production-critical
…and they’re shipped by teams optimizing for getting an AI capability out the door. As a result, exposure and hardening get less scrutiny than they’d get on a mature enterprise system.
The part that makes them worth an attacker’s time specifically: compromising one doesn’t get you a single host, it gets you the whole key ring. The same property that makes the platform useful is what makes the blast radius large, which is why “it’s just an internal tool on a dev box” ages so badly.
The quite-full ECC was about three-quarters software developers, and Pratik had just finished taking a headcount of the Java people, the TypeScripters, the Pythonistas, and of course, the Rustaceans (“most people want to rewrite everything in Rust, and I find it extremely annoying”). That’s when he said what would’ve gotten him tarred, feathered, and run out of a Java meetup if we’d been back in the days of Java SE 18 or 19:
“One of the things that you may get from this session is that the programming language doesn’t even really matter anymore.”
Of course, his paln was to prove that statement by the end of the night, as well as the conclusion that you should draw if you still plan to remain in software: If you hand the coding off to a machine, the valuable work moves up the stack.
Pratik is a longtime Java/JVM guy, one of the organizers behind the DevNexus conference in Atlanta, and as of about a month ago, he works at Hugging Face. He also polled the room on what Hugging Face actually does, and enjoyed the answers (mine: “Hang out with Jensen!”). The correct one, which he eventually told us, was that they own Transformers and the rest of the library stack that lets you download, and better still run the models.
The gamble: Building before the presentation
It takes time for an software factory to build things, even on a late-model MacBook with 64GB RAM like Pratik’s. Sp he started the demo at the beginning and then talked over it, which I suppose is the new version of “live coding”. I do that myself (both live agentic coding as well as raw-dogging the code the old-fashioned way), so I have to respect Pratik’s boldness.
He asked the room for an app idea. Someone suggested a horse-boarding scheduler, which was too big. Charles, who works for a sports organization, suggested a playoff odds calculator. Pratik assessed that it was doable within the given time, so he agreed to build that.
Pratik opened Gemini and, off the cuff, dictated a rough spec: look at the remaining season schedule for all 30 MLB teams, calculate each team’s chance of making the playoffs, and let the user tweak the metrics on the page to try different scenarios.
Gemini spat out a 255-line spec in a few seconds. He pasted that into his software factory, told it “Create a new app called MLB Odds, here’s the spec, let me know when it’s finished and what the URL will be,” hit enter, and walked away from it to start the actual presentation.
“Again,” he said, “this may be a total disaster. It may not work, but let’s see how it goes.”
So what is a software factory?
Pratik was upfront that this is the most undefined term in the industry right now: “If you ask 10 people in this room what a software factory is, you’ll probably get 10 different answers.” Here’s his:
An AI software factory is a system, not a team, that turns human-written specifications and intent into working software, using agents to perform most or all of the coding, testing, review, and release work.
The key word is system. You say “go build this,” walk away, and come back hours or a day later to working software. It’s like handing a project to a team of engineers and going off to do something else (or hey, nothing. You’re the boss!).
And critically, a software factory is not one-shotting. Telling Claude or ChatGPT “build me this website” and pasting in a detailed description relies entirely on the raw intelligence of the model. A factory applies rigorous software engineering process around the model: planning, architecture, implementation, testing, QA, security review, version control. The process is the product.
The levels: where you are on the ladder
Pratik walked through a progression that got a lot of nodding in the room:
L0: Manual coding: Open an editor, translate what’s in your brain (or in a written spec) into code, hit a wall, go read the docs or Stack Overflow, keep going.
L1: “Spicy autocomplete”: IDEs got language servers and started completing large fragments for you. A genuine step change in velocity.
L2: AI pair programmer: You describe what you want to a chatbot, it emits code, you copy it into your editor and fix it up. You’ve effectively stopped going to Stack Overflow directly, because the model already ate it.
L3: AI as a senior developer: You give it a spec, it builds, you review the code and tell it when it picked the wrong approach.
L4: You’re the PM: You write specs, you review plans, you check back later.
L5: Software factory: specs go in, software comes out. You review the product, and you may never look at the code at all.
His take: most working developers are somewhere between L2 and L4 right now, depending on how much freedom their employer gives them.
What’s left for software engineers, then?
Pratik clearly flagged this part as his own opinion. Like a lot of developers, myself included, he went through the “there’s no way AI replaces me” phase. He concluded that it doesn’t replace software engineers, but it does dramatically change what they spend their day doing, which is the remaining high-value work. I get the feeling that programmers went through something similar when going from assembly to higher-level languages.
With AI writing the code, the process of programming becomes even higher-level, with these becoming our main activities:
System design. The model knows how to write code. It does not know that you can’t build a browser front end in Python, and it doesn’t know that your real-time app needs in-memory caching to hit a sub-three-second response. And if it improvises those decisions unprompted, you will not like the implementation.
Security.
Making sure user requirements are actually met.
Writing and maintaining the specs and the architecture.
Building the validation harness that keeps the factory honest.
Agent orchestration.
He also noted, to the product managers in the room, that the line between PM and engineer is blurring fast in both directions.
The four jobs of a factory
At minimum, Pratik argues, a software factory has to do four things:
Plan the build
Code it
Test it (looping back to implementation when tests fail)
Verify with a QA/security/architecture-compliance pass (the kind of review a human QA engineer would do by using the result as a user would, not only unit tests)
The single most important artifact in all of this is architecture.md. That file is where you, the engineer, do the system design: “This is a web app, use Vue, this is the backend, these are the performance targets, this is how we build things around here.” It’s exactly what you’d tell a new team lead. The factory can’t extrapolate it from your brain.
Specs are a separate thing from architecture: specs are the features, architecture is the system. For spec format, Pratik mostly writes Markdown, though he noted that larger teams use a more rigid PRD structure, and GitHub’s Spec Kit is out there if you want something opinionated.
Pratik’s actual stack
Here’s what he’s running, top to bottom:
Hermes Agent as the front office.Hermes (the open-source, self-hosted personal agent from Nous Research, in the same category as OpenClaw) is his always-on orchestrator. It runs on a server, it’s reachable via Telegram, Discord, Slack, or email, and it has persistent memory that turns repeated requests into reusable skills. Hermes takes his request, polishes it into a formal spec, finds the right project directory, and dispatches a headless job to the factory floor. He runs it on a 35B-A3B Qwen model, which has 35 billion parameters, but only 3 billion active per token, so it fits on a modest server.
pi.dev as the factory floor.pi.dev is Mario Zechner’s minimal terminal coding harness. It ships with four tools: read, write, edit, bash. It has a tiny system prompt, and hooks for everything else. That’s it: it’s a build-your-own coding agent, not a sealed product. Pratik picked it over Claude Code, Codex, Cursor, OpenCode, Kiro, and the rest because it’s lightweight, it doesn’t burn tokens, and you can point it at any model you want.
Someone reasonably asked why he didn’t simplyu use Hermes for the coding too, since Hermes can code. His answer was about context hygiene: he wants Hermes to be the manager and pi to be the worker, and he doesn’t want to pollute the coding agent’s context with all the management and channel-routing overhead. “It doesn’t have memory. It doesn’t have learning. I don’t want all that stuff for my worker software monkey that’s going and building the code.”
A local LLM.Qwen 3.6 27B, running on a 5090 at home rather than on his laptop for the demo. (Qwen 3.8 27B had landed a few weeks earlier and he said the jump was a substantial improvement.)
A context MCP server. This is the piece I think people will underrate. Qwen 3.6’s knowledge cutoff is roughly a year stale, which means it doesn’t know current Vue and Nuxt APIs. So he plugs pi into an MCP server loaded with current docs and code samples for whatever he’s building (HTML/CSS, Vue.js, Nuxt) so that it builds against Nuxt 4, not whatever it “half-remembers” from last year.
And there isn’t one factory, there are three: a front-end one (Vue/Nuxt), a Java/Spring Boot one for performance-sensitive backends, and a Python one for utilities. Same seven-step process in all three, different tooling underneath.
The seven steps, and where they live
Inside the project’s .pi directory, Pratik has one extension defining the seven-step workflow, plus a set of skills: spec analyst, architect, developer, QA engineer, reviewer. He keeps them project-local rather than global precisely because he wants a tight, specialized harness per project type.
He opened up the spec analyst skill live, and the reveal was how short it is:
The whole thing is barely a screen’s worth of text telling the model it’s a technical product manager who reads requirements systematically and translates them into concrete action plans with features, data models, and API contracts, followed by a handful of instructions: do requirements gathering, document assumptions, create a data model, define the API contract. That’s it. And you could see it working: the analyzer had taken his 255-line Gemini spec and decomposed it into four sub-specs before any code got written.
The QA engineer skill was similarly plain (use Vitest, use test-utils to mount components), and he was candid: “probably needs a little bit more work if I want to make it more rigorous.”
For visual QA he uses a Playwright plugin. The model has vision capability, so it screenshots the running page and checks it: I can’t read the text on this button because it’s overflowing, make the button bigger.
Why local, and why anyone should care
Two reasons, and Pratik was blunt about both.
Reason one is intellectual property. Yes, there’s a checkbox that says don’t train on my data. Do you believe it? “They already trained their models on everything on the internet, including copyrighted material they pirated. I don’t know if I trust these guys with stuff I care about.” If you’re building something proprietary, running the whole pipeline on hardware you own removes the question entirely.
Reason two is cost, and this is where the harness argument lands. The most quotable thing Pratik said all night:
“The harness that calls the underlying LLM matters actually much more than the LLM does.”
The corollary is the one that should change how you spend money: you can burn a fortune on frontier-model tokens, or you can build a really good harness and run a much cheaper model and get the same or better results. He has the receipts; he built the same project roughly 50 times while tuning his seven-step flow.
He’s not a purist about it, either. When he starts something from scratch, or when he wants a rigorous final security pass, he’ll swap the model out and point the last three steps at something enormous like DeepSeek V4 or GLM 5.3 via OpenRouter or Hugging Face inference providers. Same harness, better LLM, but only where it’s worth paying for.
The hardware detour (and the bad news)
Pratik spent a useful chunk of the talk on quantization, because it’s the thing that determines whether any of this runs on your machine.
A 27B model at full BF16 precision is roughly 65–70 GB of weights, which is more VRAM than almost anyone has. Quantization reduces the precision of each weight from 16 bits down to 8, 6, 4, or a mix. Qwen 3.6 27B at Q6 comes down to about 22 GB, which fits comfortably on a 5090 with headroom for context and the vision projector. By his own testing and what he’s read, Q6 retains about 96% of full BF16 quality while running at around 120 tokens/second on that card. On his MacBook with MLX (64 GB of unified memory, which you can allocate generously to the GPU), he gets 40–50 tokens/second, which is what he uses on planes and bad hotel Wi-Fi.
Someone asked whether ~27B is the floor for useful coding models. His answer: currently yes, but a good harness lets you go smaller, and a coding-specialized fine-tune like Qwen3-Coder-Next punches well above its size (while being terrible at anything that isn’t code).
The bad news: now is a terrible time to buy hardware for this. The 5090 that cost someone in the room $2,500 a year ago is around $5,000 now. An RTX Pro 6000 with 96 GB went from roughly $8,000 to $16,000. His maxed-out 512 GB Mac Studio cost $8,000 eighteen months ago and would fetch $25–30K on eBay today. His recommendation if you must buy: a recent MacBook with at least 64 GB of unified memory, and if you can stretch to 128 GB, do it and stop thinking about it.
Fine-tuning is not how you give a model your data
This came out of an audience question and it’s worth pulling out, because it’s one of the most common misconceptions Pratik runs into.
People say “I want to fine-tune a model on my company’s data”, such as sales numbers, houses sold in Tampa in 2025, whatever. That’s the wrong tool. Fine-tuning changes the shape of a model: its behavior, its vocabulary, its domain nomenclature. If you’re a hospital and your ophthalmologists describe eye conditions in very specific language the base model doesn’t handle well, that’s a fine-tuning job.
Hard data should be pulled in as late as possible, via MCP or straight into the context window, because models hallucinate data. Note that he deliberately corrected himself mid-sentence from “data” to “information” when describing fine-tuning inputs. That distinction is the whole point.
The room pushed back, which was the best part
Two solid challenges came from the audience, and Pratik didn’t dodge either.
In their Back to the Future of Software presentations at Devnexus and Arc of AI, Baruch Sadogursky and Leonid Igolnik argue that waterfall didn’t fail because it was inherently bad, but because the cycle time was measured in months. Agentic coding shortens that time; their thesis is that the specific failure mode of waterfall was latency, and AI has changed the latency equation.Read more here.
“Isn’t this just waterfall, which we spent 20 years learning to hate?” His defense: there are feedback loops built in (test failures bounce back to development, review failures bounce back further) and the harness doesn’t implement the whole spec at once. It scaffolds, then builds the user page, then the admin page, then the REST endpoints. Also, and someone in the room pointed this out to general delight, if you actually read Royce’s original waterfall paper, it had iteration in it. The verdict was tabled for the bar.
“Every one of those artifacts is itself a product you have to maintain.” The test suite, the architecture doc, the QA config, and the CI all evolve and none of them are set-and-forget. Pratik conceded the point. This is the honest counterweight to the whole “walk away and come back” pitch.
So did the MLB odds app work?
Sort of. Which is more honest than most demos.
The first run got killed partway through because it wasn’t doing what he asked. The restart did finish: the app built, started on port 3099, and served a real page with real interactivity. Clicking around ran an actual simulation under the hood.
The problems were exactly the ones you’d predict. Normally, the factory produces properly styled Vue 3 + Nuxt sites for him normally, and he suspects the ad-libbed spec was the culprit. And the numbers were nonsense. The Rays were given a 0.3% shot; someone noted the data looked very old. Pratik’s response: “I didn’t tell it where to go get the data from. So yeah, I was very lazy.”
That’s not a failure of the factory. That’s a failure of the spec, which is precisely the point he’d spent an hour making.
Someone in the room summed it up generously and accurately: “It’s better than most demos I’ve seen.”
Limitations, stated plainly
This factory is good for small to medium projects. Larger ones need heavier machinery (he name-checked obvious.ai in Atlanta, who sell an industrial-strength factory as a service).
The demo was greenfield. You can put a factory on top of an existing brownfield codebase, but you’ll need a discovery pass and a hand-built architecture.md first.
There is no standard. What a car repair shop needs from a software factory and what an airline needs are different things. Pratik thinks some standardization is coming, but he’d put it at least a year or two out.
Everything in this space has a shelf life measured in weeks. His own words: “What I tell you today is the right way to do it will be antiquated and the wrong way to do it a month or two from now.”
Takeaways
If you only keep five things from this one:
A software factory is process, not vibes. The difference between a factory and one-shotting an app in a chat window is that the factory wraps the model in software engineering discipline: specs, architecture, staged implementation, tests that loop back on failure, QA, security review, and a git history you can roll back. Vibe coding has none of that.
The harness matters more than the model. This is the highest-leverage idea of the night. A well-built harness plus a cheap 27B local model can match or beat a frontier model driven sloppily, at a tiny fraction of the cost, and without shipping your IP to somebody else’s training run.
architecture.md is your job and nobody else’s. The factory will happily build the wrong system beautifully. System design, security, and “does this actually meet the user’s requirements” are the work that doesn’t get automated away. Be the team lead, not the typist.
Garbage spec in, garbage app out, and the MLB demo proved it live. The parts of the app that failed were the parts Pratik never specified: the styling and the data source. If you find yourself blaming the model, check the spec first.
Start small and build your own. pi.dev plus a handful of Markdown skills is a genuinely approachable starting point; Pratik’s entire spec-analyst skill is one short paragraph plus a checklist. Point it at a paid API if you don’t have the hardware, because right now is a genuinely bad moment to buy GPUs. And keep the context MCP server in mind, because your local model’s knowledge of your framework is probably a year out of date.
Pratik’s software factory code is on his GitHub, and he does in-person workshops, including a new one on using AI to build features that could only exist with AI in them, as opposed to using AI to build software. He gave that one at KCDC last week. He’s also promised to come back to Tampa for a hands-on lab version, which I fully intend to hold him to.
Oh, and DevNexus 2027 is running ten tracks, seven of them AI. If you’re looking for a conference to go deep on this stuff, that’s the one.
Here are some of the AI articles and videos I’ve been looking at this week:
HTMX: No AI Fridays. “If the productivity gains from AI are so big, spending one day a week to minimize its downsides shouldn’t be a difficult trade-off.”
AskMike.org: What my dad taught me about AI coding in the 90s. “What is clear to everyone is that if you use AI in a way where you spend little time saying what you want, and no time reading what it coded up – you are vibecoding and the resulting software is not going to last very long (if it works at all). So how much coding should you let the AI do, and how much should you control and read (and deeply understand)?”
The Internet Archive’s “Vintage Artificial Intelligence” collection: “A curated collection of early computer software claiming some aspect of artificial intelligence as a primary feature, allowing the early interactions of humans and machines before things got a little strange. Includes a variety of programs intended as therapists, conversation partners, and opponents.”
Business Insider: Dario Amodei says Anthropic is ‘not interested in destroying anyone’. “With Salesforce CEO Marc Benioff seated beside him, Amodei sought to put to rest Wall Street’s belief that Anthropic aims to conquer all things SaaS. After all, it was an Anthropic announcement in late January about plugins for Claude Cowork, not the four horsemen, that signaled the beginning of what became known as the Saaspocalypse. In one week, roughly $1 trillion in value was wiped out in the software sector.”