
Last Thursday, I was at the Machine Learning / AI / Technology / Risk + Happy Hour Prototype v0 (they’ve got to work on that name) organized by ISACA’s West Florida Chapter. It’s going to the be the first of a monthly series of get-togethers, and the first one centered around an AI panel witrh the following panelists:
From left to right, as seen from the audience
- Mikael Loefstrand — focused on data protection and data security for AI workloads
- James Gress — Accenture, 26 years this month. Splits his time between showing people how to be more productive with AI and building AI solutions for complex business problems, governance and monitoring included
- Jim Lovett — St. Pete born and raised, Gator, 26-ish years as an Oracle DBA, now an AI Ops engineer at Athena Health. Introduced himself by noting he has all of your medical records and won’t let anybody see them
- Alexei Chirokikh — PhD in physics, twenty years at IBM on Wall Street, then CIO roles in Moscow, Nairobi and across 20 countries for a global microfinance organization, then Citibank. Now CTO at a small, aggressively growing community bank
- Daniel Jarboe — five years in cyber risk management, currently wearing both the CISO hat and the business-enablement hat at an ~80-person publicly traded property and casualty insurer
Here are my notes from the panel. You might want to read to the end, because there’s a great quote…
Contents
- The room was more regulated than you might expect
- From PowerPoint to AI giving you what you asked for, but not what you wanted
- Never mind the human in the loop; where’s the leader in the loop?
- The five-minute agent, and the shadow IT it leaves behind
- Saying “no” costs political capital
- “One plus one? It depends.”
- What they’d actually let AI touch
- AI fatigue
- The value question nobody has answered
- The takeaways
- Truer final words were never spoken
The room was more regulated than you might expect
Before taking questions, the panel asked for a roll call of what industries audience members worked in. The answers shaped the discussion that followed, and they were:
- Banking and finance
- Six people in healthcare
- Somebody from Experian
- Legal
- Higher ed
- OT/SCADA and maritime security
- Healthcare
- Manufacturing
- DHS
- A former deputy for cyber governance at a combatant command
- Air Force
Nelson pointed in my direction and said “Joey, you were military,” to which I replied “No, cybersecurity!” but I then turned around, and right behind me was the Joey he was referring to: Joey Hernandez.
Clearly, this wasn’t a room of startup founders who can YOLO deploy on a whim on a Friday. Most of us did work that would be audited by one or more regulatory agencies. When a room like that says something is now possible, it’s usually possible for everyone.
From PowerPoint to AI giving you what you asked for, but not what you wanted
Before the panel started, I had a quick chat with Jim Lovett, who asked me “What are you doing now with AI that you couldn’t do six months ago?”
For me, it was a combination of more complex coding, and having it look at my “back of the envelope” cartoons and asking if I’d missed or misspelled anything (Claude is quite good at interpreting my drawings these days).
Jim posed the same question to the audience. Someone pointed out that if you have an idea you want to show someone, you can now ask a model to produce it as a single interactive HTML file in minutes, with all the data living in that one file, complete with user interface and interactivity. They’d used it in various scenarios, including raising money, and pitching an idea internally, within their organization. Pre-AI, they’d need a week, and the result would be something more static: a PowerPoint presentation, a spreadsheet, or the dreaded Word doc.
An audience member pushed back, and the pushback turned into the longest thread of the discussion. Their complaint: Sure, AI generates the text and builds the presentation, but the output still isn’t ready for public consumption. Every single slide needs tweaking, because they like things a certain way. “You can’t press a button and be done. You always have to intervene.”
(I myself have remained fully in charge of writing and designing my own slides. In my case, it’s professional pride.)
Every panelist had a different answer, and collectively they amount to a decent field guide:
Set up a persona. This was Alexei’s suggestion. He runs multiple concurrent engagements — CTO on one, chief architect on another, infrastructure on a third — and maintains a separate persona per engagement: interests, tone, language. The model learns the shape of what you want instead of relearning it every session.
Stop asking for the artifact you don’t actually want. Mikael answer was more blunt. He said he hates PowerPoint, and the AI tooling around PowerPoint was especially bad six months ago, so he stopped. He now builds single-file HTML presentations and passes them around as files that open in a browser. He told his boss, roughly, that PowerPoint sucks and he’d deliver something stunning as a single HTML file, and if PowerPoint was truly required, someone else could execute it.
Calibrate what “automation” means. One of the panel (I forgot to write down who said this, which is a shame, because this answer most closely aligns with my thinking) made the point that a lot of people expect full automation and that’s not what this is. It’s an intern. You say “go make a deck that does X”, you read it, you roll your eyes at their dumbass rookie mistakes, point them out, and say “Here’s how to fix them.”
Grill yourself before you generate. Alexei described the grill-me skill, which you drop into your agent’s workspace, and it interrogates you with thirty or forty questions about your content and structure before it attempts to write anything. It supposedly results in fewer bad outputs because it has fewer under-specified inputs.
One caveat I’d add, since nobody on the panel did: he described skills as “just text” with no security vulnerability, which is true of a plain-Markdown instruction file and not true of skills that ship executable scripts or pull in remote content. “It’s just text” is exactly the assumption prompt injection is built to exploit. Don’t trust blindly; read what you’re dropping in your workspace!
Iterate once, then capture the recipe. James’s version is the one I’ve started using: don’t over-engineer the initial prompt. Run one thread, push it until you like the output, then ask the model to write you the prompt that would reproduce this style. Start there next time.
Expectations are the actual problem. Daniel reframed the whole complaint as misaligned expectations. You expected a perfect deck from one prompt with no edits, which realistically isn’t happening, even though plenty of people out there claim it is. He was pointed about the hype: a lot of people are being negatively affected by exaggerated claims about what AI can deliver. And by the audience member’s own admission, she’s still saving an enormous amount of time. That’s the win. Take it.
An audience member — a recent government departee now consulting with small businesses — offered the framing I liked best. They tell clients it’s a relationship. It sounds ridiculous, because it’s a computer, but it’s a back-and-forth made up of conversations. If you have a slide you like, upload it and say this is how I want all my slides built from now on. And on personas: there’s him as a dad, him as a husband, him as a consultant, him as a military guy, him as a Dallas Cowboys fan. Those are the voices he wants to respond in, depending on who he’s talking to.
Never mind the human in the loop; where’s the leader in the loop?
Jim spent the evening running the most opinionated workflow in the room, and this was his thesis:
“The hell is a human in the loop? I want a leader in the loop!”
A human in the loop, he says, is a rubber stamp. A leader in the loop is someone capable of constructive criticism who brings taste, intent and accountability, and who states all three explicitly, because the model can’t infer them.
He backs this with an absurd amount of written specification. He described building two- and three-thousand-line Markdown files so that anyone could drop one in and get output shaped by that person’s intent, taste and accountability. His own standard output spec is a full page of hierarchy, and he reads back to the model exactly which rule it broke. It’s all plain text, and he makes it a point not to sound authoritative in a register he wouldn’t naturally use.
He told a story to justify all that ceremony. At a previous job, a single spreadsheet with fifteen tabs fed fifteen unique PowerPoints for twelve to fourteen executive directors, each with their own taste, intent and accountability. Fifteen people were employed to maintain it. He was told there was no way AI could ever do this. He asked for a sample; they said it had real data in it; he said he couldn’t help them then. Eventually someone sat with him and screen-shared, and he built it.
His conclusion, which got the biggest laugh of the night was not to fire the fifteen people, but fire the fifteen recipients. “Why do we mandate a spreadsheet that mandates PowerPoints? We’re using that tool because of precedent, and precedent alone.”
He also offered a method, delivered as a sixth-grade algebra story: nobody told him the solutions were in the back of the book. Ever since, he starts at the back and works forward. Don’t start at the front with the newest tool, the newest loop count, the newest memory feature. Start from the answer you need and work backward. If you have something good, let the AI rip it apart.
Then he demonstrated live. He picked up his phone, told Gemini he was on a panel with friends, described the question that had just been asked, and requested something precise, concise and not overly verbose in the next sixty seconds. His aside while doing it was sharper than the demo: talking to inanimate objects in front of other people is uncomfortable for the other people, even when it isn’t uncomfortable for you.
An audience member asked how many times you should loop before you share something. Jim’s answer was so good that I’m simply going to quote him from my recording:
“You need to beat it up 16 different ways to Sunday until your eye agrees with what comes back. And it’s layman’s terms and I can hand it to someone. Not only does it make sense, it provides a solution, and at the end of the day, damn it, it better bring value for all the shit we’re paying for, because — Lord have mercy — it’s expensive, and people want to know what we’re doing with it, and it’s the same people that never asked us what the hell we were doing with Excel?
(Jim is at his best when he’s slightly unhinged. We need Jim on more panels.)
The five-minute agent, and the shadow IT it leaves behind
Quite unsurprisingly, James provided the single most concrete answer to the “What are you doing with AI that you couldn’t do six months ago” question.
He’s been building a RAG system to answer common new-joiner questions used to mean a team of five people over months, plus ongoing maintenance. He recently encountered a team that wanted a graph database, a client AWS environment, the whole apparatus. He stopped them: the client has SharePoint and Teams. He went into M365, created an agent scoped to exactly one SharePoint site, pointed it at new-joiner information, and it worked. Five minutes. He was done before the call ended.
Two years ago, he says, it would’ve taken five people, six months, plus a maintenance burden forever.
Daniel took that exact story and turned it into the risk thesis of the evening. The entire premise of AI is speed. You’re enabling everyone to be more productive. So you can predict what happens next: everybody generates applications and use cases, and you end up with giant shadow IT. When those people leave, so does a lot of institutional knowledge. Nobody knows how those applications were built, or how to maintain them, and meanwhile the business is running on them.
Alexei’s counter was that this is the same problem we’ve always had: lack of documentation, and lack of discipline in how you build things. The difference is that pre-AI, all this happened with Excel.
Daniel rebutted with: That’s legacy, and legacy is different. Before, maybe you had one application, built by a team over ten years, where the company’s still around and so is that application, but the developers are gone. The problem now is that you have a hundred and fifty similar applications supporting the business and nobody knows anything about any of them. It’s a problem of scale.
Jim, whose job is to answer when those panicked pages come in, agreed. He added that he’s not primarily worried about security risk; he called that “his hobby”. He’s worried about operational risk. He has to support everyone’s five-minute widget. A number of them will eventually break, each break is going to end up as a ticket, and the growth of unsupportable band-aids is already getting unmanagemable.
You can call the person who created the agent, but once it goes GA and becomes part of day-to-day operations, there’s nobody to call. His organization has three teams building agents that push code autonomously, and they’ve deferred going live for two months specifically because they don’t believe they can support it operationally.
James added the enterprise version of the same caution: Accenture restricts roughly half the functionality of frontier tools, and agentic desktop tooling with broad local filesystem control is largely off-limits across the organization.
Daniel suggested this for small companies: You still need some form of center of excellence, even if it’s a part-time role for one person. Figure out your operational controls. Decide what’s allowed. His wife’s 20-person company does exactly that: a short list of permitted tools and clear guidance to staff.
In the end, there is no single “right” answer. Different organizations have different risk tolerances. Smaller companies with less to lose are going to go ahead and see what happens.
Saying “no” costs political capital
The discussion moved to a favorite ISACA topic: governance.
Somebody described arriving at an organization with a six-month backlog for building applications and no intake process at all. People asked, and the team built. The fix was unglamorous: What do you need this for? What are you using right now? Someone has to own that, which slows business down, which is the cost. If a spreadsheet does the job, use the spreadsheet.
This led to the question underneath all governance work: What is the definition of “no”, and how do you squash the squeakiest wheel?
Someone else made an excellent point: Shiny is the enemy. AI is extremely shiny compared to the legacy systems and capabilities you’re trying to build on. If you’re not modernizing infrastructure, are you actually going to be able to use AI the way you think? You’ve got a fifteen-year-old server at the back of the data center that would take the whole business down if you turned it off, and you want to onboard 20,000 GPUs.
Jim’s response was that this happens every single day at every organization, and that he’s watched three decades of it, where people ignore the technical debt on the cash cow and throw money at the shiny new thing because that might drive double-digit revenue growth. In an ideal world, the right thing to do would be to say “no”, but Jim spelled out the problem with that approach:
“Saying ‘no’ costs an incredible amount of political capital.”
He described arriving at work every morning with a truckload of it and dumping it on the step, because that’s what the job takes.
This was followed by a discussion about tool sprawl. Jim wants a single tool, and his argument isn’t technical. It’s about human friction and his own operational sanity. Models update faster that we can evaluate them. He’d rather pick one and make it work well inside the institution than maintain opinions about five.
An audience member suggested “Right tool, right time, right people.” A receptionist and a finance person do not need the same toolset. The receptionist may not need AI at all.
My take: You can declare a single tool all you like, and then the receptionist uses ChatGPT to answer customer questions anyway, and nobody’s taking inventory of how much data has been shoved into the cloud.
Jim observed that at his work, the resistance isn’t at the bottom. It’s at the top. Executives think Copilot is fine because they live in email and meetings.
“One plus one? It depends.”
Jim stopped and asked for a show of hands: “Does everyone know what ‘non-deterministic’ means?”
We all fell in love with this technology because it’s reasonably human-like. It’s human-like because it randomly samples at the end. It computes top-p and top-k and then picks. Ask it one plus one and the honest answer about the mechanism is “it depends”. Which led him to ask, a couple of years ago, how anyone could possibly deliver ambulatory AI, because it will never give you the same correct answer twice.
The best rebuttal came from the panel itself: Two doctors (or experts in any field) will never give you the exact same answer either, and they’ll sometimes differ by quite a bit.
Alexei added to this, saying that what we’re seeing today is a genuine simulation of the world we live in. The classic computer science maxim “garbage in, garbage out” applies. It means that if your processes are biased, especially in credit risk and financial services, the output carries the same bias. The world is also full of contradictions, and those contradictions live in the output.
Mikael offered the engineering response: you can constrain this. Skills, directives, instructions, templates. You can make the piece you care about deterministic, even if the prose around it varies. You can’t do this for every use case, but you can do so more than people assume.
Daniel pointed out that if you talk to an actual subject matter expert, they can usually tell you whether the output is accurate. The problem is that we’ve branched into areas where we’re not experts, and then we assume nobody can correct it. That’s not true.
What they’d actually let AI touch
Jim was the most specific person in the room about data, presumably because he’s the one holding the medical records.
His org doesn’t let AI touch production unmasked data at all. Lower environments only, everything masked before it leaves the walls, masking hundreds of millions of records weekly with Delphix so the lower environments carry no PII or PHI.
When they needed to clean PHI out of their source repositories, they evaluated multiple vendor tools and none came close to what a frontier model did from the command line. He said the model did a noticeably better job than any vendor could offer, and the vendors simply hadn’t reached that level of proficiency.
I agree with his take on vendor trust: zero trust! One vendor says they’re not looking at any clinical data. He can’t verify it. Once it’s outside his network, it’s gone. His advice: at minimum, have a legal agreement with whoever you choose.
He also took a detour into what operations looks like at scale now:
- Kafka clusters and OpenSearch on-prem
- 6 figures of pods
- Tens of thousands of Kubernetes namespaces, each carrying their own ingress and egress rules
- Hybrid nodes running the same thing on-prem and in cloud because management decided a decade ago to run two systems instead of one
- Data stretched across two platforms, usually three.
Jim’s summary of his own job: AI Ops has a lot to do with data integrity, performed by a machine with no integrity.
There was a JPMorgan story in there too, about spinning up something like a hundred thousand EC2 instances to crush a workload ahead of a market event, being asked to halve it on cost grounds with no technical reasoning, and demonstrating that you could pull information out of memory on shared servers. His note on that last part: with today’s tooling, it’s easier.
AI fatigue
Jim described the shape of his days now: “You’re always on.”
You’re either heavily reviewing something, heavily creating something, or heavily doing idea generation. The toil that used to give your brain a break between hard things is gone. He runs fifteen agent windows, moves between them, waits, moves again. As a developer he used to enjoy the two or three minutes of compile time to think about something. He never gets to wait for the compile anymore.
An audience member asked the right question: What’s your error rate? Because you cannot tell me your errors don’t go up while you’re managing fifteen things.
Alexei answered: He feels more productive. He isn’t focused anymore. Productivity used to be linked to focus; now everything is concurrent, and he thinks that’s dangerous. He can do in three hours what used to take a week, and he can no longer concentrate on one thing, which was his actual answer to the panel’s opening question about what changed in six months.
He also gave the most specific failure mode of the panel: In six months he’s given a task to the wrong agent maybe three times. Before AI, this never happened, but in the his current environment, there are multiple desktops and so much going on. It’s all to easy to copy in one window and then paste into the wrong place.
He then brought up what I think is the real risk: information that’s inaccurate but looks shiny. It’s beautiful, makes sense, and makes you feel good (that’s what it was built to do), but it may also be completely wrong.
Two mitigations came out of that exchange, both worth stealing:
- Instruct the model, up front, to tell you when it doesn’t know, and to not exercise judgment it wasn’t authorized to exercise. Treat it as a model, not a decision maker.
- Build a red team agent. The consultant in the audience does this for every small-business application build: an adversarial agent that audits everything from the requirements forward. It should come back with things like “This button doesn’t make sense”, “This color scheme won’t render”, or “You’re rounding wrong”. He does it precisely because he’s multitasking and doesn’t want to go back and check.
Daniel supplied this general rule: Responsible use of AI takes discipline. Ask what the cost of being wrong is, and how well you’d be able to detect an error. If the cost is high and your detection is poor, you don’t delegate that to an agent.
The value question nobody has answered
Running underneath everything was money. (Surpise, surprise.)
Jim talked about his frustration: Now that AI has a line item, people demand proof of value. These are the same people who never once asked anyone to justify the value of an Excel spreadsheet. The old term was “Fail fast!”, and nobody blinked at a million dollars a quarter in experiments when the mandate came from the board.
His proposed number, offered as opinion and immediately acknowledged as a big one nobody wants to touch: If you believe someone’s productivity will increase value for your firm, you should be willing to increase payroll by 12% (I assume he’s thinking one-eighth, rounded down to the nearest percent).
James countered, drawn from his own earlier history as a developer. Given a week to do something, he’d spend three and a half days playing with the code because he enjoyed it, then realize he had a day and a half left and wrap it up. Give an engineer five days with AI, and they’ll still take five days. They’ll just spend the extra time trying to get an agent to one-shot what previously took five shots. The time savings don’t automatically become business value, because engineers like playing with technology. Why wouldn’t they?
Jim brought in accountability. Developers could barely be held accountable for their output before AI. Now make them do ten times the work, and they’re still accountable for every PR. He sees real hesitation there; people don’t want to take on anything they’re not directly assigned, and story point allocations per two-week sprint haven’t moved at all.
The question was posed: If a project is a failure, is it really a failure, or is it learning? The counter: If I spend two million dollars on AI this year, do I mostly find out which employees won’t learn it?
We can make a bad process look very shiny and very efficient. We need to ask:
- What’s the optimal process?
- What’s the change management?
Sometimes the honest answer is that you had a bad process to begin with, and AI just made it beautiful.
The takeaways
Congrats if you made it this far into this writeup! You’re ahead of a lot of people these days, what with their goldfish-like attention spans.
Here are my takeaways from this panel discussion:
Leaders in the loop, not humans in the loop. Reviewing without taste, intent and accountability is just a rubber stamp with extra steps. All three have to be written down, because the model can’t infer them.
Write the spec, then reuse it. 3,000-line Markdown standards sound insane until you see the fifteen-PowerPoint story. Cheaper version: iterate once, then ask the model to write the prompt that reproduces the result.
Cost of being wrong × ability to detect the error. The cleanest delegation heuristic I’ve heard for agents.
The scale problem is real and nobody’s solved it. Everything we knew about undocumented systems is still true. There are just going to be waaaay more of them, built in five minutes each, by people who then leave.
“Beautiful and wrong” is the failure mode to fear. (It’s also the dating mode to fear) Not refusals. Not obvious hallucinations. Polished, confident, well-formatted output that feels good and isn’t right.
AI fatigue is real. It’s being reported by the people getting the most out of these tools. A panelist mentioned his boss can now see he’s working 82 hours a week, that he stopped about a month ago, and that he took up golf and feels like he’s in purgatory because he isn’t allowed to work. The room laughed. It wasn’t entirely a joke.
Truer final words were never spoken
The last words on the recording, delivered deadpan after seventy minutes of this:
“I need alcohol. Thank you for listening to this panel.”

