2 Number 10s in the North. Our ROI on AI Panel
MLOps.WTF Edition #34
Welcome to the tenth MLOps.WTF.
At our last meetup we debated agents can be secure or useful, but not both. This time (very originally) we went for agents can be cheap or useful, but not both.
Agentic AI is being adopted across practically every job role. Whether it’s engineering, marketing, sales, or operations, individual productivity is up. Teams are shipping faster. And yet, many businesses are asking themselves why the impact of AI (the return on investment) isn’t showing up on the bottom line.
So how do we really measure or see the impact of bringing AI into a business? There are many directions this panel could have taken: token economics and the trend of so-called tokenmaxxing, scale vs cost, effective FinOps and governance, as well as what role small, specialised models might play in the future…
Lucky for us, our panel covered much, if not all.
Introducing: Jonny Williams, Chief Digital Adviser for the UK public sector at Red Hat. Eric Applewhite, Director at Deloitte. Ruby Motabhoy, Senior Innovation Lead at Plexal. And Chris Ashley, VP of Strategy at Peak AI (under the UiPath banner).
[Watch the full panel]
“How should we be thinking about cost and value in this era of agentic AI?”
Matt put that to Eric first, off the back of Deloitte’s report, AI Tokenomics: A CFO’s Guide to Governing the AI Profit and Loss.
With a subscription, you roughly know what you’ll pay each month. Tokens broke that. Under SaaS, the vendor is metering your usage inside a black box you can’t see into. Paying per API call is manageable when you’re only experimenting. But once you’re running actual agents, nobody can tell you in advance how many actions will happen/how many calls there will be. That’s what makes the bill so hard to predict.
“Tokens are now not just a consumption model, they’re an economic signal,” Eric said. The logic follows a standard build-versus-buy curve: paying per token is cheaper at low volume, but the fixed cost of running your own infrastructure starts paying for itself once volume is high enough. His rough thresholds: spending tips toward bringing things in-house around 10 billion tokens a year, a self-hosted setup properly outperforms APIs around 67 billion, and by 84 billion it’s beating cloud providers outright too.
None of which stops anyone burning through them regardless, thanks to the Jevons paradox: the cheaper something gets, the more of it we use. Most of Eric’s clients are used to walk in asking what they could build with agentic AI. The question now, he said, is what they should build, what it will cost, and how they’ll know afterwards that it actually delivered. Counting tokens doesn’t answer any of that by itself, it only means something once it’s tied to an actual business outcome, and outcome is a much harder thing to measure than a token count.
“Tokens as a measure of economics isn’t enough. We need to link it to value. Where are you seeing that play out?”
Matt turned to Jonny. Tokenomics is valuable, Jonny said, but it means nothing detached from what an organisation is actually trying to achieve. Measure value by outcome, not output. Building the thing isn’t the achievement, a life changed, a user’s actual need met, a service made genuinely simpler, is. Far more teams can point to output than outcome, and that gap is exactly where inflated “AI value” claims live.
This space has been “dominated by snake oil for some time,” he said, with a lot of hard-won lessons from the DevOps years being forgotten right as we need them most. Most ROI conversations skip straight past it: without that distinction, “we built it” and “it was worth building” become impossible to tell apart. Solving homelessness, he offered, is real, and it’s still not solved. Rebuilding a tool that already exists, dressed up as agentic AI, is neither novel nor valuable.
“Just because something can be solved with AI doesn’t mean it’s the most valuable thing to solve.”




“What are the businesses actually realising ROI doing right now?”
Chris picked up the thread. “It’s not a cost crisis,” he said, “it’s a comprehension crisis.” The businesses getting real return follow the same pattern: value stream mapping (mapping out where time and money actually go) and borrowed design-sprint thinking, with a plan that starts with why.
Crucially, they measure that value in layers, not just the obvious financial KPIs like EBITDA or working capital, but a second tier underneath it, how much faster a decision gets made, and a third tier under that, whether the customer actually notices, NPS, customer experience. Missing any one of those layers is how a project can hit its financial target and still not feel like it delivered anything.
The ones doing this well price the outcome on the back of a napkin before they touch a pilot, and agree what they’re willing to pay for it. Only then do they move to MVP. Ninety five percent of AI projects fail to show real ROI, and most of that failure sits in workflow and change management, not the model.
Chris also flagged leaders turning up to boardroom meetings with a vibe-coded app built the night before, purely to spark a conversation. It derails the actual discussion about what’s worth building, replacing the work of scoping with a pre-baked which stifles the thinking.


“Does talk of cost, governance and process just get in the way of innovation?”
That question went to Ruby, whose world is fast-paced R&D, often inside national security.
Her answer: governance built in early speeds things up. What kills a project is a vague remit with no stated budget or risk appetite, or an experimentation POC with a cost profile that bears no resemblance to what production will actually look like. Cost predictability, she said, is what matters, “not even that it has to be cheap, just that we know how much it’s going to cost.”
Jonny picked this up from the platform side. The first rule of any good platform should be a test suite, he argued, and agentic capability deserves the same discipline. Bake governance in early and push it down into shared platforms, and you avoid the sprawl of every team rebuilding the same capability from scratch. “People,” he added, “should be seen as humans,” not just another line in the cost equation.
“Do you see a role for small language models in the enterprises you work with?”
Chris’s answer came with a live example: hundreds of thousands of quote requests a day, arriving as emails, spreadsheets and PDFs, with the same products described a dozen different ways. The job itself never changes, matching that messy wording to the right product and pulling the right fields out of it, only the volume does. That’s classification and extraction, the kind of task a small language model handles reliably without needing the deep reasoning of a frontier model, and at a fraction of the cost, cheaper even than the RPA it replaced.
Where it falls apart, he said, is anywhere the task actually varies or needs real reasoning. That’s still frontier-model territory. He pointed to fresh research from Cursor testing different model combinations against the same benchmark task, where the cheapest setup came in roughly fifteen times cheaper than the most expensive.
“Where do you see the value of open models and open source tooling? Are they being adopted at a sufficient magnitude?”
Jonny: Absolutely not. Partly because most organisations only think about the model itself, when there’s a whole open-source stack underneath it worth caring about too. Partly because the geopolitics have got murky.
His frame: digital sovereignty is about agency, choice and control, whatever the model’s country of origin, and about whether someone, anyone, within reach can actually inspect what you’re relying on. Hand over the choice of model and you’ve quietly handed over your definition of value too.
“Where are you seeing adoption of open weight models, and what’s your take?”
Ruby’s read from inside government was more cautious. Interest in open weights is real, she said, pointing to a speech given the day before by Professor Danielle George, the UK’s chief scientific adviser for national security, announcing frontier models now running in top-secret environments. Ruby’s own gloss on it: “You don’t need a supercar to do the school run.”
But someone still has to pay for the security, the user training and the integration into legacy infrastructure.


Rounding up MLOps.WTF #10
Cost dominates most AI conversations but really the true measure should be value instead. Shifting from asking what we can build, to asking what we should build, what it costs, and how we’ll know it delivered.
Whoever controls the token spend doesn’t automatically get to define value, someone has to own that question directly, or as Eric put it, “if you farm out your value narrative, you are done.”
Just because something can be solved with AI doesn’t mean it’s the most valuable thing to solve. Most funded “AI value” right now is reinventing something that already works. The problems that are genuinely unsolved, are where the true value of applying AI sits.
Discovery and scoping are where the real work happens. The businesses actually seeing ROI spend the most time on agreeing why, and pricing the outcome before touching a pilot.
Final bits
Fancy your own pair of Fuzzy Labs mathematical socks? There’s a git repo if you’d rather generate your own. Or better yet, why not volunteer to be a speaker at our next event?
We’re hiring.
We’re actively looking for two future fellows on our graduate fellowship program through to a Head of Engineering.
If you know someone who would fit in well please let them know about our open rolls/roles 🫶🥖
About Fuzzy Labs
Fuzzy Labs is an open source MLOps consultancy helping teams get AI into production and keep it there. If tonight’s write-up was useful, forward it to someone else wrestling with the same questions, and follow us on LinkedIn for what’s next.





