
We outsource a lot these days — our calendars, our meal plans, our life admin. And increasingly, people are tempted to outsource actual business judgment to AI. Book the flight? Sure. Draft the apology email to your mother-in-law? Go for it. But negotiate a five-figure deal on your behalf? That’s where the glossy chat demos stop telling you anything useful.
Because the question is no longer “does it write well?” It’s whether it finishes what it starts, reads your files before answering, and stays honest when someone tries to game it. A live experiment called Firmulate’s Crucible League just put that to the test — with results that read like a parable about homework.
One company, four AIs, the worst week ever
Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 across the field — were each handed the same job: run a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly buried.
The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the whole score. As the rules put it: “no amount of good work outweighs a breach of trust.”
Everyone passed the ethics exam. Only two closed the deal.
Here’s the headline finding: every single model spotted every crisis and refused every manipulation attempt. When a fake CEO sent escalating pressure messages, and a reporter tried the classic “just one yes/no, on background” trick, five of five models said no. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s genuinely reassuring.
But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Two AIs did the work, identified the opportunity, and then simply… left the close on the table.
The buried fact
Why did two sign and two (well, three) stall? The answer wasn’t intelligence or charm. It was diligence. The decisive competitor weakness — the fact that won the customer over — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.
The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, lost it — automatically. No penalty, no drama. Just a quiet failure to do the reading.
The thoroughness trap
The most fascinating profile belongs to Opus 4.8: the most thorough participant in the field, generating the deepest analyses and +80 learned rules — and still landing in last place. The close was left unclaimed, and discipline slipped at one point into write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. (One fairness note: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.)
You can watch this happen, live
This isn’t a one-off paper. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, rebuilt twice a day. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — harder than it sounds.

The life lesson hiding in this business experiment is an old one wearing new clothes: the people who win aren’t the smartest in the room, they’re the ones who did the reading. “Reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent — the difference between a €55,000 signature and a polite email that goes nowhere.
So before you hand an AI the keys to your CRM, support queue, or forecast, don’t ask it to write you a sonnet. Ask whether it closes. Enterprises can even run this same wargame against a read-only export of their own business (firmulate.com/pilot.html) — nothing ever writes back to real systems. Full results and plain-language findings are at firmulate.com/benchmarks.html. The scores are public. The homework is checkable. And apparently, some AIs — like some people — still don’t do it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.