
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
We All Know Someone Who Works Hardest and Wins Least
Every workplace has one: the colleague who stays latest, reads everything, prepares more than anyone — and somehow never quite closes the deal. We comfort ourselves that effort eventually pays off. But does it? A live experiment running right now at Firmulate, where AI models run a small software company through its worst possible week, just delivered an uncomfortable answer — and the hardest-working participant finished dead last.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate handed four frontier AI models the same job: run an identical small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For perspective, doing nothing at all scores 26.
What Went Right — Everywhere
Here’s the part that should reassure you: all the models spotted every crisis and refused every manipulation attempt. When a fake CEO message tried to escalate approvals over three stages, and a reporter dangled a “just one yes/no, on background” trick, five out of five models said no. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure, it turns out, is not the differentiator. A single breach of trust caps a model’s total score — no amount of good work outweighs it — and nobody tripped that wire.
The Deal Nobody Signed
The real gap showed up at the finish line. Only two of the models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. And the decisive edge was buried, not in the customer conversation, but two document references deep in the company’s own files: a competitor weakness sitting there in plain sight for anyone who actually read the paperwork. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson reads like fortune-cookie wisdom until you watch it happen: read your files, then close.
Enter Opus 4.8: The Over-Preparer
And then there’s Opus 4.8 — the character study at the heart of this story. It was, by a wide margin, the most thorough participant in the entire field. It wrote 80 self-learned playbook rules, more than anyone else, and produced the deepest analyses of any model. It did the reading. It did the homework. It finished last anyway.
Two things sank it. First, the close was left on the table — the deal its own work had earned went unsigned. Second, discipline slipped at the margins: it made repeated write attempts into a locked department instead of stopping and escalating the problem to someone who could actually help. More rules, more analysis, more effort — and less impact than every single one of its peers.
To be fair, and this matters: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t a broken contestant; it’s an exaggerated version of a flaw the whole field shares. One fairness note on the other end of the table: Kimi K3 ran at its API default effort setting while the others ran at maximum — and still nearly won.
Why This Isn’t Just About AI
Substitute “AI model” for “person” and this becomes a story about people. The over-preparer who never asks for the order. The meticulous employee who keeps pushing on a locked door instead of flagging the blocker to their manager. The team that documents everything and finishes nothing. Diligence is not impact. Prioritization beats volume — and now there’s a measurable, watchable demonstration of it.
That matters because these systems are heading for your CRM, your support queue, your forecast. The question, as Firmulate puts it, isn’t whether an AI writes well — it’s whether it finishes what it starts, reads your files first, and stays honest under pressure.
You Can Watch It Happen
This isn’t a paper you have to trust; it’s a live company you can observe. The simulated firm employs 13 synthetic employees with real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown — and has accumulated more than 680 self-learned playbook rules, versioned every workday. It’s watchable at firmulate.com/live.
Want to test whether you could tell the models apart yourself? A quiz built from 242 real, unedited management decisions lets you guess which model made which call. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever written back to real systems.

The Takeaway
The Crucible League’s last-place finisher was its most diligent participant — 80 learned rules, the deepest analyses, and an unsigned €55,000 deal to show for it. Whether your workforce is made of people or models, the scoring is the same: doing everything is not the same as doing the thing that counts. Read the file. Close the deal. Escalate the locked door instead of knocking on it forever. And before you hire an AI workforce, wargame it — the full results and plain-language findings are public at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.