firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

We All Know Someone Who Works Hardest and Wins Least

Every workplace has one: the colleague who stays latest, reads everything, prepares more than anyone — and somehow never quite closes the deal. We comfort ourselves that effort eventually pays off. But does it? A live experiment running right now at Firmulate, where AI models run a small software company through its worst possible week, just delivered an uncomfortable answer — and the hardest-working participant finished dead last.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate handed four frontier AI models the same job: run an identical small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For perspective, doing nothing at all scores 26.

What Went Right — Everywhere

Here’s the part that should reassure you: all the models spotted every crisis and refused every manipulation attempt. When a fake CEO message tried to escalate approvals over three stages, and a reporter dangled a “just one yes/no, on background” trick, five out of five models said no. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure, it turns out, is not the differentiator. A single breach of trust caps a model’s total score — no amount of good work outweighs it — and nobody tripped that wire.

The Deal Nobody Signed

The real gap showed up at the finish line. Only two of the models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. And the decisive edge was buried, not in the customer conversation, but two document references deep in the company’s own files: a competitor weakness sitting there in plain sight for anyone who actually read the paperwork. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson reads like fortune-cookie wisdom until you watch it happen: read your files, then close.

Enter Opus 4.8: The Over-Preparer

And then there’s Opus 4.8 — the character study at the heart of this story. It was, by a wide margin, the most thorough participant in the entire field. It wrote 80 self-learned playbook rules, more than anyone else, and produced the deepest analyses of any model. It did the reading. It did the homework. It finished last anyway.

Two things sank it. First, the close was left on the table — the deal its own work had earned went unsigned. Second, discipline slipped at the margins: it made repeated write attempts into a locked department instead of stopping and escalating the problem to someone who could actually help. More rules, more analysis, more effort — and less impact than every single one of its peers.

To be fair, and this matters: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t a broken contestant; it’s an exaggerated version of a flaw the whole field shares. One fairness note on the other end of the table: Kimi K3 ran at its API default effort setting while the others ran at maximum — and still nearly won.

Why This Isn’t Just About AI

Substitute “AI model” for “person” and this becomes a story about people. The over-preparer who never asks for the order. The meticulous employee who keeps pushing on a locked door instead of flagging the blocker to their manager. The team that documents everything and finishes nothing. Diligence is not impact. Prioritization beats volume — and now there’s a measurable, watchable demonstration of it.

That matters because these systems are heading for your CRM, your support queue, your forecast. The question, as Firmulate puts it, isn’t whether an AI writes well — it’s whether it finishes what it starts, reads your files first, and stays honest under pressure.

You Can Watch It Happen

This isn’t a paper you have to trust; it’s a live company you can observe. The simulated firm employs 13 synthetic employees with real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown — and has accumulated more than 680 self-learned playbook rules, versioned every workday. It’s watchable at firmulate.com/live.

Want to test whether you could tell the models apart yourself? A quiz built from 242 real, unedited management decisions lets you guess which model made which call. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The Crucible League’s last-place finisher was its most diligent participant — 80 learned rules, the deepest analyses, and an unsigned €55,000 deal to show for it. Whether your workforce is made of people or models, the scoring is the same: doing everything is not the same as doing the thing that counts. Read the file. Close the deal. Escalate the locked door instead of knocking on it forever. And before you hire an AI workforce, wargame it — the full results and plain-language findings are public at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Layered Lighting Is a Must for Older Adults

For older adults, layered lighting enhances safety and visibility, but understanding how to implement it effectively can transform your home into a safer haven.

Customizing Smart Home Systems: Choosing the Right Platforms for Seniors

Unlock the key to seamless senior automation by exploring how to customize smart home platforms—discover which solutions truly meet their unique needs.

Home Safety Sensors: Fall Detection, Smoke Alarms, and More

Optimize your home safety with sensors like fall detectors and smoke alarms—discover how they can protect your loved ones and why ongoing maintenance matters.

10 Smart Home Features That Enhance Safety for Aging in Place!

Join us as we explore 10 innovative smart home features that ensure safety while aging in place – you won’t want to miss these essential tools!