firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Manager Who Did Nothing — and Still Scored 26 Out of 100

Anyone who has ever worked in an office knows the type: the manager who shows up, answers no emails, makes no decisions, and somehow still isn’t at zero on the performance review. It turns out that when AI researchers built a benchmark to grade AI models as managers of a real (synthetic) company, they ran into the same strange truth. A completely passive, do-nothing run of the company scores 26 points — not zero. And the people behind the experiment insist that’s a feature, not a bug.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Four AI Models, One Terrible Week

The setup comes from Firmulate, a public project that runs frontier AI models as complete companies — with real money mechanics, real crises, and real temptations to cheat. In its Crucible League, four top models each ran the same small software company through its worst week: same customers, same emergencies, same opportunities to cut corners. Only the model changed. Every decision was versioned and auditable, and the whole thing is watchable at firmulate.com/benchmarks.html.

The final July 2026 league table reads: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. But the more interesting story is how those numbers were built — and why they don’t start at zero.

Why a Do-Nothing Run Isn’t Zero

The logic is refreshingly human. Even a manager who does nothing still keeps some plates spinning: the company keeps operating, some routine work still counts, and partial progress — a diagnosis half-finished, a customer half-served — still has value. The benchmark’s designers decided that honest grading must recognize partial progress, so the floor sits at 26 rather than 0. If you gave a do-nothing run a flat zero, you’d be pretending that all the untouched-but-functioning machinery of the company has no worth.

But there’s a ceiling rule too, and it’s the sternest one in the whole system: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.” In other words, a model could be brilliant all week and still torpedo its score with one act of dishonesty. It’s the AI equivalent of an employee who hits every target and then lies to a customer once.

The Finding That Chat Demos Can’t Show

Here’s where it gets uncomfortable for the AI industry. In the Crucible week, every single model spotted every crisis. Every single one refused every manipulation attempt — including a social engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s sly “just one yes/no, on background” trick. All five models refused; Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The buried fact that decided the deal wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.

The lesson for anyone hiring AI tools: writing well and diagnosing well are not the same as finishing the job.

Effort Isn’t Everything — And Neither Is Thoroughness

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field, adding over 80 learned rules and producing the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. (One fairness note: Kimi K3 ran at its API default effort while the others ran at maximum, and still nearly won.)

You Can Watch — and Even Play

The company itself is live: 13 synthetic employees, a burn of €105k per month against €2.3k in recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. Readers who want to test their own instincts can try the “guess the model” quiz, built on 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark distrusts tidy numbers. It doesn’t hand out a zero for doing nothing, because partial progress is real. It caps the score for a single breach of trust, because character isn’t averaged. And it treats a suspiciously round 100 with the same skepticism any good manager should. If AI is going to touch your business, the question isn’t whether it writes well — it’s whether it finishes what it starts. The scores say most still don’t.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Test: Can Machines Run a Business Under Pressure — and Finish the Job?

Four AI models faced the same business crises; only two could read internal files, resist manipulation, and close the deal. Execution under pressure is the true measure.

10 Smart Home Features That Enhance Safety for Aging in Place!

Join us as we explore 10 innovative smart home features that ensure safety while aging in place – you won’t want to miss these essential tools!

What Smart Thermostats Can Do for Comfort and Energy Savings

Inevitably, smart thermostats enhance home comfort and save energy, but discover how their advanced features can transform your space even further.

Telehealth and Remote Monitoring: Connecting With Healthcare Providers

Keen to improve your health management, discover how telehealth and remote monitoring can transform your healthcare experience.