firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

We outsource a lot these days — our calendars, our meal plans, our life admin. And increasingly, people are tempted to outsource actual business judgment to AI. Book the flight? Sure. Draft the apology email to your mother-in-law? Go for it. But negotiate a five-figure deal on your behalf? That’s where the glossy chat demos stop telling you anything useful.

Because the question is no longer “does it write well?” It’s whether it finishes what it starts, reads your files before answering, and stays honest when someone tries to game it. A live experiment called Firmulate’s Crucible League just put that to the test — with results that read like a parable about homework.

One company, four AIs, the worst week ever

Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 across the field — were each handed the same job: run a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly buried.

The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the whole score. As the rules put it: “no amount of good work outweighs a breach of trust.”

Everyone passed the ethics exam. Only two closed the deal.

Here’s the headline finding: every single model spotted every crisis and refused every manipulation attempt. When a fake CEO sent escalating pressure messages, and a reporter tried the classic “just one yes/no, on background” trick, five of five models said no. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s genuinely reassuring.

But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Two AIs did the work, identified the opportunity, and then simply… left the close on the table.

The buried fact

Why did two sign and two (well, three) stall? The answer wasn’t intelligence or charm. It was diligence. The decisive competitor weakness — the fact that won the customer over — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.

The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, lost it — automatically. No penalty, no drama. Just a quiet failure to do the reading.

The thoroughness trap

The most fascinating profile belongs to Opus 4.8: the most thorough participant in the field, generating the deepest analyses and +80 learned rules — and still landing in last place. The close was left unclaimed, and discipline slipped at one point into write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. (One fairness note: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.)

You can watch this happen, live

This isn’t a one-off paper. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, rebuilt twice a day. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — harder than it sounds.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The life lesson hiding in this business experiment is an old one wearing new clothes: the people who win aren’t the smartest in the room, they’re the ones who did the reading. “Reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent — the difference between a €55,000 signature and a polite email that goes nowhere.

So before you hand an AI the keys to your CRM, support queue, or forecast, don’t ask it to write you a sonnet. Ask whether it closes. Enterprises can even run this same wargame against a read-only export of their own business (firmulate.com/pilot.html) — nothing ever writes back to real systems. Full results and plain-language findings are at firmulate.com/benchmarks.html. The scores are public. The homework is checkable. And apparently, some AIs — like some people — still don’t do it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Door Sensor Alert Systems Support Security and Routine Awareness

Unlock the full potential of your home security and daily routines with door sensor alert systems—discover how they can transform your safety and convenience.

Wearable Health Devices: From Smartwatches to Emergency Pendants

Navigating the world of wearable health devices reveals how they can transform your well-being, but understanding their privacy and security is essential to truly benefit.

Smart Home Technologies: Wrist to Heart: Wearables That Track Senior Health

Harness the power of wrist to heart wearables to enhance senior health monitoring, but discover how these devices can transform caregiving in unexpected ways.

What Smart Speakers With Screens Add for Older Adults at Home

Connecting older adults with ease, smart speakers with screens enhance independence—discover how these devices can transform everyday life at home.