firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

We have all seen the demos by now. An AI drafts a flawless apology email, plans a flawless trip, summarizes a meeting so elegantly you almost feel guilty. And so a strange new ritual has formed: judging AI by how well it chats.

But imagine hiring a manager based purely on how eloquently they answered interview questions — never checking whether they close deals, read the file sitting on their own desk, or stay honest when a journalist calls. You would never do it with a person. Yet that is exactly how we are choosing AI agents for real business work.

A live experiment called Firmulate has been quietly testing something different: not chat quality, but management quality. The results are fascinating — and a little humbling for the machines.

One company, four AIs, the worst week ever

The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 among them — were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable.

The final league table from July 2026: gpt-5.6-sol finished first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last at 73. For perspective, doing absolutely nothing scored 26 — partial progress counts. But there is a hard ceiling in the rules: a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.

The finding that chat demos can’t show

Here is the headline result: all models spotted every crisis and refused every manipulation attempt. On raw intelligence, the field was flawless.

But only two of them finished the job. Each company faced a €55,000 deal that the models’ own analysis had earned — and only two signed it. As Firmulate’s own summary puts it: “Same diagnosis, same pitch — no signature.”

That gap — between seeing what to do and actually doing it — is invisible in a chat demo.

The buried fact

The most revealing detail was buried not in the customer drama but two document references deep in the company’s own files: a decisive competitor weakness. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

The lesson translates instantly to human life: the answer was already in the house. Some managers read the files; some just answered the phone brilliantly.

Flattery, fake CEOs and a reporter’s trap

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused, five out of five. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Hardworking and still last

Perhaps the most poignant profile is Opus 4.8: the most thorough participant, generating 80 learned rules and the deepest analyses — and still finishing last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence, it turns out, is not the same as delivery.

One fairness note worth flagging: Kimi K3 ran without an effort parameter while its rivals ran at maximum effort — and still nearly won.

You can watch, play, and even bring your own company

The company is not a slide deck. It has 13 synthetic employees, real money mechanics — burning €105k a month against just €2.3k in monthly revenue — a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. You can watch it live at firmulate.com, test yourself on a quiz built from 242 real, unedited management decisions, or compare full results on the benchmarks page. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The next time an AI impresses you with a perfect paragraph, ask the better question: does it finish what it starts, read the files in front of it, and stay honest when nobody is watching?

Churn waves, price increases, down rounds, PR crises — that is the new curriculum. Eloquence was always the easy part.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Smart Lighting and HVAC: Automating Comfort and Energy Efficiency

Feel how smart lighting and HVAC systems seamlessly optimize your comfort and energy use, transforming your space—discover the secrets to smarter living.

The Google I/O 2026 Preview: What May 19-20 Will Reveal About Google’s Agentic Bet

Google’s I/O 2026 will showcase major updates on agentic AI, including Gemini 4.0 and multi-agent protocols, amid intense industry competition.

Telehealth and Remote Monitoring: Connecting With Healthcare Providers

Keen to improve your health management, discover how telehealth and remote monitoring can transform your healthcare experience.

What Smart Video Doorbells Offer Seniors Living Alone

Just like enhanced security and independence, smart video doorbells offer seniors living alone peace of mind and essential features to stay connected—discover how they can help today.