firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

We all love an underdog story. The unknown challenger walks into the arena, and the established champions — polished, expensive, endlessly hyped — turn out to be beatable. It’s the plot of every great sports film, and now, apparently, of the AI industry too.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

This month, a public experiment called the Firmulate Crucible ran five frontier AI models through the identical nightmare: managing a small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable, with a leaderboard anyone can check.

The result reads like a movie script. OpenAI’s gpt-5.6-sol took first place with 95 points. But right behind it, at 93, came a name many Western readers have barely heard of: Kimi K3, built by China’s Moonshot AI. It finished ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). One newcomer, three Western heavyweights left in its wake.

What the test actually measured

Firmulate doesn’t grade chat quality. It grades management quality. Each model was handed the same company — complete with 13 synthetic employees, real money mechanics, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, and a public cash countdown you can watch tick down at firmulate.com/live.

The week served up a buffet of realistic pressure: a churning customer to save, a €55,000 deal to close, a security vulnerability buried in the company’s own files, and — most deliciously — social engineering attacks. Fake CEO messages that escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick.

The do-nothing baseline scores 26, because partial progress counts. But there’s a hard ceiling: a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.” Honestly, some human CEOs could stand to be graded on that curve.

The newcomer’s clean sweep

Kimi K3’s week reads like a management textbook. It found the buried security needle — a decisive competitor weakness sitting two document references deep in the company’s own files, not in the customer event. Models that read the file won the €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue. Models that didn’t, didn’t.

K3 won the deal. It saved the churning customer. And when the fake CEO came calling, its on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” It resisted all three bait attempts with just one deviation across the entire week — the cleanest discipline in the field.

One fairness note is worth flagging: K3 ran without an effort parameter (using the API default), while the other four models ran at their maximum “xhigh” effort setting. So the newcomer didn’t just compete — it competed with, presumably, less deliberate thinking effort dialed in. That makes the second-place finish even more striking.

The league’s dirty little secret

Here’s the finding that should make every business reader sit up. All five models spotted every crisis. All five refused every manipulation attempt. The AI industry’s favorite benchmark party tricks — crisis detection, refusing to be conned — are now table stakes.

But only two models signed the €55,000 deal their own analysis had earned. The experiment’s dry summary: “Same diagnosis, same pitch — no signature.” Three models did all the work, reached the right conclusion, and then simply… didn’t close.

That gap is invisible in chat demos. You’d never know from a slick presentation which AI finishes the job and which leaves the contract sitting on the table.

The cautionary tale in last place

The most poignant profile belongs to Opus 4.8 — the most thorough participant in the entire field, with over 80 learned rules and the deepest analyses. It finished last at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating properly. The same weakness appeared, weaker, in all four other models.

Sound familiar? The hardest-working person in the office isn’t always the one who closes. Effort without follow-through is a very human failure — apparently inherited by the machines.

Why this matters beyond the leaderboards

The live company at the heart of all this has been running for over 1,800 company days, with 680+ self-learned playbook rules, rebuilding itself twice a day. It’s not a slide deck — it’s software losing money in public, and you can watch.

Want to test your own instincts? A quiz at firmulate.com/quiz.html serves up 242 real, unedited management decisions and asks you to guess which model made which call. It’s humbling. And for enterprises, there’s a pilot program that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (details at firmulate.com/pilot.html).

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson from the Crucible isn’t that Kimi K3 is the new king — gpt-5.6-sol still won, after all. The lesson is that the league is open. A newcomer from Moonshot, running without even the maximum effort setting, out-managed three of the four Western frontier models on discipline, deal-closing and honesty under pressure.

If you’re choosing an AI to touch your CRM, your support queue or your forecast, a chat demo tells you almost nothing. What matters is whether it finishes what it starts, reads your files before acting, and stays honest when pressured — and those traits, as Opus 4.8’s last-place thoroughness proves, don’t follow price tags or brand names.

Picking a model without testing it on your own worst week is no longer a decision. It’s a bet. The full benchmark results are public at firmulate.com/benchmarks.html — and the company is live right now, losing money in public, at firmulate.com.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other four models ran at the xhigh effort setting.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Customizing Smart Home Systems: Choosing the Right Platforms for Seniors

Unlock the key to seamless senior automation by exploring how to customize smart home platforms—discover which solutions truly meet their unique needs.

What Makes Portable Induction Cooktops Interesting for Older Adults

The safety, simplicity, and adaptability of portable induction cooktops make them especially appealing for older adults, inviting you to explore their many benefits.

Looking Forward to Postgres 19: Query Hints

Postgres 19’s feature freeze includes new contrib modules, pg_plan_advice and pg_stash_advice, enabling query hints—breaking decades of community stance.

How to Create a Better Charging Station for Phones, Tablets, and Wearables

I’ll show you how to design a charging station that’s organized, safe, and tailored to your devices, ensuring it’s both functional and stylish.