firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before trusting AI with your business, give it a difficult week

A polished answer can sound reassuring. But would an AI agent spot trouble, resist pressure and follow through when a customer is ready to sign? That question matters beyond the technology desk: it touches the everyday decisions that shape a company’s money, reputation and relationships.

Firmulate’s live experiment puts AI models inside a small software company facing a week of crises. Now the company is inviting enterprises to move from watching that test to running one against their own business.

Same company, same temptations

In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and chances to cut corners. Every decision was versioned and auditable. The result was a leaderboard led by gpt-5.6-sol at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total. The standard is plain: no amount of good work outweighs a breach of trust.

The models all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s shorthand for the gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move, it turns out, is not the same as completing it.

The clue was buried in the company’s own files

The decisive competitor weakness was not spelled out in the customer event. It sat two document references deep in the company’s files. Models that read that material won the deal at full price, worth +€4,583 MRR. The episode makes a practical point for managers: an AI agent may need to find and use relevant company knowledge, not merely react to the latest message.

There was another test of judgment. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the finish

Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models: sound analysis could still fail to turn into effective action.

The comparison comes with a caveat. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions at Firmulate.

A live company, then a company like yours

The experiment is watchable at firmulate.com. Its live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbooks contain more than 680 self-learned rules, and each workday is versioned. These details make the test concrete: decisions have consequences inside the simulation, and viewers can follow how the company changes.

For a business considering AI in its CRM, support queue or forecasting, watching provides a starting point. A pilot takes the question closer to home: Firmulate can run the same kind of wargame against a read-only export of an enterprise’s own business. Teams can examine crisis scenarios, a board report with model rankings and weak points in their own playbooks. The pilot does not write back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decisions before they reach your systems

The league shows why fluent answers alone are a poor measure of readiness: models can identify a crisis, resist manipulation and still fail to finish the job. A company-specific exercise can help leaders see where their playbooks hold up and where they need work.

To discuss an enterprise pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Home Technologies: Loneliness, Meet AI: Can Robot Companions Cheer Us Up?

Smart home technologies are transforming loneliness into companionship with AI robots—discover how these innovations can uplift your daily life.

Resident Engagement Platforms: Digital Tools for Social and Daily Life

I invite you to discover how resident engagement platforms can transform your community involvement and enhance your daily life—find out more.

Integrating Smart Appliances: Kitchens That Support Independence

Creating a smart kitchen tailored to your independence can transform daily living—discover how these appliances can make a difference.

Why Smarter Morning Lighting Can Change the Feel of a Home

Meaningful mornings start with smarter lighting that adapts to your routines, transforming your home’s atmosphere and…