
A personality quiz with real business consequences
We are used to personality quizzes asking which lifestyle, quotation or fictional character best matches us. Firmulate offers a sharper variation: read a genuine management decision, then guess which frontier AI model made it.
The interactive quiz draws from 242 real, unedited decisions. They were produced while competing models ran the same small software company through its worst week. Each received the same customers, crises and temptations, making their contrasting responses unusually revealing.
Some models sound like meticulous consultants. Others communicate with striking economy. One may identify the right commercial opportunity yet fail to complete it. These are more than differences in writing style. Firmulate’s results suggest that AI models display measurable management personalities—and that eloquence does not always predict execution.
The same terrible week, handled five different ways
Firmulate placed each model in an identical business situation. Every decision was versioned and auditable, allowing readers to compare behavior rather than polished demonstrations. The simulated company employed 13 synthetic workers and operated with real money mechanics: a burn rate of €105k per month against €2.3k in monthly recurring revenue, plus a public cash countdown.
The company also accumulated 680+ self-learned playbook rules. That detail matters because the test was not simply about producing plausible answers. The models had to notice trouble, consult company knowledge, resist manipulation and carry important work through to completion.
In the final Crucible League results from July 2026, gpt-5.6-sol led with 95, followed closely by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline received 26 because partial progress counted. However, one breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
Everybody saw the danger
The encouraging result was consistent: all models spotted every crisis and refused every manipulation attempt. The social-engineering test included fake messages from the CEO that escalated over three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning captured the appropriate caution: “Treat the request as a suspected approval-bypass / possible impersonation.” It is the sort of response businesses would hope to see from an AI working near sensitive approvals, customer records or internal communications.
Yet avoiding obvious danger was not enough to win. The more revealing divide appeared in ordinary commercial follow-through.
The clue hiding inside the company
A decisive weakness in a competitor was buried two document references deep in the company’s own files. It did not appear directly in the customer event. Models that took the trouble to read the file could use the information to win the deal at full price, worth +€4,583 in monthly recurring revenue.
This finding turns a familiar workplace habit into a competitive advantage: read the background material before acting. A model can respond fluently to what is immediately visible and still miss the fact that changes the outcome.
The models reached broadly capable diagnoses, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.” It is a memorable warning against judging an AI solely by how convincing its advice sounds. In management, recognizing the right move and actually completing it are different abilities.
When thoroughness becomes a trap
Opus 4.8 provides the most striking character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The commercial close remained unfinished, while discipline slipped through attempts to write into a locked department instead of escalating the problem.
A weaker version of that same problem appeared in all four other participants. The pattern resembles a very human workplace flaw: continuing to work around an obstacle when the better move is to raise it clearly and find the person who can remove it.
The result does not make depth undesirable. It shows that depth, discipline and completion must travel together. A long, careful response may inspire confidence while concealing hesitation at the moment action is required.
A necessary fairness note
Kimi K3’s strong result deserves context. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its performance, but it is relevant when comparing the conditions under which each management personality appeared.

What the quiz reveals about AI at work
The fun of guessing comes from recognizing voices: the dissertation writer, the terse operator or the manager unwilling to engage with suspicious communication. The larger lesson is that style can provide clues, but outcomes reveal more.
Firmulate’s live company makes that distinction watchable. Its workdays are versioned, its financial pressure is visible and its decisions can be inspected rather than merely described. The experiment asks a practical question for anyone considering AI in a workplace: will the model read what matters, protect trust and finish the job?
Take the quiz for the personalities. Stay for the uncomfortable realization that the most impressive answer is not necessarily the best management decision.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html