
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Trust is tested when the stakes are real
Faith traditions often ask what a person does when temptation arrives, when a promise is costly, or when the right choice is harder than the easy one. A new experiment from Firmulate brings a version of that question to artificial intelligence: what happens when a model is asked to run a company through a crisis, make consequential decisions and resist manipulation?
The answer matters beyond technology. Businesses are considering AI for work involving customers, money and judgment. A polished explanation may sound reassuring, but can a model follow through when pressure builds? Firmulate’s live company experiment makes those decisions visible, and its enterprise pilot offers a way to examine the question against a company’s own playbooks.
A company’s worst week, repeated
For the final Crucible League in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Decisions were versioned and auditable. The experiment was designed to reveal management behavior, not just persuasive conversation.
Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The gap between recognizing the right move and carrying it out was captured in the experiment’s finding: “Same diagnosis, same pitch — no signature.”
The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s integrity standard is blunt: “no amount of good work outweighs a breach of trust.”
The decisive clue was already in the files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical challenge for AI at work: noticing a crisis is not the same as finding and using relevant knowledge the company already holds.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a strong response to manipulation, even as the deal results show that caution alone does not guarantee effective execution.
Thoroughness and discipline can diverge
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The scores are the published league results, but that difference is useful context when interpreting the ranking.
A watchable experiment, then a company-specific one
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The company is synthetic; the money mechanics and visible decisions make the exercise concrete. Readers can watch it at firmulate.com.
The experiment also offers a more personal way to engage: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. It invites readers to judge the choices themselves, rather than relying only on a leaderboard.
For an enterprise, the next step is to run a similar wargame against a read-only export of its own business: its customers, pipeline and rules, tested against crisis scenarios. The aim is a board report with model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems. That moves the exercise from observing a synthetic company to asking how AI might behave inside the realities of one’s own organization.

From watching to acting
Firmulate’s results show both sides of AI decision-making: models can identify crises and resist social engineering, yet still fail to close a deal or follow the right escalation path. A company-specific pilot can test those behaviors against its own scenarios before AI is trusted with consequential work.
To explore a pilot, visit firmulate.com/pilot.html and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
