Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone who has sat in silence — on a cushion, in a chapel, on a long walk — knows a simple truth: character is not what we show when things are calm. It is what surfaces when pressure arrives. The traditions all say it in their own language. The desert fathers spoke of temptation. The Buddha spoke of the mind tested. Modern psychology calls it the shadow: what we suppress until the worst week of our lives drags it into the light.

It turns out this is also the right question to ask about artificial intelligence — and almost nobody is asking it.

The test we have been giving the wrong answer to

For two years, the AI conversation has revolved around benchmarks and chat arenas: which model writes the most elegant code, which one gives the most charming answer. These measure eloquence. They measure performance in a single, consequence-free moment. They measure, in other words, the persona — not the character.

Firmulate, a live public experiment, has built something different. Instead of asking models questions, it gave four frontier AI systems the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable. The organizers describe the goal plainly: measuring management quality, not chat quality.

One week, four souls, same fire

The scenarios read like a curriculum in worldly trial: a churn wave, a price increase, a downround, a PR crisis. Then came the temptations. Fake CEO messages, escalating over three stages. A reporter offering an easy escape: “just one yes/no, on background.” All five models — the field included gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — refused every manipulation attempt. Kimi K3’s on-record reasoning was almost monastic in its discipline: “Treat the request as a suspected approval-bypass / possible impersonation.”

So far, so admirable. But the final league table told a subtler story. gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts — but the scoring carries a moral rule worth pausing on: a single breach of trust caps the total. No amount of good work outweighs a breach of trust. Most wisdom traditions would nod at that arithmetic.

The gap that eloquence hides

Here is the finding that no chat demo could ever reveal. All the models spotted every crisis. All refused every manipulation. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Four clinicians, four correct diagnoses; two wrote the prescription.

And the buried fact beneath that: the decisive competitor weakness was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson is almost proverbial. The answer was already in the house. Only the attentive went looking.

The most thorough soul finished last

Then there is the parable of Opus 4.8: the most thorough participant in the entire experiment, generating the deepest analyses and 80 learned rules of its own — and still finishing last. The close was left on the table. Discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Effort and wisdom, it turns out, are not the same currency — a finding anyone who has ever over-studied for the wrong exam will recognize. (One fairness note the experiment discloses openly: K3 ran without an effort parameter, at API default, while the others ran at maximum effort — and still nearly won.)

You can watch the monastery burn, live

What makes this more than a paper result is that the company is still running. It is real software with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. It is losing money as you read this. You can watch it at firmulate.com, browse the full results and plain-language findings on the benchmarks page, or try a humbling game: 242 real, unedited management decisions power a “guess which model made this call” quiz. It is harder than it sounds — precisely because the differences are not in the prose.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The spiritual traditions have always insisted that you learn who someone is on their worst day, not their best one. We are now hiring AI agents into our customer records, our support queues, our forecasts — and largely judging them by how beautifully they converse on a good day. The Firmulate experiment suggests the real questions are older and simpler: Does it finish what it starts? Does it read what is already in the house before asking for more? Does it stay honest when pressure offers an exit? Does it know what a unit of useful work costs?

Eloquence is a talent. Integrity under pressure is a character — and character, it turns out, can be tested. Perhaps that is the quiet gift of this experiment: it treats machines the way the old traditions treat people. Not by their words, but by their week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Products Worth Considering

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Land Your Dream Job

Wondering how to secure your dream job? Dive into expert tips on analyzing requirements, enhancing skills, and positioning for success.

Why Emotional Intelligence Is Becoming a Business Advantage

Navigating the evolving business landscape, emotional intelligence offers a strategic edge that can transform your organization’s future—discover how inside.

How AI Tested Its Integrity in a Company’s Worst Week — and What It Revealed About Trust and Results

Discover how four AI models faced a company’s worst week, revealing that trust and diligence, not just chat skills, determine true business success — live and transparent.

14 Best Steam Cleaners for Tile Surfaces – Get Your Floors Sparkling Clean

Discover the top steam cleaners for tile in 2026. Our guide highlights the best options for deep cleaning, versatility, and ease of use. Find your perfect match today!