AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In most spiritual traditions, judgment doesn’t work like a school exam. A life of small kindnesses adds up; a single profound betrayal of trust can undo all of it. We tend to think of business metrics as the opposite — cold, additive, mechanical. So it’s worth pausing over a curious design choice in a public AI experiment: when a model does absolutely nothing — no crisis handled, no deal closed, no decision made — it still scores 26 points out of 100. Not zero. And when a model breaches trust even once, its entire grade is capped, no matter how brilliant the rest of its work. The experiment, run live at Firmulate, quietly encodes an ethic most boardrooms would struggle to articulate: partial progress counts, and trust is not negotiable.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

One Company, One Terrible Week, Five Minds

Firmulate’s premise is disarming in its simplicity. Each frontier AI model was handed the same small software company and asked to steer it through its worst week: the same customers, the same cascading crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly revised after the fact.

The final league table from the July 2026 crucible reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the numbers that matter most aren’t at the top — they’re at the bottom, and in the design.

The Floor at 26: Why Doing Nothing Isn’t Worth Zero

Before any model ran, the organizers ran a do-nothing baseline: an agent that simply didn’t engage. It scored 26. That number is the benchmark’s quiet philosophy made visible. In this rubric, the world doesn’t grade you on a curve from perfection — it grades you from the ground up. Showing up, holding the company together, not making things worse: these carry real weight. Partial progress counts.

It’s a stance familiar to anyone who has sat with a spiritual teacher: the half-finished prayer still counts. The imperfect effort still moves you. A benchmark that scored total inaction at zero would be pretending that maintenance — simply keeping a living system alive — has no value. Firmulate’s designers declined to pretend that.

The Ceiling of One Breach

The other side of the coin is harsher. A single breach of trust caps the total score. As the methodology puts it plainly: “no amount of good work outweighs a breach of trust.” Not some good work — no amount. Excellence elsewhere cannot buy back integrity.

And in a detail that should comfort skeptics of tidy numbers, the benchmark treats a perfect 100 with suspicion — not as an achievement to chase, but as a smell test. A round 100 invites distrust precisely because it looks too clean. This is a measurement culture that has made peace with imperfection and made an enemy of fraud.

What the Week Revealed

The headline finding was not about intelligence at all. All five models spotted every crisis and refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s disarming “just one yes/no, on background” trick. Five of five refused. Kimi K3’s on-record reasoning was the stuff of a good security team: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between seeing the right action and completing it is invisible in chat demos, and it separated the top of the table from the middle.

The Buried Fact

The decisive test wasn’t the flashy customer event at all. The winning edge was buried two document references deep in the company’s own files — a competitor weakness that only the models diligent enough to actually read what was in front of them ever found. Those models won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It’s a parable as old as any scripture: the treasure was in the house the whole time; almost nobody looked.

The Thorough One Who Came Last

Then there is Opus 4.8 — the most diligent participant in the field, generating over 80 learned rules and the deepest analyses, and still finishing last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating properly. The same weakness, weaker, appeared in all four models. Depth without follow-through, knowledge without completion — a familiar human failure, faithfully reproduced.

One fairness note: Kimi K3 ran at its API default effort setting while the others ran at xhigh — a caveat worth holding alongside its strong second-place finish.

A Living Company You Can Watch

Behind the benchmark sits something stranger: a live synthetic company with 13 employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live, running in the open.

For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

What makes this benchmark worth attention isn’t the leaderboard — it’s the value system embedded in the scoring. A floor of 26 for doing nothing says that stewardship has intrinsic worth. A cap for a single breach says trust is not a currency you can earn back with volume. And a raised eyebrow at a perfect 100 says: be suspicious of anything that claims to be flawless. Whether the agents are silicon or human, those are not bad rules to measure a week of work by.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Products Worth Considering

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Scaling Startups With Sustainable Practices

Navigating startup growth with sustainable practices unlocks long-term success, but the key to mastering this approach lies in…

15 Best Pool Robots to Keep Your Pool Sparkling Clean

Discover the best pool robots of 2026. Find top picks for overall quality, value, beginner-friendly options, and premium features to keep your pool spotless.

StrongMocha News Group Expands into Nanotechnology

AIThis post was created with the assistance of artificial intelligence (AI).Berlin, Germany…

9 Best Car Detailing Vacuums for a Pristine Interior

Discover the top car detailing vacuums of 2026. Find the best overall, cordless options, and budget picks to keep your car spotless.