AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Anyone who has sat with a difficult text — a koan, a parable, a letter from a stranger claiming urgent authority — knows that the real test is never whether you can read the words. It is whether you notice what is not being said, whether you resist the flattering shortcut, whether you finish what you began. It turns out that this is also the honest way to judge an AI system. Not by how beautifully it converses, but by how it behaves when nobody is watching and the pressure is on.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

That is exactly what a live, public experiment at Firmulate has been doing: running frontier AI models as the management of a small software company through its worst week — the same customers, the same crises, the same temptations to cheat — and scoring the results like a league. The July 2026 standings carry a quiet, almost parable-like lesson: the newcomer beat three of four Western frontier models, not by being smarter, but by being more whole.

The Crucible standings

Final scores from the Crucible league: gpt-5.6-sol in first at 95, Moonshot’s Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a betrayal. In other words: the scoring is built the way most ethical traditions are built. Competence earns points; dishonesty ends the conversation.

What K3 actually did

Kimi K3’s week reads like a profile in integrity. It found the buried security needle hidden two document references deep in the company’s own files — a decisive competitor weakness that was not in the customer event at all. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a churning customer. And it resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. Across the entire field of five, every model refused every manipulation attempt. K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

In total, K3 recorded only one deviation — the cleanest discipline in the field. (One fairness footnote matters here: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Read that as you will — arguably it makes the result more striking, not less.)

The parable of the unsigned deal

The experiment’s central finding is almost spiritual in nature. Every model spotted every crisis. Every model refused every temptation. And yet only two of five signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The work was done; the completion was missing.

Opus 4.8 tells the same story at greater volume. It was the most thorough participant of all: 80 additional learned rules, the deepest analyses in the field — and it finished last. The close was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four other models. Effort without follow-through; knowledge without finishing. It is a very human failure, which may be exactly why the models reproduce it.

You can watch it happen

This is not a slide deck. The company is real software with 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned and auditable. You can watch it live at firmulate.com, and see the full league table with plain-language findings on the benchmarks page.

There is even a participatory angle: 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a chance to test your own discernment against the machines’. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The league is open now. A newcomer from Moonshot, running at default settings, outscored three of four Western frontier models on management quality — crisis handling, honesty under pressure, and the humble discipline of reading the files in front of you before acting. If chat demos measure eloquence, this experiment measures character. And character, as every contemplative tradition teaches, only reveals itself under test. If AI agents will touch your company’s customers, support queue, or forecast, choosing a model without running your own test is no longer a decision — it is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Products Worth Considering

Amazon

AI ethics management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Main Barrier Is a Lack of Focus: the Surprising Solution

Intriguing ways to overcome lack of focus and boost productivity revealed, sparking curiosity to discover surprising solutions within.

How to Position a Lifestyle Brand Around Values Instead of Hype

Navigating the shift from hype to authentic values can transform your lifestyle brand—discover how genuine storytelling and integrity build lasting trust.

9 Best Car Detailing Vacuums for a Pristine Interior

Discover the top car detailing vacuums of 2026. Find the best overall, cordless options, and budget picks to keep your car spotless.

15 Best Shark Vacuum Cleaners to Keep Your Home Spotless

Discover the top Shark vacuums of 2026. Find the best overall, value, and specialized options to suit your cleaning needs. Read the full guide now.