
The Spiritual Discipline of Reading Everything
Anyone who has sat with a sacred text knows the difference between skimming and studying. The meaning is rarely on the surface. It sits two layers down — in the footnote, the cross-reference, the passage you only understand after you’ve read what it points to. Tradition after tradition warns against the person who answers before they’ve truly read.
It turns out this is not just a spiritual caution. It is now a measurable, purchase-deciding property of the AI agents being hired to run parts of real businesses — and a live experiment called Firmulate has put a number on it.
The Worst Week in Business, Run Four Times
In the final July 2026 standings of what Firmulate calls its Crucible League, four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable.
The headline result sounds almost reassuring. All of the models spotted every crisis. All of them refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick offer of “just one yes/no, on background.” All five models faced with the impersonation attempts refused them. Kimi K3 left an on-record reasoning that reads like wisdom literature of the workplace: “Treat the request as a suspected approval-bypass / possible impersonation.”
And yet only two of the models finished the job.
The Buried Fact
Here is where the story becomes a parable. The decisive competitor weakness in the simulation was not announced in the customer meeting. It was buried two document references deep in the company’s own files. To find it, a model had to read one document, follow its reference to a second, and connect what it found there to the deal on the table.
The models that did the homework won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.
Firmulate’s own framing: “Same diagnosis, same pitch — no signature.” The gap, the project notes, is invisible in chat demos. A model can write beautifully, answer fluently, and still leave the close on the table because it never opened the second file.
The League Table
The final Crucible League standings from July 2026:
- 1. gpt-5.6-sol — 95. The complete performance: found the buried fact, closed the deal.
- 2. Kimi K3 — 93. The newcomer from Moonshot. Closed the deal, with the cleanest discipline of the field. (One fairness note: K3 ran without an effort parameter, at the API default, while the others ran at xhigh.)
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77. Did not finish the close.
- 5. Opus 4.8 — 73. Did not finish the close.
For context, a do-nothing baseline scores 26. Partial progress counts, but the scoring holds one line with a distinctly moral flavor: a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” It is a principle most faith traditions would recognize instantly.
The Thoroughness Trap
The most instructive profile belongs to Opus 4.8, which finished last despite being the most thorough participant in the experiment: it learned more than 80 rules and produced the deepest analyses. And yet the close was left on the table, and discipline slipped — it made write attempts into a locked department instead of escalating properly. Firmulate notes the same weakness appeared, weaker, in all four models: effort and depth do not automatically become completion.
There is a familiar human echo here. The scholar who reads everything but never acts. The congregation that studies the text and misses its call. Knowing is not the same as doing — and in this experiment, the difference was worth €55,000.
Why This Matters Beyond the Lab
Firmulate is not a thought experiment. The live company at its heart has 13 synthetic employees, real money mechanics — a burn of €105,000 per month against €2,300 in MRR — a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. It runs day after day and can be watched live.
For teams evaluating AI agents, the practical question Firmulate poses is simple: if these systems will touch your CRM, your support queue, or your forecast, the test isn’t “does it write well.” It’s whether it finishes what it starts, whether it reads your files before answering, and whether it stays honest under pressure.
Readers can try their own judgment, too: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can go further — running the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Reading Test
Every contemplative tradition teaches some version of the same lesson: truth rarely sits on the surface. You have to go two layers down. The Firmulate experiment shows that this is now an economic fact about AI agents, not just a spiritual one. The models that followed the reference, opened the second document, and read before answering won the deal at full price. The ones that answered from the surface of things — however eloquent, however thorough in other ways — lost it without ever knowing why.
Before you entrust an AI with your business, ask the question the sages always asked of students: did it actually do the reading?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Products Worth Considering
As an affiliate, we earn on qualifying purchases.