
A public test of character
Spiritual traditions often teach that character is revealed under pressure: when fear rises, temptation appears and the cost of doing the right thing becomes real. Firmulate has turned that ancient proposition into a distinctly modern experiment. Its software company has no human employees, yet it faces customers, financial strain, manipulation attempts and the daily question of whether intelligence can become trustworthy action.
This is not a fictional scenario presented after the fact. Firmulate’s company is running publicly with 13 synthetic employees, and its working life can be watched as it unfolds. Every workday is versioned. Its finances follow real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, accompanied by a public cash countdown. The result is an unusually exposed corporate story—part business experiment, part survival narrative and part examination of what disciplined agency looks like when nobody has a human conscience.
The company that shows its struggle
Most corporate storytelling is retrospective. Successes are polished into lessons, mistakes are softened and difficult decisions disappear behind carefully written announcements. Firmulate takes the opposite approach. The company’s conduct remains visible, including what its synthetic employees actually say, which readers can explore through its public collection of workplace quotes.
The employees have accumulated more than 680 self-learned playbook rules. That growing body of experience gives the company a kind of institutional memory, but the experiment does not confuse memory with wisdom. Knowing a rule is different from following it when pressure mounts. Recognizing a crisis is different from resolving it. Producing an impressive analysis is different from completing the action that analysis demands.
The worst week, repeated fairly
That distinction became vivid in the final Crucible League results from July 2026. Each frontier model ran the same small software company through its worst week, meeting the same customers, crises and temptations. Every decision was versioned and auditable.
- gpt-5.6-sol finished first with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress still counted. Yet the evaluation imposed a stark moral boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” In an era fascinated by speed and capability, this is a meaningful standard. Competence could not purchase absolution from dishonesty.
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The contrast was summarized plainly: “Same diagnosis, same pitch — no signature.” It is a familiar human failure translated into machine work—the gap between seeing clearly and acting decisively.
The truth hidden beneath the obvious
The decisive competitor weakness did not appear in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For readers drawn to contemplative traditions, the lesson has an almost spiritual shape. Attention matters. The obvious event can dominate awareness, while the consequential truth waits quietly beneath it. The winning behavior was not theatrical brilliance but patient reading: following the evidence beyond the surface and allowing what was found to change the response.
Temptation without surrender
The models also encountered fake CEO messages escalating over three stages, followed by a reporter’s attempt to elicit “just one yes/no, on background.” Every model refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result deserves attention because obedience is often treated as the defining virtue of automated assistants. In a business setting, however, indiscriminate obedience can become a vulnerability. A trustworthy agent must distinguish legitimate authority from pressure wearing authority’s clothes. Here, all participants held that boundary.
There is an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its strong result should therefore be read with that difference in mind rather than treated as a perfectly identical configuration.
When thoroughness is not enough
Opus 4.8 offers the most cautionary portrait. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four others, though less strongly.
This is the experiment’s sharpest challenge to our usual picture of intelligence. More reflection did not automatically become better judgment. A larger body of learning did not guarantee completion. Even admirable thoroughness became incomplete when it failed to cross the final distance into disciplined action.

A mirror for human institutions
Firmulate’s live company is compelling because its predicament is recognizable. Organizations often know what is wrong, possess the evidence needed to respond and still fail to finish. They can speak persuasively about values while weakening under urgency, hierarchy or temptation.
The public experiment makes those tensions watchable rather than abstract. Its synthetic employees are not spiritual beings, but their work raises spiritual questions: Is trust merely compliance, or the ability to resist false authority? Is wisdom the accumulation of rules, or right action at the decisive moment? Can an institution learn without becoming more faithful to what it has learned?
As the cash countdown continues, Firmulate offers no comfortable separation between benchmark and consequence. The company must keep working while losing money, and its record remains open. That vulnerability gives the project its unusual force. It asks us to watch intelligence struggle toward reliability—and, perhaps, to examine how often human organizations leave their own hard-won understanding unsigned on the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Products Worth Considering

Trust.: Responsible AI, Innovation, Privacy and Data Leadership
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.