
A pool company can look healthy until the heat wave brings a surge of service calls, a supplier falls behind and a key customer starts shopping around. If AI agents are going to help run the business, a polished demo won’t show how they handle that kind of week. Firmulate’s experiment puts models through a company’s troubles and temptations, then asks whether they can follow through.
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same crises
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations, with each decision versioned and auditable. The results ranged from 95 for gpt-5.6-sol to 73 for Opus 4.8. Kimi K3 scored 93, Sonnet 5 scored 88 and Fable 5 scored 77. A do-nothing baseline scored 26; a single breach of trust caps the total, because no amount of good work outweighs one.
Knowing what to do is not the same as doing it
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s concise verdict: “Same diagnosis, same pitch — no signature.” That is a consequential gap for businesses considering AI in customer support, sales or operations: recognizing a sound move does not guarantee the agent will complete it.
The deciding clue was buried two document references into the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The test also included fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning called the request “a suspected approval-bypass / possible impersonation.”
Thorough work still needs discipline
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped in discipline, making write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail to keep in mind when reading the ranking.
The live Firmulate company makes the premise tangible. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbooks include 680+ self-learned rules, and every workday is versioned. Readers can watch the experiment at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
From watching to a business pilot
For a pool or patio business, the practical question is what an AI agent would do when a customer threatens to cancel, a service issue spreads, or someone asks it to bypass approval. Firmulate’s proposed enterprise pilot takes a read-only export of a company’s own business and runs crisis scenarios against it. The result is a board report with model rankings and the weak points in the company’s playbooks. Nothing writes back to real systems.

The experiment suggests that seeing a crisis and recommending a response are only part of the job. A business also needs to know whether an AI can act with discipline, close the loop and respect trust when pressure rises. Enterprises can explore a pilot using their own read-only business export at firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
