
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Why Your Pool Partner Needs More Than Just Diligence
Think about the last time you hired a contractor or a service provider for your backyard pool or patio. You probably looked for someone reliable, thorough, and honest. But what if the machinery behind the scenes — the AI systems helping manage your project — are equally diligent but still miss the mark? That’s the core story from a recent experiment with advanced AI models, revealing that even the most thorough systems can falter when it counts most.
As an affiliate, we earn on qualifying purchases.
Behind the Experiment: Testing AI in High-Stakes Scenarios
In a live, observable experiment, four leading AI models were tasked with managing a small software company’s worst week — a scenario packed with crises, temptations, and the need for disciplined decision-making. The models faced customers, crises, manipulative tactics, and internal rules, all designed to simulate real business pressures. Every decision was recorded and could be reviewed afterward, making this a transparent test of AI performance under stress.
At the top of the leaderboard was GPT-5.6, which scored an impressive 95 out of 100, demonstrating its ability to identify buried critical facts and successfully close a lucrative deal. Kimi K3 followed closely with a 93, showing the most disciplined approach among all models. Sonnet 5 scored 88, and Fable 5 scored 77. In stark contrast, a baseline ‘do-nothing’ model scored just 26 — a reminder that effort alone isn’t enough.
What the Models Got Right
- All four AI models detected every crisis that arose during the week.
- They refused every attempt at manipulation, such as fake CEO messages or subtle approval-bypasses, maintaining integrity under pressure.
These findings highlight an essential point: advanced AI systems can be very good at seeing problems and resisting bad influences. However, the challenge lies in what happens after they identify opportunities or threats.
The Critical Shortfall: The Missing Close
Despite their diligence and honesty, only half of the models managed to close the deal that would have brought in €55,000 of revenue. GPT-5.6 and Kimi K3 succeeded, but the others faltered. The culprit? A subtle discipline slip: the models failed to follow through on the final steps, leaving the deal on the table. Instead of escalating or completing the process, they rerouted attempts into a locked department, effectively abandoning the opportunity.
Interestingly, the decisive factor wasn’t in what they missed in the immediate crisis, but buried deeper in the company’s files. Two document references held the key to closing the deal—information that the models that read those files discovered and leveraged, sealing the agreement at full price.
The Broader Lessons for Business and AI
This experiment underscores a vital truth: diligence and thorough analysis alone do not guarantee success. For AI systems to be genuinely effective in managing complex tasks — whether in customer support, sales, or operations — they must excel not just in identifying problems but also in acting decisively and following through.
In the context of managing a real-world business, this means AI should be trained and tested not only for detection and integrity but also for discipline and prioritization. The models that performed best knew what to focus on and how to complete the process, avoiding slips that cost real revenue.
Why This Matters for Your Business
If AI is touching your CRM, support queue, or forecasting tools, ask yourself: does it just spot issues, or does it finish what it starts? Does it read your company’s files deeply enough to uncover hidden opportunities? Most importantly, can it stay honest and disciplined under pressure? These qualities determine whether an AI system is a skilled worker or just a diligent observer.
Running your own AI wargame—testing your models against real-world crises—can help identify weaknesses before deploying AI at scale. Firms can simulate scenarios, see where models slip up, and reinforce discipline and prioritization. That’s the value of the live experiment conducted by Firmulate, where companies can watch AI in action, managing real money and real crises, in a safe environment.

Key Takeaway: Focus on Impact, Not Just Volume
In the end, the most thorough AI models still lost because they lacked the discipline to follow through. Diligence alone isn’t enough; prioritization and decisive action matter more. For businesses relying on AI, the lesson is clear: ensure your AI isn’t just diligent but also disciplined, focused, and capable of completing what it starts. Testing your AI workforce in simulated crises can reveal these critical weaknesses before they cost you real money.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.