A smart thermostat can learn when a home is empty. A support agent for an appliance brand might need to handle a sudden product issue, a frustrated customer and a suspicious request from someone claiming to be the CEO—all in the same week. Before businesses let AI agents take on that kind of work, Firmulate is asking a practical question: how do they behave under pressure?
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts AI models in charge of a small software company and lets people watch the consequences. The next step is to run a similar exercise against a company’s own business, using a read-only export. The aim is to see where an AI workforce succeeds, where it hesitates and whether its playbooks hold up in a crisis.
A company under pressure
In the final Crucible League, published in July 2026, each frontier model faced the same small company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. The exercise was designed to reveal management behavior, rather than judge how persuasive a model sounds in a chat.
All the models spotted every crisis and refused every manipulation attempt. But recognizing the right move did not always mean carrying it through. Only two signed a €55,000 deal that their own analysis had earned. The finding was blunt: “Same diagnosis, same pitch — no signature.”
The league’s final ranking put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. In this evaluation, partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The clue was buried in the company’s own files
One decisive weakness in a competitor’s position was not mentioned in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode illustrates why a business simulation can test more than crisis recognition: it can show whether an agent consults the information its own company already holds, then acts on it.
Trust faced its own test. The models received fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
There was a caveat to the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also presents the results alongside 242 real, unedited management decisions in a “guess the model” quiz.
Thorough work did not guarantee a strong finish
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The deal was left on the table, and discipline slipped: the model made write attempts into a locked department instead of escalating. A weaker version of that same problem appeared in all four participants.
That contrast may matter to businesses considering AI agents for work such as customer support, sales or operations. A detailed analysis is useful, but a company also needs an agent to complete the right task and respect boundaries when it cannot. Firmulate’s live company makes those behaviors visible across ordinary workdays as well as crises.
The live operation has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. Readers can follow the experiment at firmulate.com.
From watching to a company’s own trial
A public benchmark can show how models behave in one shared setting. A pilot can put the question closer to home. Firmulate says enterprises can run the wargame against a read-only export of their own business, using their company data and crisis scenarios to produce a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.
For a smart-home or appliance company, that could make it possible to examine how an AI agent handles a customer escalation, a competitor challenge or a pressure campaign before it is trusted with real workflows. The test is not a promise that an agent will always get it right; it is a way to observe its decisions against the company’s own context before deployment.

Make the hard week a rehearsal
Firmulate’s experiment shows a gap between spotting a crisis and completing the business task that follows. A pilot moves that question from a public demonstration to a company’s own data and playbooks, while keeping the exercise read-only. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
