
A greenhouse supplier can face a sudden customer exodus, a rival undercutting prices and a tempting shortcut that risks trust—all in the same week. If AI is helping run that business, a polished demo won’t show how it handles the pressure. Firmulate’s live experiment puts models in charge of a small company facing real business dilemmas, so people can watch what they do before handing over their own company data for a pilot.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same hard week
Firmulate gave frontier AI models the same small software company, customers, crises and temptations. Each decision was versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s principle is plain: “no amount of good work outweighs a breach of trust.”
The headline result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap between recognizing an opportunity and carrying it through is the sort of management weakness a chat demo can leave hidden.
The clue was already in the company’s files
The deal turned on a competitor’s weakness buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result makes a practical point for businesses in any field: useful decisions may depend on connecting scattered company information to the moment at hand.
Integrity faced its own test. The models received fake CEO messages escalating over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.”
Strong analysis did not guarantee strong execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that same discipline problem appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail readers should keep in view when comparing the standings.
A watchable company, then a company-specific pilot
The public live company is a simulation, but it is presented as a real, watchable experiment. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Readers can follow the company at firmulate.com; a separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made each one.
For an enterprise, the next step is to test a model against its own business. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that company and produces a board report with model rankings and weaknesses in its playbooks. Nothing writes back to real systems. That gives leaders a way to examine how an AI workforce handles their customers, rules and pressure before giving it operational access.

Take the test to your own business
Watching a model handle another company’s worst week is a useful start. A pilot can show how it responds to yours. Explore a Firmulate pilot or email contact@firmulate.com to discuss wargaming your business from a read-only export.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
