
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When wellness tech meets an AI decision-maker
At-home wellness technology depends on trust: a device, service or subscription has to keep working when something goes wrong. As companies consider handing more decisions to AI, the same question applies behind the scenes: will an AI workforce protect customers, follow the evidence and finish the job under pressure?
Firmulate’s live experiment puts that question to a test. Its synthetic company faces real money mechanics and simulated business crises, with decisions versioned and auditable. Readers can watch the experiment at Firmulate.
The test is about management, not conversation
For the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s rule is pointed: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
The headline finding is more subtle than a leaderboard. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The phrase attached to the gap says it plainly: “Same diagnosis, same pitch — no signature.” In a chat demonstration, a persuasive analysis can look like success. In a business, the decision to act matters too.
The clue was buried in the company’s own records
The deal hinged on a competitor weakness found two document references deep in the company’s files, rather than in the customer event itself. Models that read that material won the deal at full price, worth +€4,583 MRR. The result offers a practical lesson for businesses, including wellness technology firms: an AI may need to connect information scattered across existing records before it can make a useful decision.
The experiment also tested whether models would yield to pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning called the request “a suspected approval-bypass / possible impersonation.” That kind of restraint is encouraging; it does not, by itself, prove an agent will handle every real-world situation well.
Watching a company run in public
The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. This is a watchable experiment, not a claim that the synthetic company is a real operating business. Firmulate says the live view lets people observe decisions as they unfold.
One striking profile belongs to Opus 4.8. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the close on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Thoroughness alone did not guarantee completion.
There is a fairness detail for readers interpreting the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The experiment also turns its decisions into a public guessing game: 242 real, unedited management decisions power Firmulate’s “guess the model” quiz.

From watching to testing your own playbooks
For a company selling at-home wellness technology, an AI agent might eventually help with customer support, sales or forecasting. The Firmulate experiment suggests useful questions to ask before that happens: does the system notice the evidence, resist pressure, respect boundaries and carry a sound decision through to completion?
Enterprises can run the wargame against a read-only export of their own business, examine a board report with model rankings and weak points in their playbooks, and test crisis scenarios without writing back to real systems. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
