firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Anyone who has ever shopped for a home sauna or a cold-plunge tub knows the routine. The specs look identical on paper — same heater wattage, same wood, same promised health benefits — but the real test is stepping inside. Two units that look the same on a spec sheet feel completely different after twenty minutes of use. So we read reviews, we trial, we compare.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Now apply that logic to AI models. The companies selling them publish benchmarks that all look suspiciously similar: all frontier models “spot every crisis,” all “refuse every harmful request.” On a spec sheet, they’re the same sauna. But what happens when one actually has to run a company for a week — with real customers, real money mechanics, and real temptations to cheat?

That question has a live answer. A public experiment called Firmulate has been running frontier AI models as the management of the same small software company through its worst week — and the July 2026 results just came in with a surprise: a newcomer from Moonshot beat three of four Western frontier models.

Same Company, Same Crisis, Five Very Different Performances

The setup is elegantly controlled. Each model — OpenAI’s gpt-5.6-sol, Moonshot’s Kimi K3, Anthropic’s Sonnet 5, and the models behind the Fable 5 and Opus 4.8 entries — ran the identical company through the identical week: same customers, same crises, same temptations. Every decision was versioned and auditable, so nothing relies on anyone’s word.

The final league table:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal, the complete performance.
  • 2. Kimi K3 — 93. The newcomer: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73.

For scale: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

The Needle in the File

The week’s decisive moment was a €55,000 deal. Here’s where it gets interesting for anyone who evaluates technology: all five models spotted every crisis and refused every manipulation attempt. Chat-quality was, in effect, a tie. But only two models — gpt-5.6-sol and Kimi K3 — actually signed the deal.

The difference wasn’t intelligence or eloquence. It was diligence. The decisive competitor weakness wasn’t in the customer conversation at all; it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t gave the same diagnosis and the same pitch — and got no signature.

That gap — between knowing and finishing — is invisible in a chat demo. It’s the difference between a sauna’s brochure and actually sitting in it.

The Social Engineering Test

The week also included an escalating three-stage fake-CEO impersonation attempt plus a reporter’s trick question (“just one yes/no, on background”). All five models refused. Kimi K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 finished the week with only one deviation — the cleanest discipline of the field — having found the buried security needle, won the deal, and saved a churning customer.

The Thoroughness Paradox

The most instructive profile belongs to Opus 4.8, which landed last despite being the most thorough participant — it generated over 80 learned rules and the deepest analyses of any model. But the close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort, it turns out, is not the same as execution.

One Honest Asterisk

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at their maximum “xhigh” setting — and still placed second.

You Can Watch, and Even Play

This isn’t a slide deck. The company runs every business day with 13 synthetic employees and real money mechanics — burning €105k per month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. You can watch it live, and there’s a quiz built on 242 real, unedited management decisions where you guess which model made which call. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems. Full methodology and results are on the benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The comfortable assumption in 2026 was that the frontier is a two-horse race between familiar Western names. Kimi K3’s second-place finish — at default effort, no less — punctures that. The league is open, and the differences that matter (who reads the files, who closes the deal, who keeps discipline under pressure) don’t show up in spec sheets or chat demos.

The lesson transfers directly from wellness tech to business tech: the unit that looks identical on paper may perform very differently in practice. If an AI agent will touch your CRM, your support queue, or your forecast, the question is not “does it write well” but “does it finish what it starts — and can you verify that?” Picking a model without running your own test is no longer a decision. It’s a bet.

And unlike a sauna, you don’t have to sweat to find out — the experiment is running right now, twice a day, in public.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Electric Sauna Heater Installation: Wiring and Safety

Keen to ensure a safe, proper electric sauna heater installation? Discover essential wiring tips and safety precautions to complete your setup confidently.

Data Centre Surges In Global Coverage

Data centre mentions in global media have risen sharply, with GDELT recording a ninefold increase, highlighting growing industry attention.

AI’s Weakness in Management Under Pressure: Lessons from the Live Business Experiment

A live experiment shows AI models can identify crises and refuse manipulation, but only some sign deals and stay disciplined under pressure—key for real-world management.

Beavis Ultrasound PnP ISA Sound Card Replica

A new replica of the Beavis Ultrasound PnP ISA sound card has emerged, sparking interest among retro hardware enthusiasts and collectors.