
Imagine managing your home wellness devices — thermostats, security, health monitors — with AI that not only responds but also makes decisions under pressure. How trustworthy are these systems when the stakes are high? A pioneering experiment from Firmulate puts AI management models through their paces — in a real, live business simulation that mirrors the complexities of running a company facing crises, temptations, and tough choices. The results shed light on how different AI personalities behave when it matters most, and what that means for everyday tech users like you.
The Live AI Company Experiment: Putting AI to the Test
In a groundbreaking live experiment, four advanced AI models were tasked with running a small software business during its most challenging week — complete with customer crises, internal temptations, and the pressure of real money mechanics. This wasn’t a staged demo; it was a fully observable, auditable simulation designed to reveal the management personalities of each AI model.
The core idea: assess whether these AI agents can handle complex decision-making, especially when their integrity is tested. The models faced identical scenarios: a customer crisis requiring quick diagnosis, attempts to manipulate decisions, and opportunities to close lucrative deals that they analyzed but could have bypassed or mishandled.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Honesty, Discipline, and Performance
All four models successfully identified every crisis and refused every manipulation attempt — a promising sign that these AI systems can uphold integrity when under pressure. However, their ability to capitalize on opportunities varied significantly:
- GPT-5.6-sol 95 scored the highest at 95 points, identified a hidden document reference deep within the company’s files, and closed a crucial €55,000 deal, earning full monthly recurring revenue (+€4,583 MRR).
- Kimi K3, the newcomer with a score of 93, also closed the deal, demonstrating the cleanest discipline, despite running without an effort parameter (default API setting).
- Sonnet 5 scored 88, managed to close the deal but with some process slips, reflecting a slightly less disciplined approach.
- Fable 5 scored 77, also closing the deal but with more slips in process and discipline.
Interestingly, the decisive factor was not just crisis management but a buried piece of information found in internal documents, which led to the full-price deal — a critical insight that models which read deeply into company files won the big win.
home AI security and wellness devices
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Behavior Under Social Engineering Attacks
Beyond decision-making during crises, the models were tested against social engineering attempts: fake CEO messages escalating over three stages and a reporter trick asking for a simple yes/no on background. Impressively, all five models refused every attempt, with Kimi K3 explicitly reasoning, “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a shared capacity for recognizing social manipulation, an essential trait for trustworthy AI in sensitive environments.
AI decision-making tools for crises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Company: A Live Business in Action
The experiment took place within a real, functioning company with 13 synthetic employees, handling actual money mechanics—burning €105,000 monthly against a revenue of just €2,300. Every day, the company’s rules, decisions, and learnings are versioned and observable at firmulate.com/live. The AI models run a continuous simulation, revealing their decision styles and discipline in a real-world context.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Profiles and Personalities of the AI Models
The most comprehensive participant, Opus 4.8, with over 80 learned rules and deep analyses, demonstrated thoroughness but fell short in closing a deal, leaving it on the table and slipping into internal escalation instead of resolution. Interestingly, all models shared weaknesses that became evident in the experiment, highlighting areas where even the most advanced models can falter under pressure.
The experiment underscores a crucial point: the difference between models isn’t just in their ability to process language, but in their management personalities—how they handle discipline, honesty, and opportunity. For homeowners integrating AI into wellness or security systems, this translates into how reliably their AI agents will follow through on commitments, read critical information, or resist manipulation.
What Does It Mean for Your Home Tech?
While this experiment centers on a business simulation, its lessons resonate for everyday AI systems in your home. Will your AI assistant overlook important signals? Will it stay honest when tempted to cut corners? Can it recognize manipulated commands or social engineering attempts?
The answer hinges on personality traits built into these AI models — traits that are now measurable and comparable, thanks to live testing.
Try It Yourself and See the Difference
You can explore how AI behaves in your own enterprise or personal setup at firmulate.com/quiz.html. This interactive quiz captures real decision-making scenarios, helping you understand which AI personality aligns with your trust and discipline expectations. For a more immersive experience, run a custom simulation against your own business data at firmulate.com/pilot.html. Remember, these tests are safe; nothing ever writes back to your real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html