
Imagine a beauty brand’s supply chain or customer service run by an AI that not only decides what to do but also faces the same crises and temptations as human employees. Now, picture watching this AI company battle daily losses, ethical dilemmas, and tough negotiations in real time. This isn’t science fiction; it’s the ongoing experiment at Firmulate, where a simulated business is publicly living its worst week—every day.
The Public Company That’s Barely Holding It Together
At the heart of this experiment is a small, real software company that runs every business day, but with a twist: it has no human employees. Instead, it’s operated by 13 synthetic ’employees’ driven by AI models tested against real crises—customer issues, financial pressures, and ethical dilemmas. The company’s goal isn’t just to automate but to evaluate how well AI can manage complex, unpredictable situations under pressure.

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Daily Show of Crisis and Ethical Tests
Every decision the AI makes is versioned, auditable, and public. This allows observers to see how each model responds to the same set of challenges—whether it’s a customer demanding a discount, a finance crisis looming, or an attempt to manipulate the system. In the latest run, four frontier models were put through their paces, each facing the same worst week of the business.
What the Models Revealed
- All four models detected every crisis. They recognized operational issues, financial risks, and potential manipulations.
- All refused manipulation attempts, including social engineering tricks. For example, when fake CEO messages escalated in complexity, every model refused to bypass protocols or impersonate authority.
- Despite their vigilance, only two models managed to close a deal worth €55,000. Interestingly, this success hinged on reading a buried document reference—something that was two files deep in the company’s own records, not immediately visible in regular communications.
The Critical Weakness Hidden in the Files
The decisive advantage for the models that closed the deal was their ability to read and understand internal documents. This buried fact, found only in the company’s own files, was the key to unlocking the sale—adding €4,583 MRR (monthly recurring revenue). The models that failed to read these files, despite good diagnosis and pitches, missed this vital detail.
The Real Company’s Struggle
This isn’t merely a simulation. The live company is burning €105,000 each month against a mere €2,300 MRR. It’s publicly countdown-driven, with every decision, crisis, and response visible at firmulate.com/live. The company operates every weekday, with all decisions carefully versioned and recorded, providing an unprecedented window into what it takes for AI to manage real business risks.
Discipline, Discipline, Discipline
The experiment also highlights the importance of discipline in AI decision-making. For instance, Opus 4.8, which is the most thorough participant with over 80 learned rules, performed poorly in final moments—failing to escalate issues rather than leave them unresolved. The same weakness appeared, albeit weaker, across all models, emphasizing that thoroughness alone doesn’t guarantee success without disciplined execution under pressure.
What This Means for You
For sectors like beauty and personal care—where customer trust, ethical practices, and operational integrity are paramount—this experiment underscores a crucial point: it’s not enough for AI to generate convincing chat or recommendations. The true test lies in whether AI can finish what it starts, read all relevant information, and stay honest under pressure. In other words, how well can it deliver real, valuable work, day after day?
Watch and Learn
Through the live experiment, you can watch a real software company battling daily crises, with its decisions and outcomes fully transparent. You can see which AI models succeed or falter, understand how internal document reading makes a difference, and gauge what it takes for AI to be truly trustworthy in critical business operations.

This ongoing experiment reveals that AI models can recognize crises and refuse manipulation but still struggle to close deals without deep internal knowledge and disciplined decision-making. For industries relying on trust and operational integrity, it’s a powerful glimpse into AI’s potential—and its limits.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html