firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In the world of beauty and personal care, trust and efficiency can make or break a brand. But what if the very AI systems meant to boost your business are not always enough, even when they seem thorough? Recent live experiments with artificial intelligence reveal a surprising truth: sheer diligence and volume of effort don’t guarantee success. Instead, strategic prioritization and reading the right information matter most.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI’s Judgment Under Pressure

Firmulate, a leader in business emulation technology, ran a unique live experiment involving four advanced AI models. Each was tasked with managing a simulated small software company through its worst week—dealing with customer crises, internal temptations to cut corners, and the challenge of closing profitable deals. What set this experiment apart? Every decision was carefully versioned and auditable, mimicking real-world pressures in a transparent environment.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: All AI Models Showed Awareness, Few Delivered

Despite differences in architecture and training, all four models identified all crises and refused manipulative tactics such as fake CEO messages or reporter tricks. In fact, the models demonstrated a high level of integrity, refusing every attempt at deception. However, only two managed to secure the critical deal worth €55,000, which their own analyses had earned them—highlighting that diligence doesn’t always translate into impact.

The Hidden Weakness: Reading the Right Files Matters Most

Digging deeper, the experiment revealed a crucial insight: the decisive advantage was hidden not in reacting to customer crises but within internal company documents. The models that managed to read two document references deep into the company’s files discovered key information, enabling them to close the deal at full price. Conversely, the models that overlooked this internal data failed to capitalize on the opportunity, despite their thorough crisis management.

Human-Like Temptations and AI Integrity

Adding another layer of complexity, the experiment tested social engineering attacks. Fake CEO messages escalating over three stages and a reporter trick asking for a simple background quote were presented to all models. Every AI refused these manipulative requests, with the Kimi K3 model explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This underscores that, in controlled tests, AI can maintain ethical boundaries under pressure.

The Real-World Company: Facing Cash Burn and Rules

The experiment’s live company is a real-looking operation with 13 synthetic employees and a cash burn of €105,000 per month against a modest €2,300 monthly recurring revenue (MRR). Every day, the system self-learns and version-controls over 680 rules, simulating real business mechanics that demand disciplined decision-making. This environment is accessible for public viewing at firmulate.com/live.

The Most Thorough Participant Fails on Discipline

Among the models, Opus 4.8 was the most comprehensive, studying over 80 learned rules and performing the deepest analyses. Yet, it finished last in the final scoring—73 out of 100—because it left the critical deal on the table. The model’s discipline slipped when it failed to escalate certain write attempts into a locked department, instead leaving opportunities unpursued. This demonstrates that volume of effort alone cannot compensate for strategic focus and discipline.

Implications for Business Decision-Making and AI

What does this mean for industries like beauty and personal care, where brands increasingly rely on AI for customer engagement, support, and forecasting? The key takeaway is clear: success hinges not on how diligently an AI models work but on what they prioritize and read. A model’s ability to identify the most impactful information—reading deep into internal files—and maintain discipline under pressure can be the difference between closing a deal or leaving it on the table.

Tools for Testing Your Own AI Workforce

For enterprises eager to assess their AI readiness, Firmulate offers a unique opportunity. Running a ‘wargame’ against a read-only export of your business allows you to observe how your AI models handle crises, temptations, and priorities without risking real-world operations. This transparent, sandbox approach helps decision-makers improve their AI’s impact before deployment. More information is available at firmulate.com/pilot.html.

From the League Table: Which Model Wins?

  • gpt-5.6-sol scored 95 and managed to find the buried fact, closing the deal.
  • Kimi K3, with a score of 93, demonstrated the cleanest discipline and also secured the agreement.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Fable 5 scored 77, closing the deal with more process slips.

Surprisingly, models at different architectures showed that diligence alone isn’t enough—reading the right information and maintaining discipline matter most.

Final Thought: Diligence Isn’t the Same as Impact

For beauty brands and personal care companies, the lesson is straightforward: an AI that works hard but ignores the strategic core — the critical information in internal documents and disciplined decision-making — won’t necessarily deliver results. Prioritization, reading deeply, and maintaining integrity under pressure are the real keys to success in AI-driven business environments.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Fenty Beauty’s Newly-Launched Foundation Is Way Better Than The Original

Fenty Beauty’s new foundation has received widespread praise for its performance, surpassing the original formula according to early reviews and user feedback.

The Story Behind Elle Woods’s Most Iconic Looks in Legally Blonde’s Prequel, Elle

Details emerge on how costume designers recreated Elle Woods’s signature style for the upcoming ‘Elle’ prequel, shedding light on her fashion evolution.

Fenty Beauty Just Dropped The Replacement For Its Beloved Foundation

Fenty Beauty has introduced a new foundation replacing its beloved original. Details are confirmed, but reasons for the change remain unclear.

New Arrivals From HBX: Carne Bollente

HBX introduces new arrivals from Carne Bollente, sparking increased interest in the brand’s provocative designs amid rising search trends.