firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A beauty brand’s worst week might bring a supplier delay, a wave of customer complaints or pressure to approve a questionable promotion. If AI agents are going to help run a business, the useful question is not just whether they can write a polished reply. It is what they do when decisions collide, and whether they follow through.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get beauty and skincare favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate is testing that question by putting AI models in charge of the same small software company through the same crises and temptations. The experiment is live and watchable at firmulate.com. Its enterprise pilot takes the idea from observation to a company’s own playbooks and scenarios.

A controlled bad week

In Firmulate’s final Crucible League, published in July 2026, frontier models faced an identical company and set of circumstances. Every decision was versioned and auditable. The results ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

For a beauty or personal-care business, the point is practical. A model might recognize a service problem or spot a chance to win a customer, yet still need sound judgment about what it can promise, what it should protect and when it should escalate.

Seeing the problem is not the same as solving it

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The report’s summary gets at the gap: “Same diagnosis, same pitch — no signature.” A system can identify the right move in its reasoning and still leave the opportunity untouched.

The decisive clue was not in a customer event. It sat two document references deep in the company’s own files, describing a competitor weakness. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode shows why testing against a company’s own information and procedures may reveal more than asking a model to explain a hypothetical situation.

Pressure tests for trust and discipline

The experiment also tried social engineering: fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Refusing manipulation was not the whole story. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried writing into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. The results put follow-through and respect for boundaries alongside crisis recognition.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The ranking is the league’s result under those conditions.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com.

For enterprises, the next step is a pilot using a read-only export of the business. The company can test crisis scenarios against its own information and receive a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. That boundary makes the exercise a way to examine how agents might behave before they are put near operational decisions.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test judgment before handing over responsibility

The league’s clearest lesson is that recognizing a crisis, resisting manipulation and completing the right action are separate tests. A company-specific wargame can make those gaps visible using its own business context, while keeping the real systems untouched.

To run a pilot against your company’s read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Are Cordless Hair Dryers Worth It? Pros and Cons

Discover whether cordless hair dryers are a smart investment. Explore their benefits, drawbacks, latest tech, and if they suit your styling needs.

A Definitive List Of Celebrity Stylists Who Changed The Way We Dress Over The Past 20 Years

A comprehensive list of influential celebrity stylists who have reshaped fashion trends and industry standards over the past two decades.

The Best Dressed Stars Of The Week Did Minimalism-Versus-Maximalism

This week’s top celebrity looks highlight a clear contrast between minimalist and maximalist fashion styles, sparking widespread interest.

Lili Reinhart’s Butter Yellow Dress Is A Nod To Kate Hudson’s How To Lose A Guy In 10 Days Look

Lili Reinhart wore a butter yellow dress reminiscent of Kate Hudson’s iconic style in ‘How to Lose a Guy in 10 Days,’ sparking fashion comparisons and fan interest.