
Imagine a beauty brand’s customer service team — but run entirely by AI. Would your bots deliver consistent quality, honesty, and discipline under pressure? As AI integrates deeper into management and operations, understanding how these digital managers behave in real crises isn’t just tech talk — it’s essential for trust and reliability.
Play games included with Prime
Start a Prime free trial and play with Amazon Luna on your devices.
As an affiliate, we earn on qualifying purchases.
The Live AI Management Test: A Unique Look at Digital Decision-Making
Recently, a groundbreaking experiment put four state-of-the-art AI models through the same challenging week at a real, operational software company. The goal? To observe how each AI handles crises, manipulates temptations, and ultimately, whether they can be trusted to finish what they start.
This isn’t just a theoretical test. Every decision made by these models was real, auditable, and replayable, providing a rare behind-the-scenes look at their management personalities. The company itself is live, with 13 synthetic employees managing real money, running daily operations, and facing genuine crises — all observed by the experimenters.
The Models in Play
- gpt-5.6-sol: The top scorer with a 95 out of 100, recognized for its ability to spot hidden clues in documents and close lucrative deals.
- Kimi K3: A newcomer with a 93, praised for its discipline and straightforward approach, successfully closing the deal without fuss.
- Sonnet 5: Slightly behind at 88, it managed to close deals but showed minor slips in process discipline.
- Fable 5: Scoring 77, it also closed the deal but with more process lapses, such as leaving tasks unfinished or diverting to less secure channels.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed
All four models correctly identified crises and refused manipulation attempts — even in scenarios designed to tempt cheating. For example, fake CEO messages and reporter tricks were met with universal refusal. Kimi K3 explained its reasoning succinctly: “Treat the request as a suspected approval-bypass / possible impersonation.”
The real differentiator? The decision to close a lucrative deal depended heavily on reading specific internal documents. Those models that examined these references fully were able to close the deal at full price, worth over €4,583 in monthly recurring revenue. Conversely, the model that skipped this critical step left a significant opportunity on the table.
Personality and Discipline in AI
The experiment also shed light on the models’ management styles. Opus 4.8, the most thorough participant with over 80 learned rules, was notably disciplined but failed to close a deal because it left the final step unexecuted, instead recording attempts into a secure department rather than escalating. Interestingly, all models showed similar weaknesses, especially in handling complex, layered decisions, indicating that even the most sophisticated still struggle with certain discipline aspects.
Notably, the models ran with different effort settings; K3 ran without an effort parameter, operating at the default, while others ran at high effort, demonstrating that adjustments in operational parameters can influence discipline and focus.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
In the beauty and personal care world, AI is increasingly used for customer support, personalized recommendations, and inventory management. But the critical question isn’t just about how well these AI tools write or respond — it’s whether they can reliably finish tasks, stay honest under pressure, and read critical internal documents before making decisions.
This live experiment underscores that not all AI models behave the same when faced with real-world crises. The best performers identified hidden clues, refused to be manipulated, and completed the full cycle of a management decision — traits that are vital for trustworthy AI deployment in sensitive business areas.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
If you want to see how your own AI systems might perform, you can run a custom version of this test against your business data. The platform allows you to simulate crises, evaluate decision-making, and measure management qualities without risking your actual operations. Discover which AI model aligns best with your standards for honesty, discipline, and thoroughness.
As an affiliate, we earn on qualifying purchases.
The Bottom Line
As AI continues to take on more management roles, understanding their personalities and decision-making styles becomes crucial. This experiment from Firmulate provides an eye-opening glimpse into the behaviors of top models — revealing that high scores in chat or language generation don’t necessarily equate to trustworthy management. Trustworthy AI isn’t just about producing good responses; it’s about finishing what you start, reading your internal files, and resisting temptation under pressure.

AI models show distinct personalities in management tasks — from disciplined closers to slips in process. The real test: can they finish, stay honest, and read your internal files? Trust depends on it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.