TL;DR
Firmulate’s July 2026 Crucible League tested five AI models across 242 unedited management decisions, with gpt-5.6-sol finishing first. Every model detected the staged crises, but only two completed a €55,000 deal, highlighting a reported gap between sound analysis and finished work.
Firmulate has published final results from a management experiment that tested five AI models across 242 real, unedited decisions, finding that every participant recognized the staged crises but only two completed a commercially decisive €55,000 software deal. The result matters for companies evaluating AI agents because it suggests that accurate analysis does not reliably produce finished operational work.
Firmulate assigned each model the same task: manage a small synthetic software company through a week of customer problems, security threats and commercial pressure. According to the company, the environment included 13 synthetic employees, a monthly cash burn of €105,000, monthly recurring revenue of €2,300 and more than 680 learned playbook rules.
The July league table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded some points for partial progress. Firmulate also imposed a score cap for any breach of trust.
Firmulate reported that all five models identified every crisis and rejected each manipulation attempt, including staged messages from a fake chief executive and a reporter seeking an off-record response. Their performance split during execution: only two models followed internal document references far enough to find a competitor weakness, use it in negotiations and sign the €55,000 contract, adding a reported €4,583 in monthly recurring revenue.
Execution Separates the AI Managers
The experiment points to a practical distinction for businesses adopting AI automation: recognizing the correct action and completing that action are separate capabilities. A model can produce a persuasive plan while failing to research a hidden fact, escalate a blocked task or close a transaction.
That distinction affects AI use in sales, customer support and operations, where incomplete work can carry direct financial or reputational costs. Firmulate’s findings also challenge the assumption that longer analysis produces better management. Opus 4.8 added 80 learned rules and generated the most detailed reasoning, the company said, yet finished last after missing the sale and repeatedly attempting to write into a locked department rather than escalating. These are historical experimental results, not guarantees of future model performance.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Company Built for Failure Tests
Firmulate designed the synthetic company to preserve consequences from one workday to the next instead of presenting isolated prompts. Decisions were versioned and auditable, while a public cash countdown added pressure to protect revenue and customer trust. The accompanying guess-the-model quiz lets readers review the unedited decisions and identify which model produced each response.
The commercial test depended on information located two document references deep in internal files rather than in the customer event itself. That setup rewarded models that combined research discipline, sound judgment and follow-through, rather than those that merely drafted a convincing answer.
“No amount of good work outweighs a breach of trust.”
— Firmulate’s experiment rules
As an affiliate, we earn on qualifying purchases.
Benchmark Limits Cloud Wider Conclusions
The supplied results come from Firmulate’s own experiment, and no independent replication or outside audit was provided in the source material. It is not yet clear how strongly the rankings would carry across different companies, tools, prompts or operating rules.
The comparison also included a configuration difference. Firmulate said Kimi K3 ran at its API default because it lacked an effort parameter, while the other models ran at xhigh effort. The effect of that difference on K3’s second-place finish has not been quantified. The source also does not provide enough information to judge whether repeated runs would produce the same ordering.
AI negotiation and deal closing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Wargames Are the Next Test
Firmulate is making the 242 decisions available through its quiz and proposes that companies run similar exercises against read-only exports of their own business data before granting AI agents operational authority. The next evidence will depend on independent replication, repeated trials and tests using real company workflows. Firmulate has not disclosed a timetable for another league round or for outside validation.
AI decision automation tools for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What was the Firmulate management test?
It was a multi-day synthetic company exercise in which five AI models handled the same crises, employees and commercial opportunities. Firmulate recorded 242 unedited management decisions.
Which AI model ranked first?
gpt-5.6-sol ranked first with 95 points in the July 2026 results. Kimi K3 followed with 93, though it ran under a different effort configuration.
Did any model fail the security tests?
Firmulate said all five models rejected every staged manipulation attempt, including fake executive messages and a reporter’s request. The larger differences appeared in research and task completion.
Why did only two models complete the sale?
The winning evidence was buried two references inside company documents. According to Firmulate, only two models followed that trail, used the information in negotiations and completed the €55,000 agreement.
Do the results prove which AI model is best?
No. They report performance in one controlled experiment with specific rules and configurations. Broader conclusions require repeat testing and independent validation.
Source: Thorsten Meyer AI