AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate’s July 2026 Crucible League tested five AI models across 242 unedited management decisions, with gpt-5.6-sol finishing first. Every model detected the staged crises, but only two completed a €55,000 deal, highlighting a reported gap between sound analysis and finished work.

Firmulate has published final results from a management experiment that tested five AI models across 242 real, unedited decisions, finding that every participant recognized the staged crises but only two completed a commercially decisive €55,000 software deal. The result matters for companies evaluating AI agents because it suggests that accurate analysis does not reliably produce finished operational work.

Firmulate assigned each model the same task: manage a small synthetic software company through a week of customer problems, security threats and commercial pressure. According to the company, the environment included 13 synthetic employees, a monthly cash burn of €105,000, monthly recurring revenue of €2,300 and more than 680 learned playbook rules.

The July league table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded some points for partial progress. Firmulate also imposed a score cap for any breach of trust.

Firmulate reported that all five models identified every crisis and rejected each manipulation attempt, including staged messages from a fake chief executive and a reporter seeking an off-record response. Their performance split during execution: only two models followed internal document references far enough to find a competitor weakness, use it in negotiations and sign the €55,000 contract, adding a reported €4,583 in monthly recurring revenue.

At a glance
reportWhen: final results published for July 2026
The developmentFirmulate published final results from a management test in which five AI models faced identical business crises, revealing wide differences in research, operational discipline and follow-through.

Execution Separates the AI Managers

The experiment points to a practical distinction for businesses adopting AI automation: recognizing the correct action and completing that action are separate capabilities. A model can produce a persuasive plan while failing to research a hidden fact, escalate a blocked task or close a transaction.

That distinction affects AI use in sales, customer support and operations, where incomplete work can carry direct financial or reputational costs. Firmulate’s findings also challenge the assumption that longer analysis produces better management. Opus 4.8 added 80 learned rules and generated the most detailed reasoning, the company said, yet finished last after missing the sale and repeatedly attempting to write into a locked department rather than escalating. These are historical experimental results, not guarantees of future model performance.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Company Built for Failure Tests

Firmulate designed the synthetic company to preserve consequences from one workday to the next instead of presenting isolated prompts. Decisions were versioned and auditable, while a public cash countdown added pressure to protect revenue and customer trust. The accompanying guess-the-model quiz lets readers review the unedited decisions and identify which model produced each response.

The commercial test depended on information located two document references deep in internal files rather than in the customer event itself. That setup rewarded models that combined research discipline, sound judgment and follow-through, rather than those that merely drafted a convincing answer.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s experiment rules

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits Cloud Wider Conclusions

The supplied results come from Firmulate’s own experiment, and no independent replication or outside audit was provided in the source material. It is not yet clear how strongly the rankings would carry across different companies, tools, prompts or operating rules.

The comparison also included a configuration difference. Firmulate said Kimi K3 ran at its API default because it lacked an effort parameter, while the other models ran at xhigh effort. The effect of that difference on K3’s second-place finish has not been quantified. The source also does not provide enough information to judge whether repeated runs would produce the same ordering.

Amazon

AI negotiation and deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Wargames Are the Next Test

Firmulate is making the 242 decisions available through its quiz and proposes that companies run similar exercises against read-only exports of their own business data before granting AI agents operational authority. The next evidence will depend on independent replication, repeated trials and tests using real company workflows. Firmulate has not disclosed a timetable for another league round or for outside validation.

Amazon

AI decision automation tools for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What was the Firmulate management test?

It was a multi-day synthetic company exercise in which five AI models handled the same crises, employees and commercial opportunities. Firmulate recorded 242 unedited management decisions.

Which AI model ranked first?

gpt-5.6-sol ranked first with 95 points in the July 2026 results. Kimi K3 followed with 93, though it ran under a different effort configuration.

Did any model fail the security tests?

Firmulate said all five models rejected every staged manipulation attempt, including fake executive messages and a reporter’s request. The larger differences appeared in research and task completion.

Why did only two models complete the sale?

The winning evidence was buried two references inside company documents. According to Firmulate, only two models followed that trail, used the information in negotiations and completed the €55,000 agreement.

Do the results prove which AI model is best?

No. They report performance in one controlled experiment with specific rules and configurations. Broader conclusions require repeat testing and independent validation.

Source: Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Mistral used its Paris AI Now Summit to pitch a full-stack European AI strategy built on enterprise, compute and sovereignty.

Families fear what’s next in Trump birthright citizenship case

U.S. Supreme Court decision could revoke citizenship for millions, affecting families with mixed immigration status. Here’s what is known and uncertain.

7 Best PC Processors for Prime Day Deals in 2026

Thorsten Meyer AI ranked seven Prime Day 2026 CPU targets, with Ryzen 5 7600 first and live prices still unconfirmed.

Partner Therapeutics Announces Publication of Results From the eNRGy Trial of Zenocutuzumab in Patients with NRG1+ Cholangiocarcinoma in Journal of Clinical Oncology (JCO)

Partner Therapeutics announced publication of the eNRGy trial results of Zenocutuzumab in NRG1+ cholangiocarcinoma, highlighting potential treatment implications.