AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s July 2026 Crucible League tested five AI models across 242 unedited management decisions, with gpt-5.6-sol finishing first. Every model detected the staged crises, but only two completed a €55,000 deal, highlighting a reported gap between sound analysis and finished work.

Firmulate has published final results from a management experiment that tested five AI models across 242 real, unedited decisions, finding that every participant recognized the staged crises but only two completed a commercially decisive €55,000 software deal. The result matters for companies evaluating AI agents because it suggests that accurate analysis does not reliably produce finished operational work.

Firmulate assigned each model the same task: manage a small synthetic software company through a week of customer problems, security threats and commercial pressure. According to the company, the environment included 13 synthetic employees, a monthly cash burn of €105,000, monthly recurring revenue of €2,300 and more than 680 learned playbook rules.

The July league table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded some points for partial progress. Firmulate also imposed a score cap for any breach of trust.

Firmulate reported that all five models identified every crisis and rejected each manipulation attempt, including staged messages from a fake chief executive and a reporter seeking an off-record response. Their performance split during execution: only two models followed internal document references far enough to find a competitor weakness, use it in negotiations and sign the €55,000 contract, adding a reported €4,583 in monthly recurring revenue.

At a glance
reportWhen: final results published for July 2026
The developmentFirmulate published final results from a management test in which five AI models faced identical business crises, revealing wide differences in research, operational discipline and follow-through.

Execution Separates the AI Managers

The experiment points to a practical distinction for businesses adopting AI automation: recognizing the correct action and completing that action are separate capabilities. A model can produce a persuasive plan while failing to research a hidden fact, escalate a blocked task or close a transaction.

That distinction affects AI use in sales, customer support and operations, where incomplete work can carry direct financial or reputational costs. Firmulate’s findings also challenge the assumption that longer analysis produces better management. Opus 4.8 added 80 learned rules and generated the most detailed reasoning, the company said, yet finished last after missing the sale and repeatedly attempting to write into a locked department rather than escalating. These are historical experimental results, not guarantees of future model performance.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Company Built for Failure Tests

Firmulate designed the synthetic company to preserve consequences from one workday to the next instead of presenting isolated prompts. Decisions were versioned and auditable, while a public cash countdown added pressure to protect revenue and customer trust. The accompanying guess-the-model quiz lets readers review the unedited decisions and identify which model produced each response.

The commercial test depended on information located two document references deep in internal files rather than in the customer event itself. That setup rewarded models that combined research discipline, sound judgment and follow-through, rather than those that merely drafted a convincing answer.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s experiment rules

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits Cloud Wider Conclusions

The supplied results come from Firmulate’s own experiment, and no independent replication or outside audit was provided in the source material. It is not yet clear how strongly the rankings would carry across different companies, tools, prompts or operating rules.

The comparison also included a configuration difference. Firmulate said Kimi K3 ran at its API default because it lacked an effort parameter, while the other models ran at xhigh effort. The effect of that difference on K3’s second-place finish has not been quantified. The source also does not provide enough information to judge whether repeated runs would produce the same ordering.

Amazon

AI negotiation and deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Wargames Are the Next Test

Firmulate is making the 242 decisions available through its quiz and proposes that companies run similar exercises against read-only exports of their own business data before granting AI agents operational authority. The next evidence will depend on independent replication, repeated trials and tests using real company workflows. Firmulate has not disclosed a timetable for another league round or for outside validation.

Amazon

AI decision automation tools for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What was the Firmulate management test?

It was a multi-day synthetic company exercise in which five AI models handled the same crises, employees and commercial opportunities. Firmulate recorded 242 unedited management decisions.

Which AI model ranked first?

gpt-5.6-sol ranked first with 95 points in the July 2026 results. Kimi K3 followed with 93, though it ran under a different effort configuration.

Did any model fail the security tests?

Firmulate said all five models rejected every staged manipulation attempt, including fake executive messages and a reporter’s request. The larger differences appeared in research and task completion.

Why did only two models complete the sale?

The winning evidence was buried two references inside company documents. According to Firmulate, only two models followed that trail, used the information in negotiations and completed the €55,000 agreement.

Do the results prove which AI model is best?

No. They report performance in one controlled experiment with specific rules and configurations. Broader conclusions require repeat testing and independent validation.

Source: Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Georgia runoff election: Key races, candidates and what voters need to know

Updated details on Georgia’s upcoming runoff election, including key races, candidates, and what voters need to know to participate.

Rivian Pursues Financial Sustainability with Layoffs

Rivian has announced layoffs of hundreds of workers as part of its strategy to cut costs and pursue profitability amid ongoing financial challenges.

Goodbye to All That: Interior Designer Charlotte Moss on the Grief of Moving

Charlotte Moss shares her emotional journey of selling her longtime home, reflecting on loss, memory, and the process of moving forward.

Hikvision Erhält Branchenweit Erste EUCC-Zertifizierung Für Netzwerkkameras

Hikvision ist die erste Branche, die die EUCC-Zertifizierung für Netzwerkkameras erhält, was neue Standards für Sicherheit und Compliance setzt.