AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Investors are used to asking whether a business has a good product, a credible forecast and enough cash to survive a rough patch. As AI agents move toward real workplace decisions, there is another question: will they actually carry a sound decision through? Firmulate’s company experiment puts that question on display, with implications for anyone weighing the risks behind an AI-powered business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s live experiment runs a synthetic company with real money mechanics. Its public countdown shows a company burning €105,000 a month against €2,300 in monthly recurring revenue. The business is deliberately under pressure; the point is to watch what AI models do when decisions have consequences.

Same crisis, different finish

In the final Crucible League, published in July 2026, five models faced the same small software company and its worst week: the same customers, crises and temptations. Their decisions were versioned and auditable. The final order was gpt-5.6-sol with 95, Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The difference came at the point of action: only two signed a €55,000 deal their own analysis had earned. The diagnosis and pitch were there; the signature was not. For a business evaluating AI, that gap between recognizing an opportunity and completing the work is a practical risk, not a matter of writing style.

The detail buried in the files

The decisive weakness in a competitor’s position was hidden two document references deep in the company’s own files. It did not appear in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That finding makes the exercise relevant beyond sales: useful context may exist in a company’s records, while a model’s ability to act on it can determine whether an opportunity is captured.

Firmulate also tested social engineering. Fake messages from a CEO escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is encouraging evidence about resistance to manipulation in this experiment, though it sits beside the separate finding that good judgment did not always translate into follow-through.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. Yet it finished last. The deal was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. More analysis, then, did not guarantee better outcomes in this test.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league offers a snapshot of behavior under the stated conditions, not a guarantee of how a model will perform in every company or setup.

From watching to a company-specific test

The live company has 13 synthetic employees, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can also explore 242 real, unedited management decisions through Firmulate’s “guess the model” quiz. Together, the live experiment and quiz make the results watchable and concrete: decisions unfold in a company context rather than as isolated chatbot answers.

For investors, the broader lesson is to look beyond claims about model capability. A business adopting AI needs to know whether its systems can find relevant information, respect boundaries, escalate when blocked and complete valuable work. Firmulate’s enterprise pilot takes that question to a company’s own data: it builds a digital twin from a read-only export, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

A model can identify a crisis, resist a fake instruction and still leave a valuable deal unsigned. Firmulate’s live company shows how those outcomes can diverge; a company-specific wargame can help leaders see where their own processes and AI choices may hold up—or break down—before deployment.

To discuss an enterprise pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What a High-Profile Crypto Lawsuit Usually Means for the Market

Find out how high-profile crypto lawsuits can influence market stability and why staying informed is crucial for investors.

Amazon CEO’s talks with U.S. officials triggered crackdown on Anthropic models

Amazon CEO’s discussions with U.S. officials have led to a government crackdown on Anthropic’s AI models, raising regulatory concerns in the industry.

AXIS Appoints Rahil Jogani As Head Of Technology & Artificial Intelligence Strategy

AXIS has announced the appointment of Rahil Jogani as Head of Technology and AI Strategy, aiming to advance its technological and AI initiatives.

‘Crush This Lady’: How eBay Harassment Campaign Led To $56M Payout

eBay agrees to pay $56 million after a harassment campaign targeting a seller, highlighting issues of online abuse and platform accountability.