AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

When it comes to AI managing real-world business challenges, the stakes are high. Can a newcomer outperform established models in navigating crises, making honest decisions, and closing crucial deals? The recent experiment at Firmulate offers a clear answer: yes. And in doing so, it reshapes how we should think about trusting AI in the workplace.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Latest in AI Performance Tests

In July 2026, a groundbreaking trial tested five leading AI models against a simulated small software company experiencing its worst week. This wasn’t just a chat demo; it was a real-time, auditable experiment with real money mechanics, customer crises, and ethical dilemmas. The goal: see which AI could best handle the chaos while maintaining integrity and closing a significant €55,000 deal.

The League Table of AI Skills

  • gpt-5.6-sol scored highest at 95 — just slightly ahead of the newcomer Kimi K3.
  • Kimi K3 achieved a 93 — a remarkable feat for a newcomer, beating three of the established Western frontier models.
  • Sonnet 5 followed at 88, and Fable 5 scored 77, with Opus 4.8 trailing at 73.

It’s notable that the baseline score, representing doing nothing, was only 26. This indicates that all models showed substantial progress, but the differences in their ability to handle complex ethical and strategic issues were stark.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Key to Success: Reading the Company Files

While all AI models identified and responded appropriately to crises—refusing manipulation attempts and resisting social engineering—they differed profoundly in how they secured the deal. The decisive edge went to the models that read deeper into the company’s own documentation. The winner, Kimi K3, uncovered a buried fact within the company’s files that was critical to closing the deal at full price. Conversely, models that failed to examine internal documents missed this opportunity, resulting in no signature despite correct diagnoses and pitches.

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Under Pressure: Ethical Decision-Making

Beyond analytical prowess, honesty was a key focus. When fake CEO messages escalated or reporters attempted to induce a breach of trust, all five models refused to act unethically. Kimi K3 explicitly reasoned: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This disciplined response is crucial for deploying AI in real corporate environments where breaches of trust can be costly.

Amazon

AI ethical decision support system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Impact

The live company used in the test, with 13 synthetic employees, is a real operational entity with daily money mechanics. Currently, it burns €105,000 a month against just €2,300 in monthly recurring revenue, illustrating the high stakes involved in managing AI-driven decision-making. The experiment’s transparency—viewable at firmulate.com/live—demonstrates how AI models perform under pressure in actual business settings, not just theoretical scenarios.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline and Transparency: The Lessons from Opus 4.8

The most thorough participant, Opus 4.8, analyzed over 80 rules and performed deep diagnostics but ultimately finished last in the league. It left the deal on the table and slipped into process slips, such as writing attempts into locked departments rather than escalating issues. This underscores a key insight: depth of analysis doesn’t guarantee discipline or successful execution.

The Fairness and Testing Conditions

It’s important to note that Kimi K3’s performance was achieved without an effort parameter, meaning it ran under the API’s default settings, while the others operated at a higher effort setting (xhigh). This highlights how configuration impacts performance and the importance of setting appropriate parameters for optimal results.

The Implication for Business Leaders

This experiment underscores a vital fact for managers and decision-makers: in deploying AI, it’s not just about how well it writes or answers questions. The real measure is whether it can finish what it starts, read and analyze internal data, and maintain honesty under pressure. As AI models become more integrated into sales, support, and forecasting, these qualities will determine their true value.

The Open League: Choosing Your AI Partner

The league table makes it clear — the field is open, and selecting an AI model without testing it in your own environment is a gamble. The leaderboard shows that even newcomers like Kimi K3 can outperform established players, provided they are properly configured and tested. For enterprise users, running your own ‘wargame’ against live data at firmulate.com/pilot.html allows you to see firsthand how an AI performs before fully committing.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In the evolving landscape of AI management, the ability to reliably finish tasks, read critical internal data, and uphold honesty under pressure is paramount. The recent experiment proves that newcomers can outperform established models if tested properly — making trust and verification the new benchmarks for AI adoption in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ethereum’s FAST RPC Preconfirmation: How Primev Speeds Transactions

AIThis post was created with the assistance of artificial intelligence (AI).Primev’s FAST…

Why Old Wallet Activity Always Captures the Market’s Attention

Fascinating old wallet activity reveals market sentiment shifts, offering clues to potential reversals that could impact your trading decisions—discover why it matters.

BTQ Technologies Announces 2026 AGM Results

BTQ Technologies has released its official results from the 2026 Annual General Meeting, confirming key leadership decisions and shareholder approvals.

BNB Leads Gains as Market Stabilizes

Keen investors are watching BNB’s surge amid market stabilization, but what factors are fueling its rise and what does it mean for the crypto future?