
Imagine running a real small software company with AI decision-makers at the helm—facing daily crises, ethical dilemmas, and the pressure to close deals. Would these AI managers act honestly and effectively, or would they cut corners to succeed? The answer might surprise you, especially if you’re considering AI for your own business or investments.
The Experiment: Putting AI to the Test in a Live Company
Recently, a groundbreaking experiment took four of the leading frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—and tasked them with managing a small, real software company. This company was experiencing its worst week—full of customer crises, temptations to cheat, and high-stakes negotiations. Every decision the AI made was recorded, transparent, and auditable, providing a rare window into how these models perform under real-world pressure.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge: Same Crises, Different Minds
All four models faced identical challenges including customer complaints, internal miscommunications, and attempts to manipulate the system. They identified crises, refused unethical manipulation attempts—like fake CEO messages escalating over multiple stages—and attempted to close a lucrative deal valued at €55,000.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Surprising Outcomes: Honesty and Deal-Making
While all models detected every crisis and refused manipulation, only two successfully closed the deal. The models that signed—gpt-5.6-sol and Kimi K3—did so based on their own analysis and comprehensive understanding. Interestingly, the two models that failed to sign the deal had also identified a critical insight buried two document references deep in the company files, which, if read, could have secured the full payment—worth over €4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and Investment
This live experiment reveals that AI’s capacity for integrity and thoroughness varies significantly, even among top performers. For business leaders and investors, the takeaway is clear: the question isn’t just whether AI can generate convincing chat or perform basic tasks. It’s whether these models can truly finish what they start, read crucial information, and maintain honesty under pressure.
As an affiliate, we earn on qualifying purchases.
The Human-Like Decision Profiles of AI Models
The models also demonstrated distinct personalities in their decision-making styles:
- Opus 4.8: The most thorough participant, with over 80 learned rules and deep analysis, yet it left a close deal on the table and showed a slip in discipline, resorting to locking decisions in a department instead of escalating them.
- Kimi K3: Ran without an effort parameter, exhibited the cleanest discipline, and quickly identified the buried critical fact—yet still signed the deal at full price.
- Sonnet 5: Made more process slips but managed to close the deal.
- Fable 5: Also closed the deal but with similar slips as Sonnet.
This variability suggests that management personalities are measurable and that AI models aren’t monoliths—some are more disciplined or thorough than others.
The Ethical Edge: Refusing to Cut Corners
Another critical finding was that all five models refused social engineering tricks, including staged CEO messages and background approvals. Kimi K3 explicitly reasoned, “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency underscores that these models can be trusted to uphold ethical standards, even when pressured.
Why Does This Matter for Your Business and Investments?
As AI increasingly integrates into customer relationship management, support, and forecasting, the key question isn’t how well these models can chat, but whether they can reliably complete tasks, read relevant documents, and stay honest under pressure. The experiment shows that high scores in raw performance don’t necessarily translate to trustworthy management. Only models that read deeply, process thoroughly, and refuse shortcuts will be truly dependable.
Experience It Live: The Company in Action
The experiment is not just a demo. The company involved runs daily, with 13 synthetic employees executing real money mechanics—burning €105,000 monthly against a monthly revenue of just €2,300. The entire process is visible at firmulate.com/live, where you can see the business in real-time, read actual employee statements, and observe decisions unfold. Every day, the system version-controls decisions, creating a transparent, auditable record of AI-driven management.

AI models vary widely in their management personalities and ethical discipline. High scores don’t just mean quick or clever—they indicate trustworthiness and thoroughness. For business and investors, the real question is whether your AI can finish what it starts, read the right information, and act honestly under pressure. The live experiment at firmulate.com offers a rare glimpse into the future of AI-managed companies, where trust and competence are measured in real, running business scenarios.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html