
Imagine a business that operates entirely without employees, battles daily cash shortages, and still manages to make critical decisions under pressure—all in front of your eyes. This isn’t fiction; it’s the live experiment from Firmulate, where artificial intelligence models are running a company in real time, with every move observable and auditable.
At first glance, the scene seems impossible: a company burning €105,000 every month against a revenue of just €2,300, with its entire operation publicly documented at firmulate.com/live.html. This open display of an AI-driven enterprise is not only a spectacle of technological feat but also a stark look into what the future of management might hold—where decisions are made by algorithms, not humans.
The Experiment: Running AI Against a Crisis-Heavy Week
The setup is straightforward but intense: four state-of-the-art AI models, each representing a different decision-making approach, are challenged to steer a small software company through its worst week. They face the same customers, same crises, and same temptations—like manipulation attempts and social engineering tactics—making it a fair test of each model’s integrity and judgment.
Scores and Performance
- GPT-5.6-sol scored 95 out of 100, successfully uncovering a hidden document that led to closing a €55,000 deal, increasing monthly recurring revenue (MRR) by €4,583.
- Kimi K3 scored 93, also closing the deal and showing the cleanest disciplinary record, refusing every manipulation and suspicious request.
- Sonnet 5 scored 88, closing the deal but with some process slips.
- Fable 5 scored 77, demonstrating good rule discipline but ultimately leaving the deal unexecuted due to procedural lapses.
Quite revealing: despite all passing the crisis detection, only two models signed the deal, even though they arrived at similar diagnoses and pitches. The critical weakness wasn’t in spotting the crisis but in acting decisively—something that’s buried deep in the company’s own files, not in the superficial customer interactions.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the AI Reveals About Business Integrity
One of the most telling findings is that the models which read and analyze internal company documents discovered the hidden opportunity that others missed. This buried information was essential, and in the real world, such overlooked details could mean millions in missed revenue.
Another fascinating aspect is how the models responded to social engineering. Fake messages from a supposed CEO asking for quick approvals or background checks were refused by all five models tested, with Kimi K3 explicitly treating such requests as potential impersonation attempts. This indicates an increasing level of sophistication in AI’s ability to resist manipulative tactics, a vital trait for future management tools.
The Real-World Mechanics: The Company in Action
Behind the scenes, the live setup involves 13 synthetic employees making decisions, each governed by a set of over 680 learned rules, updated daily in versioned runs. The company’s financials are stark: it continually burns cash at a rate of €105,000/month, with only €2,300 in monthly revenue to sustain it. The public display offers a view into how AI might one day run critical parts of a business—handling crises, negotiating deals, and maintaining discipline—all in real time.
Deepest Participant and Lessons Learned
The most thorough AI participant, Opus 4.8, analyzed over 80 rules and performed the deepest analysis, yet it finished last in securing the deal. Its discipline slipped, and it failed to escalate crucial issues into the right departments. This underscores that even with sophisticated analysis, execution discipline remains a challenge—pointing to the importance of process consistency in automated decision systems.
Implications for the Automotive Industry
While this experiment is set in a tech company, the lessons resonate profoundly in automotive and garage management. As dealerships and repair shops increasingly adopt AI for customer relations, inventory, and workflow management, the questions are clear: will these systems truly complete their tasks under pressure? Will they read critical internal data before acting? And most importantly, will they maintain honesty and discipline when temptation or manipulation arises?
The experiment shows that AI can detect crises and stay honest against social engineering attempts. However, successful closure of deals and decision execution still depends heavily on how well these systems are designed and disciplined.
The Future of AI in Business Decision-Making
This live, transparent experiment from Firmulate provides a glimpse into a future where AI models are not just chatbots but full-fledged decision-makers. It highlights the importance of transparency, versioned decision-making, and rigorous testing—especially in high-stakes environments like automotive service management.
For automotive entrepreneurs and managers, the takeaway is this: before you entrust your systems to AI, you need to test and simulate how they handle crises, manipulations, and internal data. The live platform at firmulate.com/live.html offers a rare opportunity to see these processes unfold and to evaluate whether your future AI workforce can truly deliver on its promises.

Watch AI battle real business crises in real time at Firmulate’s live experiment. The test shows AI’s potential and limits—crucial insights for industries like automotive management aiming for trustworthy automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html