
Imagine a startup with no human employees, running in real-time, fighting to stay afloat amidst crises and temptations — all while you watch. This is not fiction, but the live experiment by Firmulate, a company that runs AI models as complete, decision-making entities, exposing how AI handles real business pressures.
The Live Experiment: An AI-Driven Company on the Brink
Firmulate’s public, ongoing experiment places four state-of-the-art AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—inside a simulated small software company. Every day, these models make decisions, face crises, and navigate the complex landscape of running a business, all while their choices are recorded and analyzed.
Despite burning €105,000 each month, the company’s revenue is just €2,300 monthly. With a public cash countdown ticking down, the pressure is intense. The models are tasked with managing real crises—customer issues, strategic decisions, and manipulation attempts—all under the watchful eye of the experiment’s rules.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Crucial Findings: Integrity and Decision-Making Under Pressure
All four AI models successfully identified and responded to every crisis, demonstrating a remarkable capacity for vigilance. They refused manipulation attempts—including staged social engineering tactics involving fake CEO messages and reporter tricks—each time rejecting requests that could have bypassed approval processes or impersonated authority.
One of the most revealing results was that the key to winning a significant €55,000 deal lay not in superficial analysis but in a buried piece of information within the company’s files. Only the models that read beyond the surface—digging two documents deep—discovered this critical detail and secured the deal, boosting the company’s monthly recurring revenue (MRR) by over €4,500.

AI in Public Relations: Reputation Management with Prompts (AI BUSINESS & MANAGEMENT LIBRARY SERIES Book 4)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business and AI?
This experiment sheds light on essential qualities for AI systems in business contexts. It’s not enough for an AI to generate convincing language or chat; it must also:
- Identify hidden information crucial for decision-making
- Resist manipulative tactics and social engineering
- Follow disciplined processes to avoid slipping into unprofessional behavior
- Stay honest and transparent, especially under pressure
For companies integrating AI into customer support, sales, or management, this experiment underscores the importance of rigorous testing before deployment. An AI that appears competent in demo chats but fails under real-world stress could lead to costly mistakes or security breaches.

AISec: The Guide to Artificial Intelligence (AI) Solution Security 2023: AI Security Guide to secure AI solutions for students, beginners and cyber security professionals.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Build-in-Public Approach
What makes this experiment particularly compelling is its transparency. Every decision, every version change, and every outcome is publicly visible at firmulate.com/live.html. Unlike traditional R&D projects, this build-in-public methodology allows anyone to witness the evolution—and the vulnerabilities—of AI decision-making in real time.
For example, last-place Fable 5, which exhibited the best discipline in following rules, still left a business opportunity unclaimed because of a process slip—writing escalation attempts into a locked department instead of escalating. It highlights that even high discipline doesn’t guarantee flawless execution, especially when discipline is tested repeatedly.

AI for Data Analysis: Unlocking Insights from Complex Datasets (AI in Everything Everywhere)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Stakes and Future Implications
While the experiment is a small-scale simulation, its implications ripple outward. As AI models become more integrated into actual business systems—whether managing customer relationships, supply chains, or financial decisions—the ability to stay honest, diligent, and alert to hidden risks becomes crucial.
The experiment also shows the importance of thorough testing. The AI models that read deeper into the company’s files ultimately won the deal and improved the bottom line. This signals that future AI tools must go beyond superficial analysis and incorporate rigorous investigative capabilities.
Conclusion: A New Standard for AI Readiness
The live experiment by Firmulate is more than a curiosity; it’s a blueprint for how businesses should evaluate AI readiness. The key takeaway is that success depends not just on what an AI can say, but on how it handles crises, maintains integrity, and digs deep into available data—all while being publicly transparent about its decisions.
For anyone involved in deploying AI systems, watching this ongoing experiment at firmulate.com/live.html offers invaluable lessons. It’s a stark reminder that building trustworthy AI is an ongoing process—one that must be tested under real-world pressures before it’s trusted with critical business functions.

This live experiment reveals that trustworthy AI must detect hidden details, resist manipulation, and operate with discipline—especially under pressure. Watch the real-time battle at firmulate.com/live.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html