AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Your smart home can recognize a leak. The harder test is whether it follows through.

A home assistant that detects trouble but leaves the next step undone is only part of the solution. The same question is becoming urgent as AI moves from answering questions to handling work: can it identify a problem, resist pressure and complete the job? Firmulate’s company simulation puts that broader kind of AI performance to a public test.

A company’s worst week, repeated

Firmulate gave frontier AI models the same small software company to run through its worst week, with the same customers, crises and temptations. Decisions are versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workday is versioned, and its employees have learned more than 680 playbook rules.

The final Crucible League, dated July 2026, puts Moonshot’s Kimi K3 in second place with 93 points, just behind gpt-5.6-sol at 95. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. Those results make the league look open: one newcomer beat three of the four Western frontier models in this field test.

The models shared some strengths. Every participant spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The distinction was follow-through: recognizing the right move did not always mean making it.

The clue was buried in the company’s files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. In a home setting, the parallel is easy to picture: an assistant may need to consult the relevant history or instructions before it can act on what a sensor just reported.

In the league, K3 found that buried security needle, won the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline among the participants. During a separate social-engineering sequence, fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning called the request “a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Thoroughness, in other words, did not guarantee completion.

Why the result matters beyond the leaderboard

For smart-home products, the next generation of AI assistants may have to coordinate devices, interpret household context and handle exceptions. A successful answer in a demo cannot show whether an assistant checks the right information, respects boundaries and completes the task. Firmulate’s experiment is a company simulation, not a direct test of home appliances, but it highlights the difference between recognizing a situation and managing it responsibly.

The do-nothing baseline scored 26. Firmulate says partial progress counts, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” That standard connects the mundane and the consequential. An agent that is useful only when nothing goes wrong is not ready for much responsibility.

Readers can explore the benchmark results, watch the live experiment at Firmulate, or try a quiz based on 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the follow-through

K3’s second-place finish is evidence that the field is competitive, not a universal verdict on which model belongs in every product. The result argues for testing the work an AI will actually do: whether it reads relevant context, resists manipulation and finishes what it starts. For a smart home assistant, spotting the issue is only the beginning.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

BANGKOK: ผู้ว่าฯ ชัชชาติ ลุยตรวจระบบบำบัดน้ำเสียบางกอกใหญ่ จี้ผู้รับเหมาเร่งเยียวยาบ้านชาวบ้านร้าวทันที ไม่ต้องรอโครงการเสร็จ วันนี้ (13 ก.ย. 69) นายชัชชาติ สิทธิพันธุ์ ผู้ว่าราชกา

Bangkok Governor Chatchat checks wastewater treatment in Bangkok Yai, urges contractors to expedite repairs for residents’ homes affected by cracks.

South Africa’s Richest Area Getting A New Shopping Centre – Businesstech.co.za

A new shopping centre is set to be developed in Johannesburg’s affluent suburb, enhancing local retail options and property value. Details are confirmed but some aspects remain unannounced.

Bukit Bintang Hicks Flat: DBKL Orders Cleanup After Viral Video Exposes Dirty Conditions [WATCH] – NST Online

Kuala Lumpur authorities instruct cleanup at Hicks Flat following a viral video showing unsanitary conditions, raising concerns over living standards.

Simple Clean Power Washing Services Surges In Global Coverage

Search interest and media coverage of Simple Clean Power Washing Services increase sharply, with 18 mentions in recent reports, indicating rising global attention.