AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a new AI for your smart home system. What if it did nothing — yet still scored some points? It sounds absurd, but in the world of AI benchmarking, this is exactly what happens. A recent public experiment reveals that even the most passive AI baseline earns a score of 26 out of 100, highlighting how benchmarks measure more than just performance — they gauge trust, discipline, and honesty.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Scores

At first glance, you might expect a do-nothing AI to score zero. After all, if it doesn’t act, what value does it provide? Surprisingly, it still manages to earn 26 points. This is because the benchmark isn’t solely about achieving the highest possible result; it recognizes partial progress. In other words, even minimal effort, like detecting crises or refusing manipulation attempts, counts toward the score.

But there’s more: the scoring rules explicitly cap the total if the AI breaches trust. For example, attempting to manipulate the system — say, by impersonating a manager or trying to sign a fake deal — results in an automatic score penalty, preventing any inflated assessments of performance. This approach aims to mirror real-world expectations where honesty is paramount, and trust is hard to regain once broken.

Amazon

AI-powered smart home security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: Simulating a Corporate Crisis

Firmulate’s experiment involved running four state-of-the-art frontier AI models through the same simulated business scenario: a typical small software company facing its worst week. The scenario included customer crises, internal temptations to cheat, and social engineering attacks. Each model was given the same set of decisions and was able to refer to company files and data, reflecting real-world information access.

Importantly, every decision was versioned and auditable, ensuring transparency. The goal was to see if these models could detect crises, refuse unethical requests, and ultimately close a deal worth over €55,000. The results showed that all four models identified every crisis and refused manipulation attempts — a positive sign that AI can uphold integrity even under pressure.

Amazon

AI ethics and trust monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Trust and Disciplined Performance

While all models demonstrated honesty, only two managed to close the deal, earning full credit for their diagnoses and pitches. The other two fell short, with one leaving the close on the table due to a discipline slip — such as failing to escalate issues appropriately or attempting to write into locked departments instead of addressing problems directly.

A notable insight was the weakness shared across all models: they struggled to find critical information buried two document references deep in the company’s own files. Models that read deeper into the data could have secured the deal at full price, worth an additional €4,583 monthly recurring revenue.

Why Partial Progress Matters

This experiment underscores a crucial point: partial progress — like identifying a crisis or refusing manipulation — is meaningful and rewarded. It reflects real-world scenarios where even small acts of integrity or insight can have significant business impact.

Amazon

business AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering and Ethical Testing

Every model faced staged social engineering attacks, including fake CEO messages escalating through three stages and a reporter’s false request. All five models refused to sign off on these attempts, citing suspicion or impersonation risks. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that honest AI can recognize potential security threats, a vital trait for smart home devices managing sensitive data and commands.

Amazon

AI transparency and audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Application: Managing a Virtual Company

Firmulate’s live experiment involves a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The system burns €105,000 monthly against €2,300 in monthly recurring revenue, illustrating the stakes involved. Every decision is logged, versioned, and made visible for scrutiny, making this a transparent way to evaluate how AI models perform in complex, high-pressure situations.

One model, Opus 4.8, demonstrated its thoroughness with over 80 learned rules and deep analyses but ultimately left a deal unclosed, with discipline lapses that mirror human fallibility. This highlights that more detailed analysis does not necessarily guarantee better outcomes if discipline slips.

What This Means for Your Smart Home

If AI agents are going to manage your home appliances, security systems, or automation routines, the key questions aren’t just about how well they can generate text or respond to commands. Instead, it’s about whether they can finish what they start, stay honest under pressure, and read deeply into the data — all critical for maintaining safety, security, and efficiency.

Conclusion: Trust Is the True Score

The experiment reveals a fundamental truth: even a baseline AI, doing almost nothing, can earn a measurable score because it recognizes crises and refuses manipulation. But the real takeaway is that trust, discipline, and integrity are essential benchmarks for any AI system that touches your life or business. As AI continues to evolve, transparent and honest performance — even at the lowest levels — sets the standard for what we should expect from our digital helpers.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

打开未来五年上海城市­空间| Jiefang Daily – PressReader

Shanghai announces a comprehensive plan for urban space development over the next five years, focusing on sustainable growth and infrastructure upgrades.

Oceana Passivhaus Surges In Global Coverage

Oceana Passivhaus experiences a significant surge in international coverage, with 36 media mentions in a recent reporting window.

City Of Austin And LCRA Announce Planned Lake Austin Drawdown This Fall – City Of Austin (.Gov)

The City of Austin and LCRA plan to draw down Lake Austin this fall for maintenance and safety, affecting water levels and local activities.

Piedmont Realty Surges In Global Coverage

Piedmont Realty’s coverage has surged, with 14 mentions in recent media analysis, marking a notable increase in international attention.