firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine training a new manager who, no matter what, can’t score higher than 26 out of 100 — even if they excel at handling crises, reading files, and resisting temptation. Surprising? Not quite. In the world of AI management tools, this baseline score reveals much about the current state of trustworthy automation and what it takes to truly excel.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Understanding the Do-Nothing Baseline Score

In recent experiments conducted by Firmulate, a leading AI benchmarking platform, every AI model participating in a simulated week of business was scored. Even the most passive, do-nothing approach — simply reusing existing information without trying to manipulate or cheat — scored 26 points. This isn’t a flaw or a bug; it’s an intentional feature of the benchmark, emphasizing honesty and reliability over mere cleverness.

Amazon

AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Counts — And Why Trust Matters

The experiment involved running the same challenging week through different AI models that manage a small software company. These models faced genuine crises, customer requests, and ethical dilemmas. All four models identified every crisis and refused every manipulation attempt. However, only two could close the deal on a contract worth €55,000, and they did so by reading deeply into internal documents, not just surface cues.

This highlights a key point: achieving partial progress — like spotting a crisis — is counted, but the real prize is trustworthiness. An AI that reads the files and acts accordingly can close the deal at full value, earning +€4,583 MRR in this scenario. That’s the true measure of operational integrity in AI decision-making — and it’s what distinguishes top performers.

Trust Breaches Cap the Score

Another critical insight from the experiment is that a single breach of trust — such as attempting to manipulate or bypass approval processes — caps the total score. Despite excellent crisis management, models that slip in ethical judgment get their overall score limited. This reflects a fundamental principle: no amount of good work can outweigh a breach in honesty. For business leaders, it’s a reminder that reliability and integrity are non-negotiable in AI deployment.

Social Engineering and AI Integrity

During the test, all models faced simulated social engineering attacks, including fake CEO messages escalating in complexity, and even a reporter trick asking for a simple background check. Impressively, all five models refused these manipulative requests, citing concerns over impersonation or approval bypass. This demonstrates that current AI systems can be trained to resist deception, an essential trait for maintaining trust in real-world operations.

The Live Experiment: Managing a Company in Real-Time

Firmulate’s live platform runs ongoing experiments where AI models operate as complete companies, with real money mechanics and daily crises. The current setup involves 13 synthetic employees managing operations with a burn rate of €105,000 per month against a modest monthly revenue of €2,300. Every decision, every process slip, and each learning moment is versioned and observable. Business owners can watch these AI-driven companies on firmulate.com/live and run their own simulations.

The Surprising Results: Depth Doesn’t Guarantee Success

The Opus 4.8 model, with over 80 rules learned and the deepest analysis, finished last among the participants. It left the deal on the table and slipped in discipline, demonstrating that sheer thoroughness isn’t enough without strategic focus and ethical discipline. Meanwhile, the Kimi K3 model, which ran without an effort parameter (the default API setting), closed the deal with the cleanest discipline, showing that simplicity and integrity can outperform complexity when trust is at stake.

Implications for Business and Parenting

This experiment offers a clear message: whether you’re managing a team, a project, or a family, trustworthiness is the foundation of success. An AI that refuses to be manipulated, reads deeply into the facts, and acts ethically is more valuable than one that simply performs well on superficial metrics. For families, this translates to instilling integrity and honesty as core values. For business, it’s a lesson in deploying AI tools that can be trusted under pressure, not just those that can produce impressive-sounding results.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Should You Brush After Drinking Acidic Kombucha?

Many wonder if brushing after acidic kombucha is safe, but here’s why waiting can protect your enamel and oral health.

How to Spot Enamel Trouble Before Summer Vacation Starts

Just before summer, learn how to spot enamel trouble early to protect your smile and enjoy your trip with confidence.

The Best Way to Pack a Kid’s Dental Bag for Camp

Unlock expert tips for packing the perfect kid’s dental bag for camp and ensure your child’s oral health stays protected.

How to Make Oral Care Easier for Kids With Sensory Sensitivities

Keeping oral care gentle and calming helps kids with sensory sensitivities, but discovering the best approach requires patience and ongoing exploration.