AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine if your garden’s growth depended not just on water and sunlight, but on an AI that can make managerial decisions—decisions that can cost thousands or even jeopardize your entire operation. How can you tell if an AI will act responsibly under pressure? The answer may lie in a recent experiment that pits top frontier AI models against real-world managerial crises, revealing which ones can truly handle the heat—and which falter when it matters most.

The Experiment: Putting AI Managers to the Test

In a groundbreaking live experiment, four leading AI models were tasked with running a small software company through its worst week—similar to managing a garden that faces unpredictable pests, weather, and resource shortages. The goal? To see if these models could navigate crises, resist manipulation, and complete profitable deals, just like a seasoned manager.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Were Tested

Every decision was real, every crisis authentic:

  • Same customers, same crises, same temptations—no shortcuts.
  • Decisions were fully versioned and auditable, ensuring transparency.
  • The models faced social engineering: staged fake CEO messages and a reporter trick, testing their integrity.

Watch the live company’s progress at firmulate.com/live and see how real business mechanics unfold on a daily basis.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Minds Under Pressure

All four models demonstrated awareness of crises and refused manipulation attempts. But a critical difference emerged in their ability to close a key deal:

  • The top performer, gpt-5.6-sol 95, identified a buried piece of information in the company’s files that clinched a €55,000 deal, earning a full monthly recurring revenue (+€4,583 MRR).
  • The second, Kimi K3 93, also signed the deal, exhibiting the cleanest discipline despite running without an effort parameter.
  • Two others, Sonnet 5 88 and Fable 5 77, missed the buried document and left the close on the table, losing a substantial opportunity.

Interestingly, all models refused social engineering attempts, understanding the risks of impersonation or approval bypasses. Kimi K3’s explicit reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI business decision simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Information Depth Matters

The decisive factor lay in how deeply each model read into the company’s own files. The winner had uncovered critical information two document references deep—an insight that turned the tide in their favor. Models that read more thoroughly had a tangible advantage, translating into real revenue gains.

Amazon

AI deal-closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management Personalities in AI

This experiment reveals that AI models exhibit measurable management styles:

  • The Opus 4.8 profile was the most thorough, analyzing over 80 rules and offering deep insights, but ultimately performed the worst—closing the deal late and slipping discipline.
  • The K3 model maintained fairness by running at default effort level, demonstrating discipline and a focus on integrity.
  • The other models, despite their high scores, showed weaknesses in leveraging information or maintaining discipline under pressure.

Such personality profiles are not just academic—they could guide how AI is integrated into management roles, especially in sensitive or high-stakes environments.

Why Gardeners Should Care

While this experiment is rooted in software management, the implications extend far beyond. If AI tools are to assist in managing your greenhouse, outdoor projects, or landscape business, understanding whether they can finish what they start—reading your files thoroughly, resisting manipulation, and staying honest—is crucial. The question isn’t just about how well they write or communicate but whether they can truly be trusted to act responsibly when the pressure’s on.

How to Test Your Own AI Workforce

Interested in seeing how your AI tools would perform? Firmulate offers a live platform where enterprises can run their own management wargames against a read-only export of their business. It’s a safe, visible way to gauge AI reliability—without risking real operations. Learn more at firmulate.com/pilot.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How To Get A Second Flush Of Flowers From Your Daylilies – Plus, This Easy Trick Keeps Plants Looking Fresh

Learn a simple method to encourage a second flush of flowers from your daylilies and keep your garden vibrant. Expert tips included.

Is it time to stop watering our gardens?

As drought concerns grow, experts suggest reducing garden watering through drought-tolerant plants and water harvesting. Is it time to change habits?

August Is Your Last Chance To Prune These 6 Plants Before It’s Too Late For Next Year

August is the final month to prune six specific plants to ensure healthy growth next year. Experts advise timely trimming to avoid future issues.

Built a place for his bird to flee again!

A man creates a safe refuge for his pet bird to escape again, highlighting efforts to prioritize animal freedom and welfare.