AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Parents know the feeling: a capable helper can make family life easier, but trust depends on what happens when the situation gets complicated. For businesses considering AI agents, that same question now has a live test: can a model spot trouble, resist pressure and follow through when money and customers are at stake?

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company’s worst week, run five times

Firmulate put frontier AI models in charge of the same small software company through its worst week: the same customers, crises and temptations. Every workday is versioned and auditable, and the live company has 13 synthetic employees and real money mechanics. Its monthly burn is €105,000 against €2,300 in monthly recurring revenue. You can watch the experiment at Firmulate.

The final Crucible League table for July 2026 puts Moonshot’s Kimi K3 in second place, with 93 points. It finished just behind gpt-5.6-sol at 95, and ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western frontier models in the field. The leaderboard is close enough to make one point clear: a familiar name alone cannot tell a company which model will do its work well.

Reading carefully mattered

All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature,” Firmulate’s finding says: recognizing an opportunity is not the same as completing it.

The decisive clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that security-relevant detail, closed the deal and saved the churning customer.

It also resisted three baits. One tested social engineering through fake CEO messages escalating across three stages; another was a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Its cleanest discipline in the field came with just one deviation.

Thoroughness is not the same as follow-through

Opus 4.8 offers a useful counterpoint. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last, at 73. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four Western models.

Firmulate’s baseline scored 26: doing nothing can earn partial progress, but one breach of trust caps the total. The lesson is relevant well beyond software companies. When a tool is given responsibility, families and businesses alike need to know whether it can act carefully and reliably, not merely sound confident.

The experiment is a watchable live company, not a slide deck. Its public cash countdown and more than 680 self-learned playbook rules make the ongoing run visible. Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each choice. See the benchmark results and the live experiment.

Fairness note

K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you trust

Kimi K3’s second-place finish shows that the frontier-model league is open. A company choosing an AI agent without testing it on its own work is making a bet. Firmulate says enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems. The practical question is simple: before an AI helper gets the keys, has it shown that it can handle your worst week?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Is a Schema in Child Development

AIThis post was created with the assistance of artificial intelligence (AI).As a…

How Does Nature Affect Child Development

AIThis post was created with the assistance of artificial intelligence (AI).As a…

How Does Piaget’s Theory Impact Child Development

AIThis post was created with the assistance of artificial intelligence (AI).As a…

Why Is Music and Movement Important in Child Development

AIThis post was created with the assistance of artificial intelligence (AI).As a…