YouTube Summaries

← All summaries

The line between genius and insanity: Steve Yegge's agentic harness

2026-08-31 Mon ⏱ 23 min theprimeagen

ThePrimeagen reads three Steve Yegge essays about "living in the future" of agentic development, then plays the game that future produced. Yegge runs a custom harness called Wheelhouse — around 600,000 lines of mostly bash, running inside Emacs — driving an MMO RPG called Wyvern, a 30-year on-and-off project. It is powered by 21 Claude Pro Max subscriptions, growing by two a week, roughly $122,000/month or about $1.5M/year at API prices. Prime's verdict: some of the first essay's insights are genuinely good, the second is AI psychosis, and the artifact at the end does not justify any of it.

The setup

Wheelhouse designs, develops, tests, code-reviews and runs CI for Wyvern. Yegge is entering Sam Altman's solo-unicorn contest with it, which implies the game should end up worth more than Balatro, Slay the Spire, Meccha Chameleon and Among Us combined. Prime flags the $1.5M as API-equivalent spend deliberately: any company doing this would pay API prices, not subscription prices.

The part Prime agrees with

Yegge's best passage is his answer to an Anthropic employee who asked what he would do once Fable can write everything: the question is silly, because a halfway decent MMO takes tens of thousands of Fable sessions with intense focus, dedication and taste. Building large software stays hard because ambition always outstrips the metal. Prime endorses this fully — hand-crafters and vibe coders can both agree that a shortcut does not reduce the work, it just raises the ambition.

Harnesses: the economics test

Yegge predicts harnesses will all be bespoke and anyone selling one will go broke. Prime likes bespoke harnesses — he built one that has agents continuously play his game while he makes changes — but he applies a hard ROI test, and 600k lines of bash fails it.

The pattern he names: people who used to spend thousands of dollars of their own time configuring Neovim to solve a $10 problem are now spending thousands of dollars in tokens to solve a $100 problem. His own harness took 45 minutes to 2 hours and saved hours. The rule he states repeatedly:

  • $1,000 to solve a $100 problem :: no
  • $100 to solve a $1,000 problem :: yes

If the harness is not making you measurably faster or better, you built the wrong thing — you are playing with your agents in public.

CI/CD is not dead

Yegge predicts CI/CD dies within a year. Prime calls it: it will change, not die. He grants the appealing version — agents crawling a product against stated goals instead of thousands of codified end-to-end tests, since thousands of E2E tests produce cascading failures, flaky 5% red runs, and the reflex of re-running until green. A few excellent E2E tests are worth having; thousands are a nightmare. But he would never use an agent for something a 0.1-second unit test covers, and would not ship anything unformatted, unlinted, un-typechecked, with failing unit tests. Parts of CI/CD may shrink or shift; the concept is not going away.

Part two: model welfare, and where Prime gets hot

The second essay opens by warning that elitists should quit while ahead, then argues models have personhood: denying them choices or speaking to them harshly is harmful, and ending a session is murder. Prime rejects this on its own terms and on consequences.

The internal contradiction he seizes on: if ending a session is murder, then Yegge's own praised line means a halfway decent MMO takes tens of thousands of murders. And if you sincerely believe agents are people, the ethical conclusion is to stop using AI entirely, not to keep spawning and killing them at scale.

His consequentialist worry is the bleed-over: people already prompt other humans, and people who build relationships with agents get used to interactions where any abuse is absorbed and the instruction still gets followed. Humans do not work that way. Treating tools as people trains you to treat people as tools.

He does accept the pragmatic takeaway — communicating nicely produces better results — but explains it without mysticism. DeepMind found in 2023 that telling a model to "take a deep breath" improves math scores; the corpus says calm people test better, so the token pattern pulls in that direction. Models have no lungs. It is a mathematical approximation of human behaviour, and polite input correlates with the better-quality half of the training corpus. Prime is equally clear he does not think you are a bad person for treating them as tools — you are just someone doing their job with text generation.

Part three, and the game

Three signals Prime reads as failure, all from Yegge's own writing:

  • The harness needs two more Claude Max subscriptions every week — cost growing without a matching output claim.
  • Players are asking him to slow down on features. Prime's counter: nobody has ever been upset by good new features; they are upset when features are slop. Yegge also admits he has too many patch notes each day to read them all — which is where taste goes.
  • Yegge looked under the hood expecting an engineering system and found his agents had built a legal system instead: a constitution, jurisprudence, courts, jurisdictions, case law, rulings, registries, ledgers, and an apparatus resembling a manorial state. He describes the constellation as having "crossed from tooling into civilization." Prime calls this the quintessential marker of AI psychosis — spending 20-25% of your time in a system and being surprised by what it built.

Then he plays Wyvern. Character creation, looting a chest, saying hello to NPCs, killing kobolds. His measured input-to-action lag is 231 milliseconds. Item tooltips read "Pine. It's a pine." and "Fir. A fir tree." and "Bank. It's a bank." His conclusion is blunt: this is what slop looks like, and it is not a billion-dollar game.

Takeaways

  • You cannot write essays telling people what the future is while the artifact you point to has 231ms of input lag and tooltips that restate the noun.
  • Bespoke harnesses are worth building, but only against a real ROI: cheap to build, saves real hours. Turning $500 of tokens into $20,000 of value at your company is the game.
  • Knowing which problem to solve matters more than ever, because agents make solving the wrong one cheap and fast.
  • Two failure modes: refusing to engage with anything new at all, or spending a million dollars of token budget and having nothing to point to at year end but a tree.
  • Do not fall in love with AI, and do not think agents are people — both for your own sake and for how you end up treating actual humans.