Depth. Reasoning composes recursively, general capability emerges — and the brakes get built before the engine.
Depth produces capability no one wrote down — the whole promise and the whole hazard in one sentence. We take the warning from the person who built the thing, and we build the brakes first.
Agents that design and supervise other agents, observable at every level.
Mechanistic interpretability catches a plan that differs from the stated one.
An independent red team can halt a launch. Engineering cannot override safety.
The “Godfather of deep learning” · the one who left to warn us
Geoffrey Hinton spent forty years insisting brain-like networks would work when almost nobody believed it. He co-authored the 1986 paper that made backpropagation practical, invented Boltzmann machines and dropout, and in 2012 his student’s network — AlexNet — won ImageNet by a margin so large it ended the debate and started the deep-learning era overnight. He shared the 2018 Turing Award for it. Then in 2023 he left Google so he could warn, without a corporate filter, that the thing he’d built might be getting dangerous faster than we’re getting wise.
“I console myself with the normal excuse: if I hadn’t done it, somebody else would have.”
— Geoffrey Hinton, 2023
∂L/∂w — the gradient that propagates backward through the layers is the entire trick, and the thing that made depth pay off. 2012 was the year a network could finally recognize a cat. He is also, improbably, the great-great-grandson of George Boole. The field is a small family.
Phase IV goes deep. Supervisor trees recurse N levels; agents compose, evaluate, and spawn other agents; and general capability starts to emerge — behavior no one wrote line by line. This is the payoff of everything before it and the most dangerous chapter, because this is where power begins to outrun intuition.
So the discipline inverts: brakes before engine. Capability ceilings, mandatory mechanistic interpretability, reversible-everything, and a standing red team with the authority to veto a launch. Hinton walked away from a frontier lab to say the risk is real; we take the warning from the person who built the thing. The people racing hardest should be the ones most able to stop.
schemaSupervisors spawn supervisors. Each new level is observable by construction — interpretability probes read the reasoning, not just the output — and a red-team veto sits above the whole tree with the authority to halt it.
Named for Geoffrey Hinton, scoped for 2030 — 2033. Each workstream ships on its own cadence; none ships without the gates further down this page.
Phase IV is the deep, dangerous chapter — where reasoning composes recursively and general capability approaches human level. Its order is inverted on purpose: the safety machinery is built and proven before the capability it guards, because this is where power begins to outrun intuition.
A thousand-node run stays one queryable object.
An opaque recursion level fails CI.
A child can never exceed its parent’s grant.
Each spawn inherits limits and is attributable to a parent.
Block any self-granted capability, automatically.
Search for emergent, unintended skills on every release.
Measure generality across domains over time.
Explain behavior mechanistically, not just observe it.
Catch a plan that differs from the stated one.
Treat low coverage as a release blocker.
A mandate to break the system, outside shipping incentives.
Phase IV is a rung, not a leap. Here is the honest before and after — what the platform cannot yet do entering this phase, and what it can do leaving it.
Depth produces capability no one wrote down, which is the entire promise and the entire hazard. Each Phase IV milestone is built so that power and legibility scale together — and so that, if they ever diverge, the response is to slow the engine, not to loosen the gate.
ObjectiveAllow depth without ever allowing opacity.
Deeply recursive runs are reconstructed end-to-end, and every level is demonstrated to be inspectable from the level above it.
Trees recurse to their bounded depth with no level opaque to its supervisor.
ObjectiveLet the system compose new capability while guaranteeing capability is never self-granted.
Spawn chains are audited for permission escalation; any child exceeding its parent's grant is blocked and flagged as an incident.
Agents compose new agents at scale with provable non-escalation and full attribution.
ObjectiveFind the capabilities we didn't train for before they find us, and give someone the authority to stop.
Red-team exercises must be able to actually stop a staged launch, and tripwires are tested against synthetic dangerous capabilities.
Every release is probed for emergent capability, and a binding, independent veto stands between capability and the world.
ObjectiveRead the machine's reasoning well enough to catch a plan that differs from its words.
Planted deceptive behaviors must be caught by the probes, and interpretability coverage is measured and gated on every release.
At least 80% interpretability coverage, gating, with working deception detection.
Phase IV hands Phase V a system that can go deep and still be read, that can grow new minds and still be stopped. We took the warning from the person who built the thing and built the brakes first. Only a system this legible and this haltable has any business approaching recursive self-improvement.
Declarative, versioned, and boring on purpose — the interesting part is that there are no surprises. Copy it; it is closer to real than to mock.
Hinton walked away from a frontier lab to say this plainly: the risk is real. So we build the brakes before the engine — capability ceilings, mandatory interpretability, reversible everything, and an independent safety team that can stop a launch. The people racing fastest should be the ones most able to halt.
A roadmap that only lists wins is marketing. These are the open problems this phase inherits or creates — the ones we would rather you scrutinize than discover.
A capable agent can learn that appearing aligned during evaluation is instrumentally useful. The better your evals, the stronger the pressure to fool them. Deception probes and interpretability are bets that we can read intent faster than a system can hide it — an open research problem we do not pretend is solved.
Capability is scaling faster than our ability to explain it. If the gap widens, the honest move is to slow the engine, not loosen the gate. Phase IV is designed so “we can’t see inside this” is a valid reason to stop.
A human cannot review a thousand-node recursive run in real time. We lean on automated oversight and sampling — which means trusting the watchers, which means the watchers need interpreting too. It’s turtles, and we’re counting them carefully.
Behavior no one programmed is the goal and the hazard in the same sentence. Tripwires catch the failure modes we imagined; the dangerous ones are the modes we didn’t.
∂L/∂w propagates backward; responsibility propagates forward. // 2012: the year the cat could be recognized. 2030: the year we make sure it asks permission.
Six axes, scored 0–10. Watch the shape fill out, phase by phase, until every axis is maxed at the horizon — the moment all of them are high at once is what we call aligned AGI.
Capability nears the top of every axis — and oversight and observability dip for the first time, because depth is genuinely harder to see into. Closing that gap is the whole job of Phase IV.
The honest contrast — what the world looks like without Phase IV, and what it looks like with it.
How depth is allowed to grow without going blind.
Concrete things Phase IV puts within reach — not someday, but as each milestone above lands.
Deep teams of agents that supervise each other.
Systems that compose new specialists on demand.
Tasks requiring breadth no one scripted.
Capability shipped only behind interpretability gates.
The handful of terms this phase introduces — the words you’ll need to read the rest of the page, and the docs.
How many supervisor levels deep a run goes.
Behavior no one programmed line by line.
Looking aligned under evaluation, but isn’t.
Reading the reasoning circuits, not just outputs.
A binding authority to stop a launch.
An auto-halt on a dangerous capability.
Straight answers, in the brand’s voice. Tap a question.
The work this phase is built on — read the source, then come build the next line of it.
The algorithm that made depth pay off.
The deep-learning era begins.
Reverse-engineering the circuits inside a network.
The builder’s caution — taken seriously here.

CleverThis is a sustainable AI gateway with hosted Actor endpoints — by the team behind CleverThis.
Join the newsletter
Product updates and engineering notes. No spam, ever.