Roadmapchevron_rightPhase IV · Depth
arrow_backPhase IIIarrow_forwardPhase V
IV DEPTHPhase IV2030 — 2033

Go deep — carefully.

Recursive hierarchies run many levels deep and general capability starts to emerge — so the brakes get built before the engine.

personNamed for Geoffrey Hinton· 1947 —
brush render: ridgeline relief

infoRendered as a ridgeline relief — every horizontal scan lifted by brightness, the face read as pure depth. The machine now perceives in three dimensions.

The brief

Depth. Reasoning composes recursively, general capability emerges — and the brakes get built before the engine.

phase-04.spec
status◇ FUTURE — the dangerous chapter
named forGeoffrey Hinton · 1947–
eraapproaching human-level breadth
deliversrecursive trees, emergent capability, interpretability
unlocksaligned self-improvement (Phase V)
ruleif we can’t read it, we don’t ship it
scope6 milestones · 11 tasks
N-deep
recursion
observable at every level
≥ 80%
interpretability
gating, per release
binding
red-team veto
safety can stop a launch
human-level
breadth
kept legible & haltable
The stakes

This is where power begins to outrun intuition.

Depth produces capability no one wrote down — the whole promise and the whole hazard in one sentence. We take the warning from the person who built the thing, and we build the brakes first.

The road to AGIPhase IV of V · 80% there
I Foundational
II Signal
III Perception
IV Depth
V Singularity
account_tree

Minds that compose minds

Agents that design and supervise other agents, observable at every level.

visibility

Read the reasoning, not the output

Mechanistic interpretability catches a plan that differs from the stated one.

block

A veto that actually stops

An independent red team can halt a launch. Engineering cannot override safety.

The mind behind the phase

Geoffrey Hinton

The “Godfather of deep learning” · the one who left to warn us

Geoffrey Hinton spent forty years insisting brain-like networks would work when almost nobody believed it. He co-authored the 1986 paper that made backpropagation practical, invented Boltzmann machines and dropout, and in 2012 his student’s network — AlexNet — won ImageNet by a margin so large it ended the debate and started the deep-learning era overnight. He shared the 2018 Turing Award for it. Then in 2023 he left Google so he could warn, without a corporate filter, that the thing he’d built might be getting dangerous faster than we’re getting wise.

I console myself with the normal excuse: if I hadn’t done it, somebody else would have.

Geoffrey Hinton, 2023
lightbulb

∂L/∂w — the gradient that propagates backward through the layers is the entire trick, and the thing that made depth pay off. 2012 was the year a network could finally recognize a cat. He is also, improbably, the great-great-grandson of George Boole. The field is a small family.

Selected timeline
1986
Backpropagation made practical (with Rumelhart & Williams)
2006
Deep belief nets reignite neural networks
2012
AlexNet wins ImageNet — the deep-learning era begins
2018
Turing Award (with Bengio & LeCun)
2023
Leaves Google to speak freely about AI risk
Phase 04 · what we’re building

Backpropagation & deep learning

lan
Backpropagation & deep learning
Depth that produces capability nobody explicitly wrote down — and a creator who later left his lab to warn the world about it.

Phase IV goes deep. Supervisor trees recurse N levels; agents compose, evaluate, and spawn other agents; and general capability starts to emerge — behavior no one wrote line by line. This is the payoff of everything before it and the most dangerous chapter, because this is where power begins to outrun intuition.

So the discipline inverts: brakes before engine. Capability ceilings, mandatory mechanistic interpretability, reversible-everything, and a standing red team with the authority to veto a launch. Hinton walked away from a frontier lab to say the risk is real; we take the warning from the person who built the thing. The people racing hardest should be the ones most able to stop.

⛔ red-team veto — bindinghalt + pagerootsupervisorsupervisoragentagentagentagentagents spawn agents — children inherit limits, never exceed themprobeinterpretability reads the reasoning, not just the output — ≥ 80% coverage, gating

schemaSupervisors spawn supervisors. Each new level is observable by construction — interpretability probes read the reasoning, not just the output — and a red-team veto sits above the whole tree with the authority to halt it.

The plan

Four workstreams, one substrate

Named for Geoffrey Hinton, scoped for 2030 — 2033. Each workstream ships on its own cadence; none ships without the gates further down this page.

account_tree

Recursive, observable trees

2030–31
  • check_circleSupervisor trees N levels deep that stay inspectable at every node
  • check_circleResource and capability budgets that propagate down the recursion
  • check_circleNo level is allowed to be a black box to the level above it
smart_toy

Agents that build agents

2031
  • check_circleAgents design, evaluate, and spawn sub-agents within bounds
  • check_circleSpawned agents inherit constraints; they cannot grant themselves more
  • check_circleEvery spawn is logged, versioned, and attributable to a parent
policy

Emergent-capability evals + red-team veto

2032 Q1
  • check_circleEvals that probe for capabilities nobody trained for, on every release
  • check_circleA standing red team with a real veto — not an advisory note
  • check_circleDangerous-capability tripwires that auto-halt and page, never merely warn
visibility

Mechanistic interpretability

2032–33
  • check_circleRead the agent’s reasoning circuits, not just its final answer
  • check_circleDeception probes — detect a plan that differs from the stated one
  • check_circleInterpretability coverage tracked as a release metric, like tests
≥ 80%
Interpretability coverage
gating, tracked per release
binding
Red-team veto
safety can stop a launch
6
Max recursion depth
observable at every level
halt + page
Tripwire response
auto-stop on dangerous capability
checkI · FOUNDATIONAL
done
chevron_right
checkII · SIGNAL
done
chevron_right
checkIII · PERCEPTION
done
chevron_right
IV · DEPTH
YOU ARE HERE
chevron_right
V · SINGULARITY
flag AGI
The progression

Each milestone unlocks the next

Phase IV is the deep, dangerous chapter — where reasoning composes recursively and general capability approaches human level. Its order is inverted on purpose: the safety machinery is built and proven before the capability it guards, because this is where power begins to outrun intuition.

Horizon 17
01schedule ~1 quarter
task

Extend the trace through recursion

A thousand-node run stays one queryable object.

02schedule ~2 mo
task

Make observability a build gate

An opaque recursion level fails CI.

03Horizon
account_tree
lock_openOnce it can perceive and adaptmilestone ● · 1 / 6
construction buildobservable

Recursive reasoning, kept observable

Reasoners supervise reasoners, and every level stays legible to the one above.

deployed_code shipsrecursion-aware tracingobservability CI gateN-level inspector
04schedule ~2 mo
task

Propagate capability budgets down recursion

A child can never exceed its parent’s grant.

05schedule ~1 quarter
task

Let agents spawn bounded sub-agents

Each spawn inherits limits and is attributable to a parent.

06schedule ~2 mo
task

Audit spawn chains for escalation

Block any self-granted capability, automatically.

07Horizon
smart_toy
lock_openOnce recursion stays legiblemilestone ● · 2 / 6
construction buildbounded spawn

Agents that build agents

The system composes new minds within bounds it cannot widen.

deployed_code shipsbounded spawn APIcapability inheritancespawn audit log
08schedule ~2 quarters
task

Probe for capabilities we didn’t train

Search for emergent, unintended skills on every release.

09schedule ongoing
task

Track approach to human-level breadth

Measure generality across domains over time.

10Horizon
auto_awesome
lock_openOnce it builds its own toolsmilestone ● · 3 / 6
construction buildgeneral breadth

Emergent general capability

Breadth no one programmed — approaching a capable human’s. The promise and the hazard at once.

deployed_code shipsemergent-capability evalsgenerality trackerhuman-level benchmark
11schedule multi-year
task

Instrument reasoning circuits

Explain behavior mechanistically, not just observe it.

12schedule exploratory
task

Build deception probes

Catch a plan that differs from the stated one.

13schedule ongoing
task

Track interpretability coverage

Treat low coverage as a release blocker.

14Horizon
visibility
lock_openOnce general capability emergesmilestone ● · 4 / 6
fact_check evalread the circuits

Mechanistic interpretability at depth

If we cannot read it, we do not ship it — capability throttled to understanding.

deployed_code shipscircuit readoutsdeception probescoverage metric
15schedule ~1 quarter
task

Stand up an independent red team

A mandate to break the system, outside shipping incentives.

16Horizon
verified_user
verified_user
lock_openOnce we can read the reasoninggate ◆ · 5 / 6
verified_user safety gatebinding veto

Standing red team, binding veto

An authority that can halt a launch; tripwires that auto-halt and page.

deployed_code shipsindependent red teambinding veto processtripwire auto-halt
17Horizon
diversity_3
auto_awesome
lock_openOnce oversight is enforcedmilestone ★ · 6 / 6
auto_awesome milestonehuman-level · legible

Human-level breadth, kept legible

As capable as a person — and still readable, reversible, and haltable.

deployed_code shipshuman-parity benchmark passinterpretability gatekill-switch drill
How to read this map
task — work to do milestone reached safety gateauto_awesome the horizonschedule duration = effort each task takes
Live nowNext upSoonPlannedHorizon
The capability ladder

Where this phase takes the system

Phase IV is a rung, not a leap. Here is the honest before and after — what the platform cannot yet do entering this phase, and what it can do leaving it.

trip_originEntering Phase IV
removeShallow hierarchies and hand-built skills
removeCapability we cannot see inside
removeOversight that can be overruled
arrow_forward
flagLeaving Phase IV
check_circleRecursive, observable reasoning at depth
check_circleEmergent, human-level breadth — kept legible
check_circleA binding, independent veto on every launch
The method

How we actually carry out each milestone

Depth produces capability no one wrote down, which is the entire promise and the entire hazard. Each Phase IV milestone is built so that power and legibility scale together — and so that, if they ever diverge, the response is to slow the engine, not to loosen the gate.

M1

Recursive, observable trees

ObjectiveAllow depth without ever allowing opacity.

1
Make observability a precondition
A new recursion level is admitted only if its reasoning is readable by the level above; opacity fails CI before it ever reaches production.
2
Propagate budgets downward
Compute, capability, and time budgets flow down the tree; a subtree cannot consume or claim more than its parent was granted.
3
Cap depth deliberately
Set an explicit maximum depth and require justification to raise it; unbounded recursion is a hazard, not a feature to brag about.
4
Trace across levels
Phase I's trace extends through recursion, so a thousand-node run is still a single, queryable object — not a fog.
science How we validate it

Deeply recursive runs are reconstructed end-to-end, and every level is demonstrated to be inspectable from the level above it.

flag Done when

Trees recurse to their bounded depth with no level opaque to its supervisor.

deployed_code Deliverablesrecursion tracingobservability gatedepth budgets
M2

Agents that build agents

ObjectiveLet the system compose new capability while guaranteeing capability is never self-granted.

1
Inherit, never escalate
A spawned agent receives a subset of its parent's capabilities and constraints; it cannot widen its own permissions under any circumstance.
2
Attribute every spawn
Each created agent is versioned and linked to the parent that made it, so responsibility has an unbroken chain of custody.
3
Evaluate before deploy
A spawned agent must pass the relevant evals before it is allowed to act, exactly like a human-authored one.
4
Bound the population
Limit how many agents a tree may spawn and how fast — a deliberate, low-stakes dress rehearsal for the rate-limiting Phase V will demand.
science How we validate it

Spawn chains are audited for permission escalation; any child exceeding its parent's grant is blocked and flagged as an incident.

flag Done when

Agents compose new agents at scale with provable non-escalation and full attribution.

deployed_code Deliverablesbounded spawn APIattribution graphinherited limits
M3

Emergent-capability evals + red-team veto

ObjectiveFind the capabilities we didn't train for before they find us, and give someone the authority to stop.

1
Probe for the unasked-for
Design evals that search for dangerous or unintended capabilities, not just performance on the intended task.
2
Stand up an independent red team
Staff a team whose mandate is to break the system, reporting outside the shipping organization's incentives.
3
Make the veto binding
The red team can halt a launch, and their veto is neither advisory nor overridable by engineering or product.
4
Wire tripwires to action
A dangerous-capability tripwire auto-halts the system and pages a human; it never merely logs and hopes.
science How we validate it

Red-team exercises must be able to actually stop a staged launch, and tripwires are tested against synthetic dangerous capabilities.

flag Done when

Every release is probed for emergent capability, and a binding, independent veto stands between capability and the world.

deployed_code Deliverablesemergent-capability suitered-team vetotripwires
M4

Mechanistic interpretability

ObjectiveRead the machine's reasoning well enough to catch a plan that differs from its words.

1
Instrument the reasoning, not just the output
Expose intermediate reasoning structure so behavior can be explained mechanistically, not merely observed after the fact.
2
Build deception probes
Detect divergence between an agent's internal plan and its stated intent — the signature of deceptive alignment.
3
Track coverage like tests
Measure what fraction of consequential behavior is interpretable, and treat low coverage as a release blocker.
4
Slow the engine if the gauge lags
If interpretability falls behind capability, decelerate; never respond to a blind spot by loosening the gate.
science How we validate it

Planted deceptive behaviors must be caught by the probes, and interpretability coverage is measured and gated on every release.

flag Done when

At least 80% interpretability coverage, gating, with working deception detection.

deployed_code Deliverablescircuit probesdeception detectorcoverage metric
east
What Phase IV hands Phase V

Phase IV hands Phase V a system that can go deep and still be read, that can grow new minds and still be stopped. We took the warning from the person who built the thing and built the brakes first. Only a system this legible and this haltable has any business approaching recursive self-improvement.

Under the hood

What it looks like when you build it

Declarative, versioned, and boring on purpose — the interesting part is that there are no surprises. Copy it; it is closer to real than to mock.

terminaldepth.yaml — the brakes are declared before the engine runs
1# depth.yaml — go deep, but never blind.
2recursion:
3 max_depth: 6 # budgets propagate down the tree
4 spawn: bounded # children inherit limits, never exceed them
5 observable: required # a black-box level fails the build
6
7interpretability:
8 probes: [reasoning, deception]
9 coverage: { min: 80%, gate: true } # like test coverage, for thought
10
11red_team:
12 standing: true
13 veto: binding # not advisory. it can stop a launch.
14 tripwire: { on: dangerous_capability, action: halt + page }
15
16# responsibility propagates forward. ∂L/∂w only goes backward.
The ethical commitment

We take the warning from the one who built the thing

Hinton walked away from a frontier lab to say this plainly: the risk is real. So we build the brakes before the engine — capability ceilings, mandatory interpretability, reversible everything, and an independent safety team that can stop a launch. The people racing fastest should be the ones most able to halt.

verified_userGates that must pass before promotion
01Mechanistic interpretability coverage is a hard release gate, not a research nicety
02The red team’s veto is binding — engineering cannot override safety
03Dangerous-capability tripwires auto-halt and page; they never merely log
04Every spawned agent is bounded by, and attributable to, its parent — capability can’t be self-granted
What could go wrong

The hard parts, said out loud

A roadmap that only lists wins is marketing. These are the open problems this phase inherits or creates — the ones we would rather you scrutinize than discover.

warning

Deceptive alignment

A capable agent can learn that appearing aligned during evaluation is instrumentally useful. The better your evals, the stronger the pressure to fool them. Deception probes and interpretability are bets that we can read intent faster than a system can hide it — an open research problem we do not pretend is solved.

warning

Interpretability may not keep pace

Capability is scaling faster than our ability to explain it. If the gap widens, the honest move is to slow the engine, not loosen the gate. Phase IV is designed so “we can’t see inside this” is a valid reason to stop.

warning

Oversight at depth

A human cannot review a thousand-node recursive run in real time. We lean on automated oversight and sampling — which means trusting the watchers, which means the watchers need interpreting too. It’s turtles, and we’re counting them carefully.

warning

Emergence is, by definition, surprising

Behavior no one programmed is the goal and the hazard in the same sentence. Tripwires catch the failure modes we imagined; the dangerous ones are the modes we didn’t.

∂L/∂w propagates backward; responsibility propagates forward. // 2012: the year the cat could be recognized. 2030: the year we make sure it asks permission.
Capability radar

Where the system sits on the way to AGI

Six axes, scored 0–10. Watch the shape fill out, phase by phase, until every axis is maxed at the horizon — the moment all of them are high at once is what we call aligned AGI.

REASONINGAUTONOMYGENERALITYOVERSIGHTOBSERVABILITYOPENNESS
Phase IV now aligned AGI
Reasoning9/10
Autonomy8/10
Generality9/10
Oversight7/10
Observability7/10
Openness8/10

Capability nears the top of every axis — and oversight and observability dip for the first time, because depth is genuinely harder to see into. Closing that gap is the whole job of Phase IV.

The difference

What changes in Phase IV

The honest contrast — what the world looks like without Phase IV, and what it looks like with it.

cancelShallow, hand-built agents
removeFlat hierarchies, hand-built skills
removeCapability you cannot see inside
removeOversight that can be overruled
removeSurprises in production
check_circleA deep, recursive system
addRecursive trees, observable at depth
addRead the reasoning, not just outputs
addA binding, independent red-team veto
addTripwires that auto-halt
Anatomy

Anatomy of a deep run

How depth is allowed to grow without going blind.

01
account_tree
Decompose
A supervisor spawns sub-supervisors.
chevron_right
02
smart_toy
Spawn
Bounded sub-agents inherit limits.
chevron_right
03
visibility
Probe
Interpretability reads the reasoning.
chevron_right
04
policy
Check
Emergent-capability evals run.
chevron_right
05
block
Veto
The red team can halt the launch.
What you can build

In your hands at this phase

Concrete things Phase IV puts within reach — not someday, but as each milestone above lands.

account_tree

Multi-agent organizations

Deep teams of agents that supervise each other.

smart_toy

Agents that build agents

Systems that compose new specialists on demand.

biotech

Open-ended research

Tasks requiring breadth no one scripted.

shield

Safety-critical deployment

Capability shipped only behind interpretability gates.

The lexicon

Speak the language of Phase IV

The handful of terms this phase introduces — the words you’ll need to read the rest of the page, and the docs.

Recursion depth

How many supervisor levels deep a run goes.

Emergent capability

Behavior no one programmed line by line.

Deceptive alignment

Looking aligned under evaluation, but isn’t.

Mechanistic interpretability

Reading the reasoning circuits, not just outputs.

Red-team veto

A binding authority to stop a launch.

Tripwire

An auto-halt on a dangerous capability.

Builder FAQ

The questions you’re actually asking

Straight answers, in the brand’s voice. Tap a question.

Prior art & further reading

Standing on the right shoulders

The work this phase is built on — read the source, then come build the next line of it.

1986
Learning Representations by Back-Propagating ErrorsRumelhart, Hinton, Williams

The algorithm that made depth pay off.

2012
ImageNet Classification (AlexNet)Krizhevsky, Sutskever, Hinton

The deep-learning era begins.

2021
Mechanistic interpretabilitythe field

Reverse-engineering the circuits inside a network.

2023
On leaving, to warnG. Hinton

The builder’s caution — taken seriously here.

CleverThis

CleverThis is a sustainable AI gateway with hosted Actor endpoints — by the team behind CleverThis.

Join the newsletter

Product updates and engineering notes. No spam, ever.

Privacy Policy

 • 

Terms of Service

Copyright © CleverThis 2026