The leap from autocomplete to understanding. Phase II is GISM — the engine that makes the machine actually reason.
Capability without faithfulness is a confident liar. GISM makes reasoning explicit, grounded, and measurable — the difference between a model that sounds right and one that can show its work.
Plans, reflects, and revises mid-thought — the jump single-shot models can’t make.
The chain of thought it shows is provably the one it used.
Everything Phase III learns, it learns on top of trustworthy reasoning.
Father of information theory · the man who measured meaning
Claude Shannon’s 1937 master’s thesis — still arguably the most consequential ever written — showed that Boolean logic could be built from electrical switches, handing us the digital circuit. Then in 1948 he did it again: “A Mathematical Theory of Communication” defined the bit, proved how much information a noisy channel can carry, and quietly founded the field every modem, codec, and language model rests on. In his spare time he juggled on a unicycle and built a machine whose only function was to switch itself off.
“The fundamental problem of communication is that of reproducing at one point a message selected at another.”
— A Mathematical Theory of Communication, 1948
Yes — Claude Shannon, and Claude the model. The naming is a homage the field is quietly proud of. Phase II takes the “information” in information theory literally and scores agents in bits. Entropy isn’t a vibe; it’s an integral.
Phase II turns runs into signal. Evaluation harnesses become first-class, versioned artifacts — graded on every run, gating every promotion. Telemetry, durable memory, and context compression let agents carry state across long horizons without drowning in it. We score capability, faithfulness, and drift in numbers, not adjectives.
Shannon’s frame is the whole design: an agent is a channel between what you intended and what it did, and hallucination is noise on that channel. The job isn’t to wish the noise away — it is to measure it, drive it toward its entropy floor, and protect the signal that survives. A system you can’t trust to report itself accurately can’t be governed, however capable it looks in a demo.
schemaIntent enters; the agent encodes it into actions; the world adds noise; we decode what actually happened and compare. The gap is hallucination — measured, not described. Evals tap the channel at every stage.
Named for Claude Shannon, scoped for 2026 — 2028. Each workstream ships on its own cadence; none ships without the gates further down this page.
Phase II is GISM — the leap from a model that completes patterns to one that genuinely reasons and understands. These milestones build reasoning that is explicit, grounded, faithful, and measurable: a mind whose thoughts you can read, check, and trust. This is the difference between an autocomplete and an understanding.
A model that plans and revises mid-thought instead of decoding in one pass.
Make the thinking an explicit, inspectable artifact.
Tasks that single-shot models fail by construction.
Each assertion links to a source or is flagged as ungrounded.
The system marks what it is certain of versus guessing.
Does the shown reasoning match the reasoning actually used?
Catch post-hoc stories that do not drive the answer.
A durable store where every memory traces to its origin.
A structured model it reasons over across runs.
Does understanding persist between sessions, not reset?
Self-critique that catches its own errors.
Escalate genuine uncertainty instead of confidently guessing.
Stated confidence matches measured accuracy.
Capability and faithfulness graded continuously.
No untrustworthy reasoning is allowed to ship.
Phase II is a rung, not a leap. Here is the honest before and after — what the platform cannot yet do entering this phase, and what it can do leaving it.
Information theory gives us the discipline: define the signal, name the noise, and measure the gap. Each Phase II milestone is a way of putting a defensible number on something the field too often describes with a shrug.
ObjectiveMake evaluation a reproducible, adversarial, versioned artifact that the whole system is accountable to.
Two independent graders plus one adversarial grader must agree within tolerance; eval-on-eval drift is itself monitored over time.
Every capability claim is backed by a versioned, reproducible eval that anyone with the repo can re-run and challenge.
ObjectivePut a continuous, claim-level number on how much of what an agent says is actually grounded.
Sampled human adjudication calibrates the automatic scorer, and the scorer's own error bars are published alongside its verdicts.
Hallucination rate ≤ 0.5% per release at the claim level, with calibration error ≤ 0.05 — both gating.
ObjectiveLet agents act over long horizons by remembering faithfully and forgetting deliberately.
Long-horizon tasks are replayed with and without compression; faithfulness must not degrade beyond a published threshold.
Agents sustain multi-day tasks with auditable memory and no measurable loss of faithfulness from compression.
ObjectiveMake capability and its erosion visible across time, so improvement is a trend you can defend, not a launch-day boast.
Synthetic regressions are injected and must be caught by the drift alarms before any human notices the change in behavior.
Every agent carries a defensible capability/faithfulness history, and no silent regression survives a release.
Phase II hands Phase III a system that can measure itself honestly. Learning without measurement is just drift with good intentions — so the meter had to come before the motor. Now the platform can be allowed to change itself, because we can finally tell whether a change made it better.
Declarative, versioned, and boring on purpose — the interesting part is that there are no surprises. Copy it; it is closer to real than to mock.
Hallucination is noise on the channel between intent and act. We drive it toward its entropy floor and defend the faithful signal — because a system you can’t trust to report itself accurately can’t be governed, no matter how capable.
A roadmap that only lists wins is marketing. These are the open problems this phase inherits or creates — the ones we would rather you scrutinize than discover.
“When a measure becomes a target, it ceases to be a good measure.” The moment an eval gates promotion, agents — and the people tuning them — optimize the eval. Held-out adversarial sets and rotating graders push back, but the arms race is permanent.
Some uncertainty is irreducible — the world is genuinely ambiguous. The goal isn’t a 0% hallucination rate (that’s a confident liar); it’s a calibrated one that knows the floor and reports it.
Heavy instrumentation has cost and can alter behavior — the observer effect, in software. Telemetry budgets and sampling keep the channel from being dominated by the meter.
SNR ≥ 0 dB or it does not ship. // Shannon juggled on a unicycle; we settle for juggling correctness and capability.
Six axes, scored 0–10. Watch the shape fill out, phase by phase, until every axis is maxed at the horizon — the moment all of them are high at once is what we call aligned AGI.
Reasoning jumps from a flat 2 to a 7 — the GISM leap. Autonomy and generality stay deliberately low: a mind that reasons, but does not yet act or generalize on its own.
The honest contrast — what the world looks like without Phase II, and what it looks like with it.
How GISM turns an intent into a reasoned, auditable answer.
Concrete things Phase II puts within reach — not someday, but as each milestone above lands.
Agents whose every claim is grounded and auditable.
Tasks that need planning, not just recall.
Assistants that remember and reason across sessions.
Agents that catch their own errors before you do.
The handful of terms this phase introduces — the words you’ll need to read the rest of the page, and the docs.
The reasoning engine — plans, reflects, revises.
Tying a claim to a verifiable source.
The shown reasoning matches the used reasoning.
Stated confidence matches measured accuracy (ECE).
An internal model the system reasons over.
The irreducible uncertainty a faithful system reports.
Straight answers, in the brand’s voice. Tap a question.
The work this phase is built on — read the source, then come build the next line of it.
Entropy, the bit, and the noisy channel.
Reasoning steps unlock multi-step ability.
Does the explanation reflect the computation?
When does stated confidence track accuracy?

CleverThis is a sustainable AI gateway with hosted Actor endpoints — by the team behind CleverThis.
Join the newsletter
Product updates and engineering notes. No spam, ever.