We Gave Two AI Systems the Same Four Questions. One Read Everything. One Knew Where to Look.
GISM-0 and Claude Code CLI received identical questions about our live internal workspace. Both got all four right. One used 5.8× fewer tokens, finished 18% faster, and cost roughly 83% less to run. We built the efficient one — a detail we are disclosing upfront so you can calibrate how much to trust us before you get to the part where the numbers are impressive.
Jeffrey Phillips Freeman
If you have ever watched someone search an entire filing cabinet for a document that was sitting right on top of it — methodically reading each folder label aloud, checking anything that looked promising, narrating their progress to no one in particular — you already have an intuitive grasp of what token inefficiency looks like in an AI retrieval system. The search is thorough. The search is also expensive. And the search was, technically, unnecessary.
This is not a slight against the searcher. It is a design observation. If you do not know where things are, searching is the rational strategy. The question is whether your system can be designed to know where things are — and what happens to your inference bill when it does.
GISM-0 is CleverThis's answer to that question. This is the first numerical look at what the answer is worth.
It knew where to look.First, the part where we disclose our bias
We built GISM-0. We designed this benchmark. We ran it on our own internal workspace. We evaluated the results. We are now writing about the results in our own newsletter. If you are scanning this paragraph looking for a conflict of interest to flag, please save yourself the effort: yes, all of them apply, we are aware, we are telling you preemptively so you can adjust your credence accordingly and move on to the interesting part, which is the numbers.
What we can offer, in lieu of the independence we obviously do not have, is specificity. Every methodological choice is documented in the whitepaper linked at the end of this article. Every limitation is stated in this article before you reach the part where the numbers look impressive. The token counts are from the actual run, not estimated or extrapolated. If you find an error in the methodology, we want to know. If the numbers survive scrutiny, great. That is how this is supposed to work.
What these two systems actually are
A brief orientation for readers arriving without a technical glossary — which is not a prerequisite for finding any of this interesting:
Claude Code CLI is Anthropic's flagship agentic AI assistant, one of the most capable and well-engineered systems publicly available. When it needs to answer a factual question about a project or codebase, it searches files, reads the relevant ones, assembles the resulting context into a prompt, and passes that prompt to its underlying language model. It is an excellent system and handles a wide range of tasks — code generation, iterative problem-solving, analysis, research — with considerable sophistication. We are not here to diminish it.
GISM-0 is CleverThis's graph-native retrieval layer. Rather than storing workspace knowledge as files to be searched and read, GISM-0 maintains a structured graph of facts and relationships. When a question arrives, it traverses that graph directly to the relevant facts — the way a seasoned investigator with a well-organized case file pulls the two relevant index cards rather than re-reading the entire dossier from the beginning. Both approaches find the answer. One of them bills you for considerably less time getting there.
Both systems in this benchmark used the same underlying language model: Claude Sonnet 4.6. The intelligence is identical. The retrieval is not.
The test
Four factual questions, each requiring retrieval from multiple files. Our live internal workspace. Same underlying model on both sides (Claude Sonnet 4.6). Single run. Total token counts measured in aggregate — no per-question breakdown was captured. One run is one run. Calibrate accordingly.
We designed four questions that each require locating facts distributed across multiple files in a live workspace — the kind of question a real agent handles dozens of times per session: who is responsible for this account, what does this configuration value do, what was the outcome of the last decision on a particular topic. Real factual retrieval from a moderately large internal workspace. Not a synthetic dataset constructed to flatter the system we built.
Both systems received the same questions. Both had access to the same workspace. Both used the same model. The only variable was retrieval architecture.
What happened
Both systems answered all four questions correctly.
We are stating this clearly before the efficiency numbers take over the conversation: there was no accuracy trade-off. The score was four out of four on both sides. GISM-0 did not achieve its efficiency by cutting corners on the answer quality. The answers were equivalent. The cost of reaching those answers was not.
5.8×
Fewer Tokens (GISM-0)83,749 vs 486,650 aggregate18%
Fasterwall-clock, same model∼83%
Lower Estimated Costlist prices, October 20264/4
Correct (Both Systems)no accuracy trade-offGISM-0: 83,749 tokens. Claude Code CLI: 486,650. The ratio is 5.81. We ran this calculation several ways because 5.81 is a suspiciously tidy number, the kind of result that makes a reasonable person check whether they divided by the wrong thing. We had not divided by the wrong thing. The number is 5.81, and it has been 5.81 for the several weeks since the run, regardless of how many times we have recalculated it.
The honest part: Question 1 went the wrong way
On the first question, GISM-0 was slower than Claude Code CLI. Not slower in aggregate — the 18% speed advantage is the real aggregate figure — but specifically on Question 1, where GISM-0 took a bash/grep traversal path rather than its native semantic graph retrieval.
To stay with the case-file analogy: instead of pulling the relevant index cards, it ran a full search across the filing system. It found the answer. Correctly. Just via the slow route. We are including this because four questions is not enough data to characterize a system's behavior, and readers who think carefully about benchmarks deserve to know the aggregate hides variance. One question went poorly on speed. Three did not. The aggregate reflects all four.
Per-question token counts were not captured separately in this run — only the totals for each system survive. Future instrumentation would give us that resolution, and we intend to build it. We note the gap here so the aggregate number is not mistaken for uniform performance across all four questions.
Why the gap exists — the structural explanation
Token cost in AI retrieval is not arbitrary. It is a direct function of how much text gets passed into the model's context. When you retrieve facts by searching and reading files, the cost of retrieval is the cost of reading — file searches, file reads, context assembly — all of it enters the prompt before the model sees the first token of your actual question.
GISM-0's graph traversal retrieves only the specific facts the question requires. The model receives a small, precise context instead of a large, assembled one. The question determines the size of the prompt, not the workspace. This means the cost of a query is bounded by what the query needs, not by how much the workspace contains.
This is why we believe the efficiency gap grows — rather than closes — as workspaces scale. A file-reading approach reads more as more files exist. A graph traversal approach reads the same amount regardless of workspace size; what changes is the graph, not the cost of navigating it per question. We have not benchmarked this at large scale. That is the honest answer. We have the architectural argument and four data points. Rigorous large-scale evidence is what we are building next.
What this result is not
Four questions, one run, internal workspace, evaluated by the people who built the system. This is directional evidence from a use-case test, not a comprehensive evaluation. Claude Code CLI is a sophisticated system with strengths far beyond factual retrieval. This benchmark measures one specific category. Please draw specifically-scoped conclusions.
We are not claiming GISM-0 is a superior system in any general sense. Claude Code CLI handles code generation, iterative reasoning, tool use, and a wide range of agentic workflows with a depth and polish that reflects years of investment by a large and talented team. “Factual multi-file retrieval from a structured workspace” is the specific thing this test measured. It is a real and practically important category. It is not the whole map of what either system does.
What would turn a use-case test into a real claim: a wider question set, per-question token instrumentation, multiple workspace sizes to characterize how the efficiency gap changes as data volume scales, and evaluation by someone without a financial stake in the outcome. We are working toward that. We are not there yet. We are publishing what we have now because directional evidence, clearly labeled as such, is more useful than waiting for perfection.
The documentation, for those who want to go deeper
Two documents accompany this article:
- GISM-0 Benchmark Whitepaper — Full methodology, raw token counts, architecture overview, and analysis of what drives the efficiency gap.
- GISM-0 Benchmark Deck — Slide-format summary for presentations and stakeholder review.
The whitepaper is the primary technical record. The deck is for the meeting where not everyone has forty minutes, or the stamina.
Why we think the number matters even at small scale
GISM-0 is not a research prototype built specifically to win a benchmark. It is the retrieval layer running in production under CleverThis's workspace tools — the same system processing inference requests on our platform today. In the Agent Amnesia Tax, we documented reducing CleverThis's own AI infrastructure bill by 95.5% in a single quarter. GISM-0's retrieval efficiency is a material part of how that number was reached.
At enterprise scale, 83% lower cost on factual retrieval is not a performance footnote. It is a structural difference in unit economics, the kind that compounds quietly and then suddenly appears as a very large line item advantage on someone's budget comparison. A four-question benchmark does not prove the advantage at scale. The architecture does. The benchmark is evidence that the architecture is performing as designed. We think that is worth reporting, with the appropriate caveats attached — which we have endeavored, at some length, to supply.

CleverThis is a sustainable AI gateway with hosted Actor endpoints — by the team behind CleverThis.
Join the newsletter
Product updates and engineering notes. No spam, ever.
