prism
Ranks a codebase by graph centrality to decide what a small local model gets to see.
- Sole engineer
- 2026
- in development
- Python, Tree-sitter, NetworkX, Ollama
- Source
- 4/8 vs 7/8
- 26×
- 365
- 7
PRISM stands for PageRank-Indexed, Symbol-aware Model. It was built for the problem that a local model with a 4K or 16K context window cannot read your codebase, so something has to decide what it sees. Most tooling answers with embedding similarity. This answers with the call graph.
It runs against Ollama by default, with Anthropic available behind a flag, since the point was making small local models useful rather than making a large hosted one cheaper.
It did not work, in the sense that it lost to a much simpler baseline on the benchmark I built to test it. The architecture is below, then the measurement.
Pass one: ranking a codebase with no model involved
Every run starts deterministically. Tree-sitter parses the source, a symbol relationship graph is built in NetworkX, and PageRank ranks it. No LLM is involved, the result is reproducible, and it is cached per file with hash invalidation.
Ranking by PageRank rather than by text similarity is the bet. A function that half the codebase calls is architecturally central whether or not its name resembles the task. Symbols land in three tiers by percentile: CORE is the top ten percent, SUPPORT the next thirty, PERIPHERAL the rest.
The task is not ignored. Keywords from the task description build a personalisation vector, so the same repository ranks differently depending on what you asked for:
personalization = self._build_personalization(graph, task_keywords or [])
pagerank_scores = nx.pagerank(
graph, alpha=self.config.pagerank_alpha,
personalization=personalization, max_iter=100, weight='weight',
)
Centrality decides the shape, the task tilts it. If PageRank fails to converge the ranker falls back to uniform scores and carries on, since pass one returning nothing is worse than pass one returning something unranked.
This part is fast: about 50 to 200 files per second, roughly 100 ms to build and rank a thousand symbols, under five seconds for a 50,000-line codebase.
Pass two: keeping working memory bounded
The compact index goes to a model, which picks the symbols it wants to see in full. Those get fetched as token-budgeted source slices and fed into a synthesis loop.
Output can exceed the context window as easily as input can, so OutputLog
keeps working memory proportional to the number of chunks rather than the size
of everything produced so far: each synthesis call sees the compact index, the
current slices, and prior log entries, never the full prior output. Documenting
a large project costs the same per call as documenting a small one.
Every budget decision uses tiktoken, with a four-characters-per-token
approximation only as a fallback. Character counts are never used to decide
whether something fits, because being wrong about that shows up as a truncation
at the far end of a long run.
There are two model roles. A lightweight reader model does context planning, section evaluation and fallback chunk reading; a stronger synthesis model does the output and the worker tool loops. On local presets both point at the same model, but the split means the cloud preset does not spend a frontier model on deciding which file to open.
The coding agent
On top of the two passes sits a session pipeline. A planner breaks a task into
sections sized to roughly one context window of work. An evaluator scores them
and splits any that are over budget, recursively, to a maximum depth of three.
Worker agents then run a tool loop per section against ten tools, including
apply_diff, grep_symbol, query_graph for callers and callees, and
run_tests.
Shell access is allowlisted rather than open, and scoped to the project root.
After every file-mutating tool call the test suite runs against the changed files and the result is appended to the tool output, so the model gets a correctness signal on the next turn without having to ask for one. That turned out to be both the most useful idea in the project and the source of its worst bug.
Sessions are persisted to JSON after every completed section, so an interrupted run resumes rather than restarting, and interface changes in one section propagate as revision flags to sections downstream.
Then I measured it
Everything above was architectural reasoning. The 365 unit tests mock the LLM,
so they verify plumbing and say nothing about effectiveness. So I built an eval
harness: a seed repository of 469 lines across seven modules, hidden-test
scoring, and three arms. oneshot pastes the files into one prompt. toolloop
is a plain agent loop of about forty lines with none of PRISM’s machinery.
prism is the full pipeline.
At a 16K window, over eight tasks and 24 runs:
| arm | solved | regressions | LLM calls | tokens |
|---|---|---|---|---|
oneshot |
4/8 | 0 | 1.0 | 1,742 |
toolloop |
7/8 | 0 | 14.9 | 68,226 |
prism |
4/8 | 1 | 10.8 | 46,293 |
The plain tool loop solved every task PRISM solved, plus three more. PRISM won zero tasks the tool loop lost, and produced the only regression in the run. Pasting the files into a single prompt matched PRISM’s score using 26 times fewer tokens.
The compression claim did not hold either. A 10 to 20 times reduction assumes long function bodies. On the benchmark repository the rendered index came to 7,380 tokens describing 3,346 tokens of source, a 2.2 times expansion, because per-symbol overhead exceeds short function bodies.
The caveat, and why it is weak
This benchmark is unfair to PRISM. The entire seed package is 3,346 tokens inside a 16,384-token window, about twenty percent of it. There was no context problem to solve, which is the only problem PRISM exists for, and the compression machinery fired zero times across eight runs. The table says do not reach for this on a small repository. It does not evaluate the design.
So I ran the regime it was built for, a 4K window, and the sign flipped: PRISM solved 2 of 3 where the tool loop solved 1 of 3, including both tasks PRISM had failed at 16K.
That is a hypothesis with weak support rather than a finding. It is one discordant pair, paired significance p=1.000. PRISM still spent 1.5 times the tokens and 2.1 times the calls. It burned thirty calls and 84,000 tokens failing a task it never solved. Run-to-run variance is large enough to swamp all of it: the tool loop solved one task and then failed the identical task twenty minutes later, same arm, same config, temperature 0.2.
What would settle it is written down: three repeats across eight or more tasks at the 4K window, a seed repository large enough that context is scarce at any window size, and an arm that keeps pass one but drops the planner, to separate the index’s contribution from its cost. Partial data suggests the index hurts at 16K, where it is larger than the source it describes.
Seven bugs a green suite could not see
Building the benchmark found seven production bugs on first contact, four of which had been silently disabling the auto-test feature entirely. The unit suite, all 365 tests, was green throughout.
The two worst:
run_tests shelled out to bare pytest, so the project root never reached
sys.path and any healthy repository without a src layout exited with an
error. The automatic post-edit test ran pytest against the edited source file,
which collects zero tests, and reported that to the worker as “Tests FAILED, fix
the issue” after every single edit. So the auto-test mechanism was giving the
model a false failure signal on every turn.
The rest are the same kind of thing. Neither tool loop passed temperature
through, so worker loops ran at the provider default of 1.0 while plain
completions ran at 0.2, making every arm comparison partly a comparison of
sampling temperature. Unknown Ollama models were assumed unable to call tools
and were silently routed into a weaker JSON fallback path. An incremental
indexer compared OS-native paths against forward-slashed ones, so stale symbols
accumulated on every re-index. One preset requested 4,096 completion tokens
inside a 4,096-token window.
All of them are invisible to a test suite that mocks the model, because each one lives at the boundary the mock replaces. A green unit suite that mocks the LLM says the plumbing is connected and nothing about whether the agent works.
What I would keep
The eval harness. It is the part that generalises and the only reason I know any of the above. It also carries the ablation arms that matter more than the headline: one that drops auto-test, one that drops the index, one that gives the plain loop the index without the planner. Those are the questions I would run first if I picked this up again.
I am less sure about the architecture. Graph centrality is a better selection signal than embedding similarity for code, and I still think that is true. Whether it is worth ten times the token cost of pasting the files in is a separate question, and on everything measured so far the answer is no.