← All work

prism

Ranks a codebase by graph centrality to decide what a small local model gets to see.

Role
Sole engineer
Year
2026
Status
in development
Built with
Python, Tree-sitter, NetworkX, Ollama
Go look
  • 4/8 vs 7/8Benchmark, 16K window
  • 26×Token cost vs one-shot
  • 365Unit tests, all green
  • 7Bugs the eval found anyway

PRISM stands for PageRank-Indexed, Symbol-aware Model. It was built for the problem that a local model with a 4K or 16K context window cannot read your codebase, so something has to decide what it sees. Most tooling answers with embedding similarity. This answers with the call graph.

It runs against Ollama by default, with Anthropic available behind a flag, since the point was making small local models useful rather than making a large hosted one cheaper.

It did not work, in the sense that it lost to a much simpler baseline on the benchmark I built to test it. The architecture is below, then the measurement.

Pass one: ranking a codebase with no model involved

Every run starts deterministically. Tree-sitter parses the source, a symbol relationship graph is built in NetworkX, and PageRank ranks it. No LLM is involved, the result is reproducible, and it is cached per file with hash invalidation.

Ranking by PageRank rather than by text similarity is the bet. A function that half the codebase calls is architecturally central whether or not its name resembles the task. Symbols land in three tiers by percentile: CORE is the top ten percent, SUPPORT the next thirty, PERIPHERAL the rest.

The task is not ignored. Keywords from the task description build a personalisation vector, so the same repository ranks differently depending on what you asked for:

personalization = self._build_personalization(graph, task_keywords or [])
pagerank_scores = nx.pagerank(
    graph, alpha=self.config.pagerank_alpha,
    personalization=personalization, max_iter=100, weight='weight',
)

Centrality decides the shape, the task tilts it. If PageRank fails to converge the ranker falls back to uniform scores and carries on, since pass one returning nothing is worse than pass one returning something unranked.

This part is fast: about 50 to 200 files per second, roughly 100 ms to build and rank a thousand symbols, under five seconds for a 50,000-line codebase.

Pass two: keeping working memory bounded

The compact index goes to a model, which picks the symbols it wants to see in full. Those get fetched as token-budgeted source slices and fed into a synthesis loop.

Output can exceed the context window as easily as input can, so OutputLog keeps working memory proportional to the number of chunks rather than the size of everything produced so far: each synthesis call sees the compact index, the current slices, and prior log entries, never the full prior output. Documenting a large project costs the same per call as documenting a small one.

Every budget decision uses tiktoken, with a four-characters-per-token approximation only as a fallback. Character counts are never used to decide whether something fits, because being wrong about that shows up as a truncation at the far end of a long run.

There are two model roles. A lightweight reader model does context planning, section evaluation and fallback chunk reading; a stronger synthesis model does the output and the worker tool loops. On local presets both point at the same model, but the split means the cloud preset does not spend a frontier model on deciding which file to open.

The coding agent

On top of the two passes sits a session pipeline. A planner breaks a task into sections sized to roughly one context window of work. An evaluator scores them and splits any that are over budget, recursively, to a maximum depth of three. Worker agents then run a tool loop per section against ten tools, including apply_diff, grep_symbol, query_graph for callers and callees, and run_tests.

Shell access is allowlisted rather than open, and scoped to the project root.

After every file-mutating tool call the test suite runs against the changed files and the result is appended to the tool output, so the model gets a correctness signal on the next turn without having to ask for one. That turned out to be both the most useful idea in the project and the source of its worst bug.

Sessions are persisted to JSON after every completed section, so an interrupted run resumes rather than restarting, and interface changes in one section propagate as revision flags to sections downstream.

Then I measured it

Everything above was architectural reasoning. The 365 unit tests mock the LLM, so they verify plumbing and say nothing about effectiveness. So I built an eval harness: a seed repository of 469 lines across seven modules, hidden-test scoring, and three arms. oneshot pastes the files into one prompt. toolloop is a plain agent loop of about forty lines with none of PRISM’s machinery. prism is the full pipeline.

At a 16K window, over eight tasks and 24 runs:

arm solved regressions LLM calls tokens
oneshot 4/8 0 1.0 1,742
toolloop 7/8 0 14.9 68,226
prism 4/8 1 10.8 46,293

The plain tool loop solved every task PRISM solved, plus three more. PRISM won zero tasks the tool loop lost, and produced the only regression in the run. Pasting the files into a single prompt matched PRISM’s score using 26 times fewer tokens.

The compression claim did not hold either. A 10 to 20 times reduction assumes long function bodies. On the benchmark repository the rendered index came to 7,380 tokens describing 3,346 tokens of source, a 2.2 times expansion, because per-symbol overhead exceeds short function bodies.

The caveat, and why it is weak

This benchmark is unfair to PRISM. The entire seed package is 3,346 tokens inside a 16,384-token window, about twenty percent of it. There was no context problem to solve, which is the only problem PRISM exists for, and the compression machinery fired zero times across eight runs. The table says do not reach for this on a small repository. It does not evaluate the design.

So I ran the regime it was built for, a 4K window, and the sign flipped: PRISM solved 2 of 3 where the tool loop solved 1 of 3, including both tasks PRISM had failed at 16K.

That is a hypothesis with weak support rather than a finding. It is one discordant pair, paired significance p=1.000. PRISM still spent 1.5 times the tokens and 2.1 times the calls. It burned thirty calls and 84,000 tokens failing a task it never solved. Run-to-run variance is large enough to swamp all of it: the tool loop solved one task and then failed the identical task twenty minutes later, same arm, same config, temperature 0.2.

What would settle it is written down: three repeats across eight or more tasks at the 4K window, a seed repository large enough that context is scarce at any window size, and an arm that keeps pass one but drops the planner, to separate the index’s contribution from its cost. Partial data suggests the index hurts at 16K, where it is larger than the source it describes.

Seven bugs a green suite could not see

Building the benchmark found seven production bugs on first contact, four of which had been silently disabling the auto-test feature entirely. The unit suite, all 365 tests, was green throughout.

The two worst:

run_tests shelled out to bare pytest, so the project root never reached sys.path and any healthy repository without a src layout exited with an error. The automatic post-edit test ran pytest against the edited source file, which collects zero tests, and reported that to the worker as “Tests FAILED, fix the issue” after every single edit. So the auto-test mechanism was giving the model a false failure signal on every turn.

The rest are the same kind of thing. Neither tool loop passed temperature through, so worker loops ran at the provider default of 1.0 while plain completions ran at 0.2, making every arm comparison partly a comparison of sampling temperature. Unknown Ollama models were assumed unable to call tools and were silently routed into a weaker JSON fallback path. An incremental indexer compared OS-native paths against forward-slashed ones, so stale symbols accumulated on every re-index. One preset requested 4,096 completion tokens inside a 4,096-token window.

All of them are invisible to a test suite that mocks the model, because each one lives at the boundary the mock replaces. A green unit suite that mocks the LLM says the plumbing is connected and nothing about whether the agent works.

What I would keep

The eval harness. It is the part that generalises and the only reason I know any of the above. It also carries the ablation arms that matter more than the headline: one that drops auto-test, one that drops the index, one that gives the plain loop the index without the planner. Those are the questions I would run first if I picked this up again.

I am less sure about the architecture. Graph centrality is a better selection signal than embedding similarity for code, and I still think that is true. Whether it is worth ten times the token cost of pasting the files in is a separate question, and on everything measured so far the answer is no.