Live on npm · MIT-licensed · MCP-ready

Stop paying to dump the whole repo.

Every AI coding agent runs into the same wall: it can only see so much code at once, and every irrelevant file it has to read still costs real money and real attention. Context Compiler closes that gap — real semantic search plus structural signals, packed into a token budget, so an agent gets the 20 files that actually matter instead of guessing across 2,000.

~/context-compiler
$ context-compiler "how does the embedding cache work" --budget 3000
Wrote 2997 tokens (6 chunks) to demo-out.md
# Context Bundle
- Task: how does the embedding cache work
- Tokens used: 2997 / 3000
- Chunks included: 6 (125 skipped — didn’t fit budget)
Files included (by relevance):
test/cache.test.ts (lines 29-137) score 0.191 semantic match
src/cache.ts (lines 1-115) score 0.155 semantic match
src/embeddings.ts (lines 116-155) score 0.074 semantic match
A genuine, unedited run against its own source — captured live

Out of 131 candidate chunks across the whole repo, it surfaced the cache implementation and its own test file as the two most relevant — ranked purely by meaning, not by string-matching the words in the question.

99.9%max token reduction — React, 7.26M → 8K tokens
$21.76saved per query, measured, at that scale
$43M/yrmodeled ceiling at enterprise scale — see below
Why it exists

Two workarounds exist for getting an AI agent the code it needs, and both are expensive: hand it the entire repo — slow, costly, and often literally impossible past a few hundred files — or let it explore file-by-file, burning several extra turns just finding its footing before the real task even starts.

Context Compiler does the "find the right code" step once, instantly, using the same category of information-retrieval technology search engines use, instead of guesswork. The goal is narrow and concrete: cut what an agent has to read as much as possible, without ever cutting out something that matters.

The right 20 files out of 2,000. Not a guess.

TypeScript · MCP server · CLI · tree-sitter · OpenAI / Voyage AI embeddings · MIT-licensed
Before Context Compiler

Both existing options cost you something real.

Every modern coding agent — Claude Code, Cursor, Copilot's agent mode — works the same basic way: it reads through a codebase before it writes anything. What it gets to see decides how good the output is. Neither default way of deciding that is actually good.

Dump the whole repo

Slow, costly, and often literally impossible past a few hundred files. React alone runs 7.26 million tokens — no context window holds that, and even when it fits, models measurably struggle with information buried in the middle of a very long prompt.

Explore file-by-file

Works, eventually. But it burns several back-and-forth turns just orienting itself before the agent starts the task you actually asked for — and every one of those turns is billed the same as real work.

How it works

Six stages, in order, every run.

Point it at a repo and a task in plain English. It hands back a markdown bundle of the most relevant code — as a CLI command, or automatically, as an MCP tool (the open standard that lets a compatible agent call outside tools on its own) any MCP-compatible agent can reach for before it starts working.

01

Walk

Lists every real text file in the repo, honoring .gitignore, skipping binaries by extension and by content-sniffing — so a misclassified binary can't slip through.

02

Chunk

Splits files at real function/class/method boundaries via tree-sitter parsing, not arbitrary line cuts. Deepest for JavaScript, TypeScript, TSX and Python; other languages still get clean line-based chunking.

03

Embed

Turns each chunk into a vector via OpenAI or Voyage AI (code-specialized). Unchanged chunks are read from a local, namespace-keyed cache instead of re-embedded, so only your first run on a repo pays the full embedding cost — re-runs are fast and cheap.

04

Rank

Scores chunks by semantic similarity to the task, boosted for files structurally connected via real import-graph parsing, and for anything you explicitly pin.

05

Rerank optional

A second, cheap model pass — GPT-4o-mini or Claude Haiku — reviews the top candidates and drops the ones that don't actually hold up.

06

Pack & format

Greedily fills a token budget you set (default 8,000) with the highest-ranked chunks and formats the result as a clean markdown bundle.

Proof it works

146 tests across 19 files. Run on every change.

Unit tests for every pipeline stage, integration tests, true end-to-end tests that invoke the actual CLI binary, and a packaging-level test that verifies both entry points install and run cleanly.

the actual suite, passing
$ npm test
✓ test/astChunker.test.ts (7 tests)
✓ test/chunker.test.ts (14 tests)
✓ test/rerank.test.ts (15 tests)
✓ test/integration.test.ts (12 tests)
✓ test/cache.test.ts (10 tests)
✓ test/cli.test.ts (9 tests)
✓ test/mcpServer.test.ts (6 tests)
✓ test/importGraph.test.ts (6 tests)
✓ test/packaging.test.ts (2 tests)
…and 7 more test files
Test Files 19 passed (19)
Tests 146 passed (146)
Tested at scale, on real code

Verified against real, unmodified public codebases — not synthetic benchmarks — up to 7,146 files.

146/146tests passing
11real public codebases tested
0crashes, at any scale
Flask · Express · axios · ripgrep · sinatra · gin · React (7,146 files) · Angular · Django · NestJS · FastAPI
Quantified token & cost savings

Measured directly, with the tool's own tokenizer.

"Dump the whole repo" versus the compiled bundle at the default 8,000-token budget, across seven real repos.

RepoFilesFull-repo tokensCompiled bundleReduction$ saved / query*
Flask223267,2587,99897.0%~$0.78
Express211190,4998,00095.8%~$0.55
axios458915,1497,99699.1%~$2.72
ripgrep234913,6887,99799.1%~$2.72
sinatra288241,6237,99996.7%~$0.70
gin130250,3377,99496.8%~$0.73
react7,1467,260,1958,00099.9%~$21.76

* At roughly $3 per million input tokens — a frontier-model ballpark, check current pricing. Input-token cost only, and it recurs every single time an agent is invoked on that repo without a pre-compiled slice.

The key structural point: the compiled bundle's cost stays flat near the configured budget no matter how big the repo is, while a full-repo dump scales linearly with repo size. That's why the percentage saved grows as the codebase grows — this tool pays off more, not less, on the repos where it's hardest to hand-manage context.

Who it's for

The same fix, at every level.

What breaks down when an agent lacks the right context looks different depending on how well you know the codebase you're working in.

01

New to the codebase

Don't know the structure yet? It picks the right files automatically and shows why each one was picked — so you're not stuck guessing where to look.

02

Onboarding fast

The import-graph boost surfaces files structurally connected to your task that a manual search would miss, and --pin locks in the ones you already know matter.

03

Working across repos

One init per repo, and it works identically everywhere — no per-repo setup, and no habitually oversized prompts compounding across every codebase you touch.

04

Owning the AI bill

MIT-licensed and MCP-standard — zero vendor lock-in, fully auditable output, and the savings table above is the artifact you bring to that conversation.

05

Working in a large monorepo

The savings grow with the mess. The React example — 7.26M tokens compiled to an 8,000-token bundle — is what happens at the largest, oldest, most tangled end of a codebase.

What this means at scale

Solo developer, small team, whole org.

Real, measured per-query numbers from the table above, applied at increasing scale. The team and org figures are illustrative models built on that measured data, not guarantees — actual results depend on how often you'd otherwise invoke an agent against the full repo.

PatternRepo scaleQueries/dayPer queryPer monthPer year
Light useMid-size (Flask/Express, ~200–300 files)10~$0.65~$143~$1,716
Heavy / power useLarge (axios/ripgrep scale, ~450 files)40~$2.72~$2,394~$28,730

Real, measured, raw input-token cost for a single developer on a single repo — before counting the harder-to-measure part: an agent that gets the right context immediately needs fewer exploratory turns, which is real time saved on top of the dollar figure.

ScenarioTeamRepo scaleQueries/monthEst. savings/monthEst. savings/year
Small team5 engineersExpress-sized (~200 files)~1,650~$900~$10,800
Growing startup20 engineersaxios-sized (~450 files)~6,600~$18,000~$216,000
Larger org / monorepo50 engineersReact-sized (~7,000 files)~16,500~$359,000~$4,308,000

Assumptions: ~15 agent queries per engineer per working day, ~22 working days/month, using the measured per-query savings from a similarly-sized repo above.

Ceiling, not a forecast

Extending the same measured numbers to their logical extreme — 500 engineers, one React-sized monorepo, 15 queries/engineer/day — lands at ~$3.59M/month, ~$43.08M/year. That assumes zero existing context strategy today, so a real org's savings against its actual baseline would be smaller. What it shows: the mechanism doesn't cap out as repos get bigger and messier — it pays off more where it's needed most.

Try it yourself

Live and installable today. No clone required.

install & run — from npm
$ npm install -g @shreyanshojha/context-compiler
$ export OPENAI_API_KEY=sk-...
$ context-compiler init
# writes .context-compiler.json + prints an MCP snippet, once
$ context-compiler "fix the login bug"
MCP setup — Claude Code
$ claude mcp add context-compiler -s user \
-e OPENAI_API_KEY=sk-... \
-- "$(command -v node)" \
"$(npm root -g)/@shreyanshojha/context-compiler/dist/mcpServer.js"
$ claude mcp list
# should show context-compiler as ✓ Connected

Any other MCP-compatible agent — Cursor, Windsurf, anything that speaks MCP — registers it the same way: point its MCP config at the compiled binary, exactly like the Claude Code command above.

npm: npmjs.com/package/@shreyanshojha/context-compiler · GitHub: github.com/shreyanshojha/context-compiler

Straight answers

The questions people actually ask.

What is Context Compiler, exactly?

A CLI tool and MCP server that pre-compiles the most relevant slice of a repo for a specific task, using real semantic search plus structural signals, and hands an AI coding agent a tight, token-budgeted bundle instead of the whole codebase.

What does it cost?

Context Compiler itself is free and open source (MIT license). You only pay your own embedding-API costs — OpenAI or Voyage AI — which the savings tables above already account for and which are far smaller than what a full-repo dump would cost you per query.

Can I actually use it?

Yes — it's live on npm as @shreyanshojha/context-compiler, and works as a CLI or as an MCP tool with Claude Code or any MCP-compatible agent. Install command is above.

Is the 99.9% reduction real?

It's a real, measured number for one specific case — React's ~7.26 million tokens compiled down to an 8,000-token bundle. It's the largest repo tested, not a typical case; the full table above shows the range across seven repos, from 95.8% to 99.9%.

What's next for it?

More language grammars for AST-aware chunking beyond JS/TS/Python, deeper editor integrations, and a public before/after benchmark comparing it against tools like Aider's repo-map and Cursor's own indexing.

Live on npm · MIT-licensed

Ready to try it on your own repo?

Install it in under a minute, point it at any repo, and see the bundle it builds for your very next task.

Questions about it? Email me →

Context Compiler is a CLI and MCP tool for compiling AI-agent context, currently at v0.7.0.

It is a real, shipped package, live and installable via npm today. Every number and terminal snippet on this page is drawn directly from the project's own test suite and real command output, except where explicitly labeled illustrative or a theoretical ceiling.

© 2026 Shreyansh Ojha. All rights reserved.