Skip to content

Pack context as passages instead of whole chunks; add multi-repo BM25 comparison - #15

Draft
charan-rathore wants to merge 2 commits into
mainfrom
feat/passage-packing-beats-bm25
Draft

charan-rathore wants to merge 2 commits into
mainfrom
feat/passage-packing-beats-bm25

Conversation

@charan-rathore

Copy link
Copy Markdown
Owner

What this changes

After ranking, the top chunks are split into sentences. Each sentence is scored against the question (idf-weighted term overlap, a neighbor bonus, and text under a matching heading counts). The same token budget is then filled with the best sentences instead of whole chunks. Chunks that already fit the budget are untouched.

  • New: web/src/lib/rag/compact.ts (+ compact.test.ts), wired into retrieve-core.ts. This changes what the retriever hands to the answering model, so it is a behavior change in the live retrieval path, not a side module.
  • New: eval/multi-repo/ (dataset, 6 README fixtures, builder script) and web/scripts/multi-repo-eval.ts, a no-LLM, no-network comparison of IntelliRAG against plain BM25 at the same 1,400-token budget. Small edits to repo-support-eval.ts and its README.

Why

Chunks are about 500 tokens, so only 2-3 fit the budget and most of each is unrelated text. On the p-queue set, required-evidence recall was 70.0% for IntelliRAG against 73.3% for plain BM25. Across 4 more repos it trailed BM25 by about 8 points on average before this change.

Results (eval/multi-repo, 82 answerable questions, 7 READMEs, anchors validated as exact substrings)

  • Development (5 repos, 58 questions): BM25 81.0% -> IntelliRAG 93.2%.
  • Holdout (p-map, uuid; 24 questions written before any change; run twice in total, baseline then final): BM25 79.2% -> 91.7%, 3 wins, 0 losses, 95% bootstrap interval 0.0 to 25.0 points. Mean tokens 1254 vs 999.
  • All: 80.5% -> 92.8%, interval 3.9 to 20.9.
  • p-queue set: 73.3% BM25 vs 87.2% IntelliRAG (was 70.0%).
  • Lean mode (compactMinUnitShare 0.2): about 79% recall (BM25 level) at about 450 tokens, roughly 60% fewer tokens.

The numbers above are from the author of the patch. What I re-ran for this PR: the development split of multi-repo-eval.ts reproduced 81.0% vs 93.2% (58 questions). I did not re-run --final, to keep the holdout at two runs.

Limits (please keep these with any claim)

  • Questions and anchors were written by one author (the project). README-only corpora, 7 sources.
  • Keyword mode only, no embeddings.
  • Measures evidence retention, not generated-answer correctness.
  • Refusal on unanswerable questions is unchanged and weak (holdout 0 of 4).
  • Remaining misses are vocabulary gaps (for example "randomly" vs "randomizes") that need dense retrieval.
  • Holdout interval touches 0.0, so the holdout result alone is not conclusive.
  • Licence notices for the checked-in third-party README fixtures have not been checked. Do that before treating the fixtures as publishable.

Checks I ran locally

  • RAG tests (npx tsx --test src/lib/rag/*.test.ts): 122 passed, 0 failed (4 new).
  • npx tsc --noEmit: clean.
  • eslint on the changed files: 0 errors, 1 warning (packedIds assigned but never used in retrieve-core.ts).
  • Not run: full build, Playwright, the answer-quality evaluation in eval/repo-support (needs an API key and spend limit).

Draft, not for merge.

… comparison

Same 1,400-token budget, plain BM25 baseline: required-evidence recall 80.5% -> 92.8% over 82 questions on 7 READMEs (holdout of 2 unseen repos: 79.2% -> 91.7%). Evidence retention only, not answer accuracy.
@vercel

vercel Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
intellirag Ready Ready Preview Oct 4, 2026 11:50am UTC
intellirag-live-own-track Ready Ready Preview Oct 4, 2026 11:50am UTC
irag-fix-0904 Ready Ready Preview Oct 4, 2026 11:50am UTC

This branch was successfully deployed

3 active deployments
Preview – irag-fix-0904 — 0169ffe5 Deployed Oct 4, 2026 by vercel[bot]
Preview – intellirag — 0169ffe5 Deployed Oct 4, 2026 by vercel[bot]
Preview – intellirag-live-own-track — 0169ffe5 Deployed Oct 4, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant