Repository navigation
Pack context as passages instead of whole chunks; add multi-repo BM25 comparison - #15
Draft
charan-rathore wants to merge 2 commits into
Draft
charan-rathore wants to merge 2 commits into
charan-rathore wants to merge 2 commits into
Conversation
… comparison Same 1,400-token budget, plain BM25 baseline: required-evidence recall 80.5% -> 92.8% over 82 questions on 7 READMEs (holdout of 2 unseen repos: 79.2% -> 91.7%). Evidence retention only, not answer accuracy.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
After ranking, the top chunks are split into sentences. Each sentence is scored against the question (idf-weighted term overlap, a neighbor bonus, and text under a matching heading counts). The same token budget is then filled with the best sentences instead of whole chunks. Chunks that already fit the budget are untouched.
web/src/lib/rag/compact.ts(+compact.test.ts), wired intoretrieve-core.ts. This changes what the retriever hands to the answering model, so it is a behavior change in the live retrieval path, not a side module.eval/multi-repo/(dataset, 6 README fixtures, builder script) andweb/scripts/multi-repo-eval.ts, a no-LLM, no-network comparison of IntelliRAG against plain BM25 at the same 1,400-token budget. Small edits torepo-support-eval.tsand its README.Why
Chunks are about 500 tokens, so only 2-3 fit the budget and most of each is unrelated text. On the p-queue set, required-evidence recall was 70.0% for IntelliRAG against 73.3% for plain BM25. Across 4 more repos it trailed BM25 by about 8 points on average before this change.
Results (eval/multi-repo, 82 answerable questions, 7 READMEs, anchors validated as exact substrings)
compactMinUnitShare0.2): about 79% recall (BM25 level) at about 450 tokens, roughly 60% fewer tokens.The numbers above are from the author of the patch. What I re-ran for this PR: the development split of
multi-repo-eval.tsreproduced 81.0% vs 93.2% (58 questions). I did not re-run--final, to keep the holdout at two runs.Limits (please keep these with any claim)
Checks I ran locally
npx tsx --test src/lib/rag/*.test.ts): 122 passed, 0 failed (4 new).npx tsc --noEmit: clean.packedIdsassigned but never used inretrieve-core.ts).eval/repo-support(needs an API key and spend limit).Draft, not for merge.