Coding task
PonyC
66% less input and 1m 10s faster in this test
Raw search
653,348
input tokens
- Task time
- 3m 05s
Codebase Context
223,558
input tokens
- Task time
- 1m 55s
Open-source code-search benchmark
I compared five tools on 20 coding tasks. Each task had a reference list of files expected to matter. Codebase Context found 25.7% of those files on average; 11.5% of the files it returned were on the list. The full results and limits are below.
Share of each task's reference files found by each tool, averaged across 20 tasks.
Share of each tool's returned files that were on the reference list, averaged across 20 tasks.
The bars count returned files. Raw search uses fixed text-search commands; its result is separate from the two-task agent test. Finding a file does not tell us whether the agent's change was correct.
For two open-source tasks, each tool ran three times. The numbers below are the middle result from those three runs for each tool and task. Cached tokens are included in the input total.
Coding task
66% less input and 1m 10s faster in this test
653,348
input tokens
223,558
input tokens
Coding task
About 15% less input; Codebase Context was slightly slower in this test
290,186
input tokens
247,868
input tokens
Tokens are pieces of text sent to the AI. These totals include cached text. The full breakdown is in the test data. Two tasks cannot tell us whether every task will use less text or finish faster.
These counts answer different questions, so I kept them separate instead of combining them into one score.
100
99 completed, 1 failed. We measured files found and returned files on the reference list.
300
288 completed, 12 Repowise attempts had no usable context. These stay separate as test history.
30
Two tasks, five tools, and three runs of each. We recorded text used and time taken.
100
100 older records. I could not confirm that the tools were used as intended, so they are excluded.
All five tools searched the same 20 coding tasks. I compared the files they returned with the reference list for each task and kept failed runs in the results. The other tests cover repeated searches and agent runs.
The links below show the earlier public report, test method, rules, and 20 test tasks.
Download test data used for this page. The exact run files remain local, so the tables cannot yet be reproduced from one command in the public checkout.
Codebase Context took about 12 minutes 31 seconds to build its index in these tests (the median). It used the local Xenova/bge-small-en-v1.5 embedding model. That time excludes installation and searches. Later updates may take a different amount of time. The tools built different indexes, so their setup times cannot be compared directly.
I could not confirm whether every tool started with an empty model cache or used the GPU.
Usage analytics: page views and project interactions via Vercel and PostHog. No session replay or visitor profiles.