Skip to main content

Open-source code-search benchmark

Testing five code-search tools

I compared five tools on 20 coding tasks. Each task had a reference list of files expected to matter. Codebase Context found 25.7% of those files on average; 11.5% of the files it returned were on the list. The full results and limits are below.

Relevant files found

0-100% scale

Share of each task's reference files found by each tool, averaged across 20 tasks.

Raw search4.3%
Codebase Context25.7%
jCodeMunch27.1%
Repowise20.1% (failed run = 0)
context-mode16.8%

Useful results returned

0-100% scale

Share of each tool's returned files that were on the reference list, averaged across 20 tasks.

Raw search4.0%
Codebase Context11.5%
jCodeMunch11.0%
Repowise12.0%
context-mode11.7%

The bars count returned files. Raw search uses fixed text-search commands; its result is separate from the two-task agent test. Finding a file does not tell us whether the agent's change was correct.

Two tasks: tokens and time

For two open-source tasks, each tool ran three times. The numbers below are the middle result from those three runs for each tool and task. Cached tokens are included in the input total.

Coding task

PonyC

66% less input and 1m 10s faster in this test

Raw search

653,348

input tokens

Task time
3m 05s

Codebase Context

223,558

input tokens

Task time
1m 55s

Coding task

fmt

About 15% less input; Codebase Context was slightly slower in this test

Raw search

290,186

input tokens

Task time
2m 22s

Codebase Context

247,868

input tokens

Task time
2m 24s

Tokens are pieces of text sent to the AI. These totals include cached text. The full breakdown is in the test data. Two tasks cannot tell us whether every task will use less text or finish faster.

How I tested it

Test counts

These counts answer different questions, so I kept them separate instead of combining them into one score.

100

Main test runs

99 completed, 1 failed. We measured files found and returned files on the reference list.

300

Repeat tests

288 completed, 12 Repowise attempts had no usable context. These stay separate as test history.

30

Two-task test

Two tasks, five tools, and three runs of each. We recorded text used and time taken.

100

Older token-count tests

100 older records. I could not confirm that the tools were used as intended, so they are excluded.

What I tested

All five tools searched the same 20 coding tasks. I compared the files they returned with the reference list for each task and kept failed runs in the results. The other tests cover repeated searches and agent runs.

Sources and data

The links below show the earlier public report, test method, rules, and 20 test tasks.

Download test data used for this page. The exact run files remain local, so the tables cannot yet be reproduced from one command in the public checkout.

Limits

  • This test counted returned files. It did not test whether an agent wrote a correct patch.
  • The token comparison uses two tasks. It does not show a general saving.
  • The 100 older token-count tests stay out because I could not confirm that the tools were used as intended.
  • The reported memory readings have no verified unit. They remain in the data but are not used to rank tools.
  • Install caches and indexing work differed, so this is not a fair cost comparison.
  • Query time was not measured the same way for every tool.
  • The agents were asked to use at most 12 tool calls, but that limit was not enforced.

How long did indexing take?

Codebase Context took about 12 minutes 31 seconds to build its index in these tests (the median). It used the local Xenova/bge-small-en-v1.5 embedding model. That time excludes installation and searches. Later updates may take a different amount of time. The tools built different indexes, so their setup times cannot be compared directly.

I could not confirm whether every tool started with an empty model cache or used the GPU.

Codebase Context · Patrick Rossello

Usage analytics: page views and project interactions via Vercel and PostHog. No session replay or visitor profiles.