Benchmarks

Embedding Models

Nine open embedding models compared inside jscpd's experimental --semantic mode, with thresholds for each one.

jscpd's --semantic mode looks for functions that do the same job without sharing code: renamed and restructured, or written in another language. It embeds every function with a code embedding model and reports the pairs whose vectors point the same way, so the model decides much of what gets reported. We ran nine open models through the same pipeline to find out which one reports the most real duplicates, and whether the thresholds tuned for the default model of the time, jina-embeddings-v2-base-code, hold for the others. After this comparison CodeRankEmbed became the default, and jscpd now keeps thresholds for each of the nine models.

--semantic is experimental. It is on jscpd's master branch and not in a release yet, and neither are the flags on this page (--semantic-url, --semantic-model, --semantic-models, --semantic-threshold, --semantic-same-threshold). The Rust CLI docs describe the mode and its rules.

Summary

  • CodeRankEmbed, a 137M code model from Nomic AI under the MIT license, beats the previous default, jina-embeddings-v2-base-code, on every quality measure below. On the 14 projects that every model finished, it reports about a quarter more real duplicates: an estimated 2,739 against 2,178. It is jscpd's default model now.
  • jina-code-embeddings-0.5b finds the most known clones, but its CC-BY-NC-4.0 license does not allow commercial use, so jscpd cannot ship it.
  • The thresholds of jina-v2, 0.6 for a pair across languages and 0.75 for a pair within one language, fit that model only. At the same precision, the cross-language threshold of the nine models ranges from 0.31 to 0.86, so jscpd now picks the thresholds by the model.
  • A threshold for each pair of languages did worse on held-out data than one threshold for pairs across languages and one for pairs within a language.

The models

The candidates come from the code sections of the MTEB and CoIR leaderboards and from 2026 reviews of embedding models. All of them are open and run on a CPU. embeddinggemma-300m requires a Hugging Face login, so bge-m3 took its place. The 7B and 8B models (nomic-embed-code, Qwen3-Embedding-8B) did not fit in the test machine's memory, and the paid APIs (voyage-code, codestral-embed, OpenAI and Gemini) were not tested.

ModelSizeLicenseTrained forFunctions per second
jina-embeddings-v2-base-code161MApache-2.0code4.3
CodeRankEmbed (now the default)137MMITcode2.6
jina-code-embeddings-0.5b494MCC-BY-NC-4.0code1.3
Qwen3-Embedding-0.6B596MApache-2.0general text and code0.9
SFR-Embedding-Code-400M_R434MCC-BY-NC-4.0code1.0
gte-modernbert-base149MApache-2.0general text and code3.2
codesage-small-v2130MApache-2.0code4.2
granite-embedding-english-r2149MApache-2.0general text and code3.3
bge-m3568MMITgeneral text1.4

Speed is measured on four CPU cores without a GPU, for texts of up to 1,024 tokens, with PyTorch and sentence-transformers. Inside jscpd both CodeRankEmbed and jina-v2 run on candle, and the time includes parsing the files. On four Linux cores jina-v2 managed 2.2 functions per second there. On jscpd's semantic demo, CodeRankEmbed takes 8 seconds and jina-v2 6.

How we compared the models

  • Every model ran behind an OpenAI-compatible server that jscpd called through --semantic-url. jscpd parsed the files, found the functions and sent every model the same texts. The server added the prefix that the model's card recommends for code, where the card names one, and cut texts at 1,024 tokens, the limit jscpd uses for the models it runs itself.
  • Rosetta Code solves one task in many languages, so it is known which functions should pair; we took 100 tasks with 2,388 functions from RosettaCodeData. The models score similarity on different scales, so each one got the thresholds at which its precision equals that of jina-v2 at 0.6 and 0.75, and the models are compared by recall at that precision.
  • In an earlier run of --semantic over 100 open-source projects, reviewers labelled 393 pairs as a duplicate, related code or unrelated. For each model we computed the AUC on them: the chance that a real duplicate scores higher than a pair that is not one.
  • The project set has 30 open-source projects with 8,667 functions: 5 full-stack apps, 6 ports of one library to several languages, 3 mixed projects and 16 in a single language. The four largest models needed more than a day, so the project comparison uses the 14 projects that every model finished, which are the full-stack, port and mixed ones. Language-model reviewers judged 30 pairs from each model without knowing which model had found them, 258 distinct pairs in all.

Results

Known clones and judged pairs

ModelRosetta recall, across languagesRosetta recall, within a languageAUC, across languagesAUC, within a language
jina-v20.5380.3740.8240.805
CodeRankEmbed0.5650.3800.8400.825
jina-code-0.5b0.5870.4600.8510.802
Qwen3-0.6B0.5760.4230.8130.785
SFR-Code-400M0.5320.4170.8620.816
gte-modernbert0.5350.3250.8000.824
codesage-small-v20.5790.3560.8360.796
granite-r20.5070.3500.7990.868
bge-m30.4400.2880.8390.867

Every model's precision on Rosetta Code is matched to jina-v2's: about 0.90 across languages and 0.68 within a language.

Known clones found on Rosetta Code at equal precisionRecall at the precision jina-v2 has at its thresholds, 0.6 and 0.75.
Across languagesWithin a language
jina-code-0.5bcodesage-small-v2Qwen3-0.6BCodeRankEmbedjina-v2gte-modernbertSFR-Code-400Mgranite-r2bge-m3

Real projects

ModelDuplicates among judged pairsPairs on 14 projectsEstimated real duplicatesPairs shared with jina-v2
jina-v283% ± 72,6132,178
CodeRankEmbed93% ± 52,9352,7390.73
jina-code-0.5b83% ± 73,0062,5050.72
Qwen3-0.6B80% ± 73,1332,5060.69
SFR-Code-400M90% ± 52,3942,1550.64
gte-modernbert87% ± 62,6452,2920.68
codesage-small-v267% ± 93,0742,0490.66
granite-r270% ± 82,7701,9390.62
bge-m380% ± 72,3841,9070.64

The estimate of real duplicates is the number of pairs times the share of duplicates among the judged ones. Pairs shared with jina-v2 is the Jaccard index: the pairs both models report, divided by the pairs either of them reports. With 30 judged pairs per model, the share of duplicates is accurate to 5 to 9 percentage points.

Estimated real duplicates on 14 projectsPairs reported × the share of duplicates among 30 blind-judged pairs. jscpd’s default model is highlighted.
CodeRankEmbed2,739Qwen3-0.6B2,506jina-code-0.5b2,505gte-modernbert2,292jina-v22,178SFR-Code-400M2,155codesage-small-v22,049granite-r21,939bge-m31,907

What stands out

  • CodeRankEmbed is ahead of jina-v2 on recall at equal precision, on both AUCs and among the judged pairs, and it reports the most real duplicates. It is slower in the same setup, 2.6 functions per second against 4.3.
  • jina-code-embeddings-0.5b has the best recall on Rosetta Code, 0.587 across languages and 0.460 within a language, where jina-v2 has 0.374. It is three times slower than jina-v2, and its license limits it to non-commercial use.
  • SFR-Embedding-Code-400M ranks cross-language duplicates best among the judged pairs (AUC 0.862), and 90% of its judged pairs are duplicates. It carries the same non-commercial license, and only Qwen3 is slower.
  • Qwen3-Embedding-0.6B does well on Rosetta Code but ranks the judged pairs lower than jina-v2, at about a fifth of its speed.
  • bge-m3, granite-r2 and gte-modernbert find fewer known clones within a language at the same precision: 0.288 to 0.350, where jina-v2 finds 0.374. granite-r2 and bge-m3 still rank same-language duplicates best among the judged pairs (AUC 0.868 and 0.867).
  • codesage-small-v2 is as fast as jina-v2 and finds many cross-language pairs, but only 67% of its judged pairs are duplicates.
  • Each model shares between 62% and 73% of its pairs with jina-v2.

Before the rules

The rules of --semantic (mutual best matches, thresholds and z-scores) work on each model's raw similarities. On Rosetta Code, without the rules:

ModelPartner ranked first, acrossPartner ranked first, withinPartner cosineImpostor cosineGapBest F0.5, acrossBest F0.5, within
jina-v278%73%0.690.510.170.7930.589
CodeRankEmbed77%73%0.540.340.180.8160.620
jina-code-0.5b82%76%0.700.460.230.8240.648
Qwen3-0.6B83%80%0.700.510.180.8180.657
SFR-Code-400M73%67%0.800.720.070.7970.636
gte-modernbert74%69%0.750.650.090.7970.602
codesage-small-v280%76%0.490.250.220.8230.600
granite-r276%71%0.880.830.040.7810.610
bge-m367%60%0.710.660.050.7530.545

The partner is the function that solves the same task in the other language; the impostor is the most similar function from another task. The numbers are medians over all language pairs. The two cosine columns show each model's scale: granite-r2 puts partners at 0.88 and impostors at 0.83, codesage-small-v2 puts them at 0.49 and 0.25. A threshold set for one model does not carry over to another.

Where each model puts similar and dissimilar codeMedian cosine of a function and its partner in another language, and of the most similar function from another task. The tick is the model’s threshold across languages.
PartnerImpostorThreshold across languages
codesage-small-v2CodeRankEmbedjina-code-0.5bQwen3-0.6Bjina-v2gte-modernbertbge-m3SFR-Code-400Mgranite-r2

Thresholds for each model

These thresholds give each model the precision that jina-v2 has at 0.6 and 0.75 on Rosetta Code. jscpd keeps them for all nine models: --semantic-model picks the model, the thresholds follow it, and jscpd --semantic-models prints the list. --semantic-threshold and --semantic-same-threshold still override them. Set only the first, and jscpd keeps the model's gap between the two.

--semantic-modelAcross languagesWithin a languageGap
CodeRankEmbed (default)0.41250.63750.225
jina-embeddings-v2-base-code0.60.750.15
jina-code-embeddings-0.5b0.56250.71250.15
Qwen3-Embedding-0.6B0.58750.76250.175
SFR-Embedding-Code-400M_R0.73750.83750.1
gte-modernbert-base0.68750.850.1625
codesage-small-v20.31250.56250.25
granite-embedding-english-r20.86250.9250.0625
bge-m30.70.83750.1375

jscpd runs the first two models itself, after jscpd --semantic-download (with --semantic-model jina-embeddings-v2-base-code for the second). The others need an OpenAI-compatible server. jscpd sends the server the model name as typed and knows the models by their Hugging Face ids and their Ollama names too, so this gets the Qwen3 thresholds:

ollama pull qwen3-embedding:0.6b
jscpd . --semantic --semantic-url http://localhost:11434/v1 --semantic-model qwen3-embedding:0.6b

Two of the models expect an instruction before the code: jina-code-embeddings-0.5b and Qwen3-Embedding-0.6B. The comparison put the prompt from the model card before every function, and jscpd does the same. A server that adds the prompt itself, such as the Jina API with a task, needs "prefix": "" in the config file so the prompt is not added twice.

The rule for groups of copies has two more settings: how close to a function's best match another match may score and still count as one, and the similarity a pair needs when its functions are not each other's best match. Both were tuned with jina-v2 only, at 0.05 and 0.8. For the other models jscpd scales them by the model's gap, which shows how widely the model spreads its scores: CodeRankEmbed gets 0.075 and 0.7125. These values are derived from the thresholds and were not measured in this comparison.

One threshold for each language pair?

--semantic uses one threshold for pairs across languages and one for pairs within a language. We checked with jina-v2, on the labelled sets, whether each pair of languages (Rust with TypeScript, Python with Python and so on) should get its own. We fitted the thresholds on half of the tasks and scored them on the other half, in three splits with two passes each. The numbers are F0.5.

ThresholdsRosetta, acrossRosetta, withinCodeNet, acrossCodeNet, within
0.6 and 0.75, the thresholds of jina-v20.8120.7030.3840.392
the best pair fitted on the training half0.8120.7030.3830.415
one fitted for each language pair0.8090.6650.3800.407

On Rosetta Code the best pair of thresholds came out at exactly 0.6 and 0.75 in five splits out of six. The best threshold for a single language pair moves around between splits: from 0.55 to 0.85 for Python with Python, from 0.53 to 0.78 for C# with Python. The data shows no stable value for any language pair; the differences are noise. The z-score rule of --semantic already accounts for the language, since it measures each pair against the background of functions in that language.

What this means for jscpd

  • CodeRankEmbed is jscpd's default model now. jscpd runs its NomicBERT architecture on candle; the vectors match those of sentence-transformers to six decimals, and the download is pinned by revision and SHA-256. Its thresholds, 0.4125 and 0.6375, are the defaults. The 100-project study with blind review has not been repeated with it yet, so the switch rests on the two labelled sets and the 14 projects above.
  • The thresholds belong to the model. jscpd picks them from --semantic-model, and --semantic-threshold, --semantic-same-threshold (or threshold and sameThreshold in the config file) override them. A model outside the list gets jina-v2's 0.6 and 0.75, and jscpd warns that they are not calibrated for it.
  • jina-code-embeddings-0.5b and SFR-Embedding-Code-400M are better at single measures, but their license rules them out as defaults. jscpd --semantic-models lists them with their licenses, for use through --semantic-url.
  • The 7B and 8B models and the paid APIs rank highest for code search in published reviews. Testing them needs a GPU or API keys.

Limitations

  • With 30 judged pairs per model, the share of duplicates is accurate to 5 to 9 percentage points. The gap between CodeRankEmbed and jina-v2 among the judged pairs, 93% against 83%, is at the edge of that. The conclusion rests on CodeRankEmbed being ahead on Rosetta Code, on the AUC and among the judged pairs at once.
  • The 14 shared projects are mostly ports, where cross-language pairs are most common. Their share of duplicates is higher than across the 100-project study: 83% against 57% for jina-v2.
  • jina-v2, the default model at the time, picked the 393 judged pairs. They measure how each model ranks the candidates that jina-v2 found, and say nothing about pairs it missed.
  • All texts were cut at 1,024 tokens. Models with a longer window might gain on long functions, at a higher cost in time.
  • The runs used transformers 4.51 and sentence-transformers 6.1 on a CPU. With transformers 5 the large models ran about twice as fast, but two of the models with custom code do not load there.