Semantic Clones
Token matching finds code that someone copied. It cannot find two functions that do the same job with different code, such as a validation rule that a Rust backend enforces and a Svelte frontend repeats, or two helpers that two people wrote for the same task. These are Type-4 clones. --semantic looks for them with a code embedding model. jscpd turns every function into a vector and reports two functions as a clone of kind semantic when their vectors point the same way.
Quick start
jscpd --semantic-download # once: fetches the model, 548 MB, into the jscpd cache
jscpd --semantic src/ # pairs within one language and across languages
The model runs inside jscpd on the CPU, so a scan makes no network call. The first scan embeds every function; later scans take the vectors of unchanged functions from the cache. jscpd reads the functions of JavaScript, TypeScript, JSX, TSX, Vue, Svelte, Astro, Python, Rust, Go, Java, Kotlin, C#, C, C++, PHP, Ruby, Scala and Swift files that clear --min-tokens and --min-lines.
A pair appears in the same report as the token clones:
Clone found (rust, semantic ~0.73)
- backend/src/pricing.rs [34:1 - 58:2] (25 lines, 174 tokens)
frontend/src/lib/components/CartSummary.svelte:typescript [19:3 - 31:4]
To try it on a ready-made project, run it on fixtures/semantic-demo: a Rust API and a SvelteKit frontend with eight rules written on both sides and two features written twice in one language.
How jscpd pairs two functions
jscpd embeds each function's code without comments, starting at the name the function is declared under. It then reports two functions as a pair when all of these hold:
- They are in different files, neither calls the other, and the clones the token passes already found do not cover both. A function and the helper it calls are related, not duplicated.
- Each is the other's closest match among the functions of its language, or close to it. A feature written three times makes three pairs, while a function that resembles many others, such as a request handler, pairs once per language at most.
- Their similarity reaches the threshold: one for a pair across languages and a higher one for two functions of one language, because code in one language resembles other code in that language whatever it does.
- The similarity stands at least 3 standard deviations above each function's similarity to the rest of the other function's language, so a family of generated look-alikes does not pair with itself.
--skip-local and --skip-isolated apply to semantic pairs as they do to token clones.
Ways to use it
The same feature written twice in one language
Two people solve one problem their own way, and the code base ends up with two slug functions or two e-mail checks. --semantic-scope same keeps only the pairs within one language:
jscpd --semantic --semantic-scope same src/
The bar within one language is higher on purpose. If you know a pair of duplicates that the report leaves out, look at its score and lower --semantic-same-threshold a little, as described in Thresholds.
A rule repeated across a backend and a frontend
A project with a Rust or Python backend and a TypeScript or Svelte frontend often checks the same rule on both sides. --semantic-scope cross keeps only the pairs across languages. When the two halves live in their own folders, scanning them as two paths with --skip-local keeps only the pairs that cross from one folder to the other:
jscpd --semantic --semantic-scope cross .
jscpd --semantic --skip-local backend/ frontend/
Code ported to another language
When a library exists in two languages, or a service is moving from one language to another, the pairs across languages show which functions have a counterpart on the other side:
jscpd --semantic --skip-local --kind semantic python/ rust/
A function without a counterpart does not appear in the report. jscpd does not embed functions smaller than --min-tokens or --min-lines, so lower those two for a code base with many short functions.
Only semantic clones, for an agent or a script
--kind semantic keeps only the semantic pairs. The ai reporter prints one line per pair, and JSON marks each pair with "kind": "semantic" and its "similarity":
jscpd --semantic --kind semantic -r ai src/
# backend/src/pricing.rs:34-58 ~ frontend/src/lib/components/CartSummary.svelte:typescript:19-31 [~0.73 semantic]
jscpd --semantic --kind semantic -r json -o report src/
In CI
The model download is 548 MB, so keep the jscpd cache between runs. The cache also holds the vectors, and with it a run embeds only the functions that changed. On a Linux runner the cache is ~/.cache/jscpd:
- uses: actions/cache@v4
with:
path: ~/.cache/jscpd
key: jscpd-${{ runner.os }}-${{ github.run_id }}
restore-keys: jscpd-${{ runner.os }}-
- run: npx jscpd --semantic-download
- run: npx jscpd --semantic --kind semantic -r ai src/
--semantic-download does nothing when the model is already in the cache. Semantic clones count in the statistics like any clone, so --threshold and --exit-code see them. While the mode is experimental, keep it out of the step that fails the build on duplication and run it as a report of its own.
Choosing a model
Each model scores similarity on its own scale: two functions that one model scores 0.9 another scores 0.5. jscpd keeps calibrated thresholds for nine models, and jscpd --semantic-models lists them:
MODEL CROSS SAME LICENSE RUNS
CodeRankEmbed (default) 0.4125 0.6375 MIT in jscpd, 548 MB
jina-embeddings-v2-base-code 0.6 0.75 Apache-2.0 in jscpd, 324 MB
jina-code-embeddings-0.5b 0.5625 0.7125 CC-BY-NC-4.0 API
Qwen3-Embedding-0.6B 0.5875 0.7625 Apache-2.0 API (Ollama: qwen3-embedding:0.6b)
SFR-Embedding-Code-400M_R 0.7375 0.8375 CC-BY-NC-4.0 API
gte-modernbert-base 0.6875 0.85 Apache-2.0 API
codesage-small-v2 0.3125 0.5625 Apache-2.0 API
granite-embedding-english-r2 0.8625 0.925 Apache-2.0 API
bge-m3 0.7 0.8375 MIT API (Ollama: bge-m3)
CROSS is the default threshold for a pair across languages and SAME the one for a pair within one language. --semantic-model takes a name from the list, its Hugging Face id or its Ollama name, in any letter case, and the thresholds follow the model. Embedding Models compares the nine models.
jscpd runs the first two models itself. CodeRankEmbed, the default, found the most real duplicates in the comparison. jina-embeddings-v2-base-code is faster. jscpd's demo takes 6 seconds with it and 8 with CodeRankEmbed.
jscpd --semantic-download jina-embeddings-v2-base-code
jscpd --semantic --semantic-model jina-embeddings-v2-base-code src/
The other seven need an embeddings API. jina-code-embeddings-0.5b and SFR-Embedding-Code-400M_R are licensed for non-commercial use only.
Using an embeddings API
--semantic-url sends the functions to a server that speaks the OpenAI embeddings API instead of running the model in jscpd: Ollama, LM Studio, llama-server --embedding, text-embeddings-inference or a hosted API. The server gets the model name as typed, so give the name the server knows.
With Ollama, the default model of an API is jina-embeddings-v2-base-code under its Ollama name, because Ollama's library has no copy of CodeRankEmbed:
ollama pull unclemusclez/jina-embeddings-v2-base-code
jscpd --semantic --semantic-url http://localhost:11434/v1 src/
ollama pull qwen3-embedding:0.6b
jscpd --semantic --semantic-url http://localhost:11434/v1 --semantic-model qwen3-embedding:0.6b src/
A hosted API needs a key. jscpd reads it only from the JSCPD_SEMANTIC_API_KEY environment variable and sends it only to a URL given with --semantic-url or to a server on this machine, and never over plain http to another machine. With OpenAI:
JSCPD_SEMANTIC_API_KEY=sk-… jscpd --semantic --semantic-url https://api.openai.com/v1 \
--semantic-model text-embedding-3-small src/
OpenAI's models are not in the list of calibrated models, so jscpd uses 0.6 and 0.75 for them and warns that these thresholds are not calibrated. See Thresholds for how to set your own.
With an API, the code of every function goes to that API. When the code must not leave the machine, use the model that runs in jscpd, or run the API on the machine itself, like Ollama on localhost.
A config file can arrive with the code being scanned, for example in a pull request, so it cannot send that code or your key anywhere on its own. A url in the config file that is not on this machine is used only when --semantic itself is on the command line, and it never receives the key.
Three config keys matter only for some APIs and models:
dimensionsasks the API for shorter vectors. That works with models trained for it, such as OpenAI's text-embedding-3 models and jina-code-embeddings-0.5b. When a server ignores the setting and sends the full vector, jscpd cuts it to that length.paramsgo into every request as they are, for an API that needs more than the model name.prefixis the text jscpd puts before every function. Qwen3-Embedding-0.6B and jina-code-embeddings-0.5b get the instruction they were calibrated with. Setprefixfor another model that expects one, or to""when the API adds the instruction itself.
If the server cannot be reached, the model is missing or the key is rejected, the run fails with exit code 1 and a hint, such as the ollama pull command to run. It never reports such a run as a clean scan.
Thresholds
--semantic-threshold is the lowest similarity of a pair across languages, and --semantic-same-threshold the lowest similarity of a pair within one language. Both default to the calibrated values of the model: 0.4125 and 0.6375 for CodeRankEmbed. When you set only --semantic-threshold, the bar within one language keeps the model's gap above it: 0.225 for CodeRankEmbed, 0.15 for a model jscpd has not calibrated.
jscpd --semantic --semantic-threshold 0.5 src/ # stricter; 0.725 within one language
jscpd --semantic --semantic-threshold 0.5 --semantic-same-threshold 0.7 src/ # a bar for each kind of pair
To set thresholds for a model jscpd does not know, run it on code where you know a few real pairs, read their similarity in the JSON report, and set both thresholds a little below them. Scores are comparable only within one model, so a threshold tuned for one model says nothing about another.
The vector cache
jscpd caches the vectors, keyed by the model, its settings and the function's text, so a change to a comment embeds nothing. Each set of scanned paths has its own cache folder, so jscpd . and jscpd src keep separate files. When the vectors of changed and deleted functions make up more than a quarter of a file, the next run that embeds something rewrites the file without them.
| System | Cache directory |
|---|---|
| macOS | ~/Library/Caches/jscpd |
| Linux | $XDG_CACHE_HOME/jscpd, or ~/.cache/jscpd |
| Windows | %LOCALAPPDATA%\jscpd\cache |
JSCPD_CACHE_DIR moves the cache elsewhere. The models folder holds the downloaded models and the embeddings folder the vectors. --semantic-rebuild-cache embeds every function again and replaces the cached vectors of the scanned paths, and "cache": false in the config file turns the vector cache off. To free the space, delete the embeddings folder; deleting models means --semantic-download has to fetch the model again.
What it does not do
jscpd compares whole functions only, so it does not embed top-level code, templates or styles. Only the clone scan embeds: --history, --dashboard, --health, --complexity and --mcp ignore --semantic.
A score says that two functions look alike to the model, and it cannot tell whether they behave alike. Related code, such as a client call and the endpoint it calls or a route and its test, shows up next to real duplicates.
Options
Command line
| Option | Description | Default |
|---|---|---|
--semantic | Find semantic clones (Type-4, experimental) | off |
--semantic-download [MODEL] | Download a model jscpd runs itself into the cache directory, checked against its pinned SHA-256: MODEL, the one --semantic-model names, or CodeRankEmbed. Alone it exits after the download; with --semantic it goes on to scan with that model | — |
--semantic-scope | Which pairs to report: all, same (within one language) or cross (across languages) | all |
--semantic-threshold | Lowest cosine similarity of a pair across languages, in (0, 1] | the model's calibrated value, 0.4125 for CodeRankEmbed; 0.6 for a model jscpd has not calibrated |
--semantic-same-threshold | Lowest cosine similarity of a pair within one language, in (0, 1] | the model's calibrated value, 0.6375 for CodeRankEmbed; with --semantic-threshold set, that value plus the model's gap |
--semantic-model | Embedding model: a name from --semantic-models, its Hugging Face id, or any name an API serves | CodeRankEmbed; unclemusclez/jina-embeddings-v2-base-code with --semantic-url |
--semantic-models | List the models jscpd has calibrated, with their thresholds, licenses and where they run, and exit | — |
--semantic-url | OpenAI-compatible embeddings API, e.g. http://localhost:11434/v1 for Ollama; selects the http provider | — |
--semantic-provider | Where embeddings come from: local (the model run in jscpd) or http (an embeddings API) | local; http when a URL is given |
--semantic-rebuild-cache | With --semantic, embed every function again and replace the cached vectors of the model in use | off |
--kind semantic | Report only semantic clones | all kinds |
Config file
The semantic key in .jscpd.json takes true, or an object with these keys. An object turns the mode on only with "enabled": true, so a project can keep its settings in the file and choose the run on the command line.
{
"semantic": {
"enabled": true,
"scope": "all",
"model": "CodeRankEmbed",
"threshold": 0.4125,
"sameThreshold": 0.6375,
"cache": true
}
}
| Key | Description | Default |
|---|---|---|
enabled | Run the semantic pass | false |
provider | local or http | local; http when url is set |
scope | all, same or cross | all |
threshold | Lowest similarity across languages | the model's calibrated value |
sameThreshold | Lowest similarity within one language | the model's calibrated value |
model | Embedding model | CodeRankEmbed |
url | Embeddings API; jscpd uses a URL that is not on this machine only when --semantic is on the command line | — |
dimensions | Vector length to ask the API for | the model's own |
params | Extra fields for every API request | — |
prefix | Text put before every function | the model's calibrated instruction, or none |
cache | Keep vectors in the cache directory | true |
jscpd stops the run when the config file holds an API key.
Environment variables
| Variable | Effect |
|---|---|
JSCPD_SEMANTIC_API_KEY | The key for an embeddings API |
JSCPD_CACHE_DIR | The directory for models and vectors, instead of the user cache directory |
HF_ENDPOINT | A Hugging Face mirror for --semantic-download |
See also
- Embedding Models: how the nine models compare, and where the thresholds come from.
- Types of Code Clones: where semantic clones sit next to exact, renamed and near-miss clones.
fixtures/semantic-demo: a runnable example with the expected output of every command.- Rust CLI docs: the full rule set.