Guides

Semantic Clones

Find functions that do the same job with different code, in one language or across languages, with jscpd --semantic and a code embedding model.

Token matching finds code that someone copied. It cannot find two functions that do the same job with different code, such as a validation rule that a Rust backend enforces and a Svelte frontend repeats, or two helpers that two people wrote for the same task. These are Type-4 clones. --semantic looks for them with a code embedding model. jscpd turns every function into a vector and reports two functions as a clone of kind semantic when their vectors point the same way.

Available from jscpd 5.3.3 and experimental. Besides real duplicates it reports related code, such as a client function and the server endpoint it calls, so review a pair before you merge two functions.

Quick start

Terminal
jscpd --semantic-download     # once: fetches the model, 548 MB, into the jscpd cache
jscpd --semantic src/         # pairs within one language and across languages

The model runs inside jscpd on the CPU, so a scan makes no network call. The first scan embeds every function; later scans take the vectors of unchanged functions from the cache. jscpd reads the functions of JavaScript, TypeScript, JSX, TSX, Vue, Svelte, Astro, Python, Rust, Go, Java, Kotlin, C#, C, C++, PHP, Ruby, Scala and Swift files that clear --min-tokens and --min-lines.

A pair appears in the same report as the token clones:

Clone found (rust, semantic ~0.73)
 - backend/src/pricing.rs [34:1 - 58:2] (25 lines, 174 tokens)
   frontend/src/lib/components/CartSummary.svelte:typescript [19:3 - 31:4]

To try it on a ready-made project, run it on fixtures/semantic-demo: a Rust API and a SvelteKit frontend with eight rules written on both sides and two features written twice in one language.

How jscpd pairs two functions

jscpd embeds each function's code without comments, starting at the name the function is declared under. It then reports two functions as a pair when all of these hold:

  • They are in different files, neither calls the other, and the clones the token passes already found do not cover both. A function and the helper it calls are related, not duplicated.
  • Each is the other's closest match among the functions of its language, or close to it. A feature written three times makes three pairs, while a function that resembles many others, such as a request handler, pairs once per language at most.
  • Their similarity reaches the threshold: one for a pair across languages and a higher one for two functions of one language, because code in one language resembles other code in that language whatever it does.
  • The similarity stands at least 3 standard deviations above each function's similarity to the rest of the other function's language, so a family of generated look-alikes does not pair with itself.

--skip-local and --skip-isolated apply to semantic pairs as they do to token clones.

Ways to use it

The same feature written twice in one language

Two people solve one problem their own way, and the code base ends up with two slug functions or two e-mail checks. --semantic-scope same keeps only the pairs within one language:

Terminal
jscpd --semantic --semantic-scope same src/

The bar within one language is higher on purpose. If you know a pair of duplicates that the report leaves out, look at its score and lower --semantic-same-threshold a little, as described in Thresholds.

A rule repeated across a backend and a frontend

A project with a Rust or Python backend and a TypeScript or Svelte frontend often checks the same rule on both sides. --semantic-scope cross keeps only the pairs across languages. When the two halves live in their own folders, scanning them as two paths with --skip-local keeps only the pairs that cross from one folder to the other:

Terminal
jscpd --semantic --semantic-scope cross .
jscpd --semantic --skip-local backend/ frontend/

Code ported to another language

When a library exists in two languages, or a service is moving from one language to another, the pairs across languages show which functions have a counterpart on the other side:

Terminal
jscpd --semantic --skip-local --kind semantic python/ rust/

A function without a counterpart does not appear in the report. jscpd does not embed functions smaller than --min-tokens or --min-lines, so lower those two for a code base with many short functions.

Only semantic clones, for an agent or a script

--kind semantic keeps only the semantic pairs. The ai reporter prints one line per pair, and JSON marks each pair with "kind": "semantic" and its "similarity":

Terminal
jscpd --semantic --kind semantic -r ai src/
# backend/src/pricing.rs:34-58 ~ frontend/src/lib/components/CartSummary.svelte:typescript:19-31 [~0.73 semantic]

jscpd --semantic --kind semantic -r json -o report src/

In CI

The model download is 548 MB, so keep the jscpd cache between runs. The cache also holds the vectors, and with it a run embeds only the functions that changed. On a Linux runner the cache is ~/.cache/jscpd:

.github/workflows/semantic.yml
- uses: actions/cache@v4
  with:
    path: ~/.cache/jscpd
    key: jscpd-${{ runner.os }}-${{ github.run_id }}
    restore-keys: jscpd-${{ runner.os }}-
- run: npx jscpd --semantic-download
- run: npx jscpd --semantic --kind semantic -r ai src/

--semantic-download does nothing when the model is already in the cache. Semantic clones count in the statistics like any clone, so --threshold and --exit-code see them. While the mode is experimental, keep it out of the step that fails the build on duplication and run it as a report of its own.

Choosing a model

Each model scores similarity on its own scale: two functions that one model scores 0.9 another scores 0.5. jscpd keeps calibrated thresholds for nine models, and jscpd --semantic-models lists them:

MODEL                         CROSS   SAME    LICENSE       RUNS
CodeRankEmbed (default)       0.4125  0.6375  MIT           in jscpd, 548 MB
jina-embeddings-v2-base-code  0.6     0.75    Apache-2.0    in jscpd, 324 MB
jina-code-embeddings-0.5b     0.5625  0.7125  CC-BY-NC-4.0  API
Qwen3-Embedding-0.6B          0.5875  0.7625  Apache-2.0    API (Ollama: qwen3-embedding:0.6b)
SFR-Embedding-Code-400M_R     0.7375  0.8375  CC-BY-NC-4.0  API
gte-modernbert-base           0.6875  0.85    Apache-2.0    API
codesage-small-v2             0.3125  0.5625  Apache-2.0    API
granite-embedding-english-r2  0.8625  0.925   Apache-2.0    API
bge-m3                        0.7     0.8375  MIT           API (Ollama: bge-m3)

CROSS is the default threshold for a pair across languages and SAME the one for a pair within one language. --semantic-model takes a name from the list, its Hugging Face id or its Ollama name, in any letter case, and the thresholds follow the model. Embedding Models compares the nine models.

jscpd runs the first two models itself. CodeRankEmbed, the default, found the most real duplicates in the comparison. jina-embeddings-v2-base-code is faster. jscpd's demo takes 6 seconds with it and 8 with CodeRankEmbed.

Terminal
jscpd --semantic-download jina-embeddings-v2-base-code
jscpd --semantic --semantic-model jina-embeddings-v2-base-code src/

The other seven need an embeddings API. jina-code-embeddings-0.5b and SFR-Embedding-Code-400M_R are licensed for non-commercial use only.

Using an embeddings API

--semantic-url sends the functions to a server that speaks the OpenAI embeddings API instead of running the model in jscpd: Ollama, LM Studio, llama-server --embedding, text-embeddings-inference or a hosted API. The server gets the model name as typed, so give the name the server knows.

With Ollama, the default model of an API is jina-embeddings-v2-base-code under its Ollama name, because Ollama's library has no copy of CodeRankEmbed:

Terminal
ollama pull unclemusclez/jina-embeddings-v2-base-code
jscpd --semantic --semantic-url http://localhost:11434/v1 src/

ollama pull qwen3-embedding:0.6b
jscpd --semantic --semantic-url http://localhost:11434/v1 --semantic-model qwen3-embedding:0.6b src/

A hosted API needs a key. jscpd reads it only from the JSCPD_SEMANTIC_API_KEY environment variable and sends it only to a URL given with --semantic-url or to a server on this machine, and never over plain http to another machine. With OpenAI:

Terminal
JSCPD_SEMANTIC_API_KEY=sk-… jscpd --semantic --semantic-url https://api.openai.com/v1 \
  --semantic-model text-embedding-3-small src/

OpenAI's models are not in the list of calibrated models, so jscpd uses 0.6 and 0.75 for them and warns that these thresholds are not calibrated. See Thresholds for how to set your own.

With an API, the code of every function goes to that API. When the code must not leave the machine, use the model that runs in jscpd, or run the API on the machine itself, like Ollama on localhost.

A config file can arrive with the code being scanned, for example in a pull request, so it cannot send that code or your key anywhere on its own. A url in the config file that is not on this machine is used only when --semantic itself is on the command line, and it never receives the key.

Three config keys matter only for some APIs and models:

  • dimensions asks the API for shorter vectors. That works with models trained for it, such as OpenAI's text-embedding-3 models and jina-code-embeddings-0.5b. When a server ignores the setting and sends the full vector, jscpd cuts it to that length.
  • params go into every request as they are, for an API that needs more than the model name.
  • prefix is the text jscpd puts before every function. Qwen3-Embedding-0.6B and jina-code-embeddings-0.5b get the instruction they were calibrated with. Set prefix for another model that expects one, or to "" when the API adds the instruction itself.

If the server cannot be reached, the model is missing or the key is rejected, the run fails with exit code 1 and a hint, such as the ollama pull command to run. It never reports such a run as a clean scan.

Thresholds

--semantic-threshold is the lowest similarity of a pair across languages, and --semantic-same-threshold the lowest similarity of a pair within one language. Both default to the calibrated values of the model: 0.4125 and 0.6375 for CodeRankEmbed. When you set only --semantic-threshold, the bar within one language keeps the model's gap above it: 0.225 for CodeRankEmbed, 0.15 for a model jscpd has not calibrated.

Terminal
jscpd --semantic --semantic-threshold 0.5 src/                                 # stricter; 0.725 within one language
jscpd --semantic --semantic-threshold 0.5 --semantic-same-threshold 0.7 src/   # a bar for each kind of pair

To set thresholds for a model jscpd does not know, run it on code where you know a few real pairs, read their similarity in the JSON report, and set both thresholds a little below them. Scores are comparable only within one model, so a threshold tuned for one model says nothing about another.

The vector cache

jscpd caches the vectors, keyed by the model, its settings and the function's text, so a change to a comment embeds nothing. Each set of scanned paths has its own cache folder, so jscpd . and jscpd src keep separate files. When the vectors of changed and deleted functions make up more than a quarter of a file, the next run that embeds something rewrites the file without them.

SystemCache directory
macOS~/Library/Caches/jscpd
Linux$XDG_CACHE_HOME/jscpd, or ~/.cache/jscpd
Windows%LOCALAPPDATA%\jscpd\cache

JSCPD_CACHE_DIR moves the cache elsewhere. The models folder holds the downloaded models and the embeddings folder the vectors. --semantic-rebuild-cache embeds every function again and replaces the cached vectors of the scanned paths, and "cache": false in the config file turns the vector cache off. To free the space, delete the embeddings folder; deleting models means --semantic-download has to fetch the model again.

What it does not do

jscpd compares whole functions only, so it does not embed top-level code, templates or styles. Only the clone scan embeds: --history, --dashboard, --health, --complexity and --mcp ignore --semantic.

A score says that two functions look alike to the model, and it cannot tell whether they behave alike. Related code, such as a client call and the endpoint it calls or a route and its test, shows up next to real duplicates.

Options

Command line

OptionDescriptionDefault
--semanticFind semantic clones (Type-4, experimental)off
--semantic-download [MODEL]Download a model jscpd runs itself into the cache directory, checked against its pinned SHA-256: MODEL, the one --semantic-model names, or CodeRankEmbed. Alone it exits after the download; with --semantic it goes on to scan with that model—
--semantic-scopeWhich pairs to report: all, same (within one language) or cross (across languages)all
--semantic-thresholdLowest cosine similarity of a pair across languages, in (0, 1]the model's calibrated value, 0.4125 for CodeRankEmbed; 0.6 for a model jscpd has not calibrated
--semantic-same-thresholdLowest cosine similarity of a pair within one language, in (0, 1]the model's calibrated value, 0.6375 for CodeRankEmbed; with --semantic-threshold set, that value plus the model's gap
--semantic-modelEmbedding model: a name from --semantic-models, its Hugging Face id, or any name an API servesCodeRankEmbed; unclemusclez/jina-embeddings-v2-base-code with --semantic-url
--semantic-modelsList the models jscpd has calibrated, with their thresholds, licenses and where they run, and exit—
--semantic-urlOpenAI-compatible embeddings API, e.g. http://localhost:11434/v1 for Ollama; selects the http provider—
--semantic-providerWhere embeddings come from: local (the model run in jscpd) or http (an embeddings API)local; http when a URL is given
--semantic-rebuild-cacheWith --semantic, embed every function again and replace the cached vectors of the model in useoff
--kind semanticReport only semantic clonesall kinds

Config file

The semantic key in .jscpd.json takes true, or an object with these keys. An object turns the mode on only with "enabled": true, so a project can keep its settings in the file and choose the run on the command line.

.jscpd.json
{
  "semantic": {
    "enabled": true,
    "scope": "all",
    "model": "CodeRankEmbed",
    "threshold": 0.4125,
    "sameThreshold": 0.6375,
    "cache": true
  }
}
KeyDescriptionDefault
enabledRun the semantic passfalse
providerlocal or httplocal; http when url is set
scopeall, same or crossall
thresholdLowest similarity across languagesthe model's calibrated value
sameThresholdLowest similarity within one languagethe model's calibrated value
modelEmbedding modelCodeRankEmbed
urlEmbeddings API; jscpd uses a URL that is not on this machine only when --semantic is on the command line—
dimensionsVector length to ask the API forthe model's own
paramsExtra fields for every API request—
prefixText put before every functionthe model's calibrated instruction, or none
cacheKeep vectors in the cache directorytrue

jscpd stops the run when the config file holds an API key.

Environment variables

VariableEffect
JSCPD_SEMANTIC_API_KEYThe key for an embeddings API
JSCPD_CACHE_DIRThe directory for models and vectors, instead of the user cache directory
HF_ENDPOINTA Hugging Face mirror for --semantic-download

See also