Concepts

Types of code clones

Exact, renamed, near-miss and semantic clones, and which of them jscpd finds.

"Duplicated code" covers more than copy and paste. Clone-detection research sorts duplicates into four types by how much the copies have drifted apart, and every detector, jscpd included, draws its line somewhere in that scale. Knowing the types tells you what a scan can and cannot show, and which flag to reach for.

TypeAlso calledThe copies differ injscpd
Type-1exact clonewhitespace, layout, commentsdefault
Type-2renamed, parameterized cloneidentifier names, literal values, annotations--ignore-identifiers, --ignore-literals, --ignore-annotations
Type-3near-miss, gapped cloneadded, removed or changed statements--max-gap-lines N merges the pieces around a gap; --similarity RATIO compares whole JS/TS functions by syntax tree
Type-4semantic cloneeverything but the behavior--semantic compares functions with a code embedding model (experimental, 5.3.3+)
One function, four ways to copy it: exact by default, renamed with --ignore-identifiers, an inserted line with --max-gap-lines 2, scattered edits with --similarity 0.7.

Type-1: exact clones

Two fragments are Type-1 clones when their tokens are identical. Whitespace, line breaks and, depending on the mode, comments do not count.

// a.js
function total(items) {
  let sum = 0;
  for (const item of items) { sum += item.price; }
  return sum;
}

// b.js
function total(items) {
  let sum = 0;   // running total
  for (const item of items) {
    sum += item.price;
  }
  return sum;
}

jscpd finds these by default. The --mode option decides what "identical" ignores:

  • mild (default) drops whitespace, so layout changes are invisible.
  • weak also drops comments, so a re-commented copy still matches.
  • strict keeps everything except blocks marked with jscpd:ignore-start / jscpd:ignore-end.

Type-1 is the category to gate in CI: an exact clone is nearly always a copy that should be a shared function, and the result is stable enough for a baseline.

Type-2: renamed clones

A Type-2 clone is a Type-1 clone whose identifiers, literals or annotations were changed while the structure stayed the same. This is what a copy looks like after someone adapted it to a second use.

// cart.js
export function cartTotal(items, taxRate) {
  let subtotal = 0;
  for (const item of items) {
    subtotal += item.price * item.quantity;
  }
  const tax = subtotal * taxRate;
  return { subtotal, tax, total: subtotal + tax };
}

// basket.js
export function basketTotal(entries, vatRate) {
  let net = 0;
  for (const entry of entries) {
    net += entry.price * entry.quantity;
  }
  const vat = net * vatRate;
  return { net, vat, gross: net + vat };
}

A default scan reports nothing here, because every other token differs. jscpd 5.2 adds three flags that normalize token classes before hashing:

FlagConfig keyEffect
--ignore-identifiersignoreIdentifiersevery identifier hashes as $id; keywords are kept, so for still has to match for
--ignore-literalsignoreLiteralsstrings hash as $str, numbers as $num; a string never matches a number
--ignore-annotationsignoreAnnotations@Name, @a.b.Name and @Name(...) are dropped in Java, Kotlin, Scala, Groovy, Python, Dart, Swift, JavaScript and TypeScript
jscpd --ignore-identifiers src/
# Clone found (javascript, renamed)
#  - basket.js [1:1 - 9:2] (9 lines, 57 tokens)
#    cart.js [1:1 - 9:2]

Clones found this way are labelled so they never blend into the exact ones. Every clone carries a kind: exact when the raw tokens match, renamed when they match only after normalization. The console prints Clone found (javascript, renamed), the JSON report has a "kind" field, and the SARIF report files renamed clones under the rule jscpd/renamed-code instead of jscpd/duplicate-code, so GitHub code scanning shows them as a separate rule.

Normalized runs find more and longer clones than exact runs, so their fingerprints differ. Keep a separate --baseline file for a normalized configuration; the one from an exact scan does not match it.

--ignore-case is a much smaller step in the same direction: it folds Total and total together and nothing else.

The repository ships a runnable demo of each flag in fixtures/type2-demo, with the expected output of every command.

Type-3: near-miss clones

Type-3 clones are copies with statements added, removed or changed in the middle: a validation line inserted here, a logging call dropped there.

// original                          // copy with an inserted check
function save(user) {                function save(user) {
  const row = toRow(user);             const row = toRow(user);
  row.updatedAt = Date.now();          if (!row.id) throw new Error('id');
  db.put(row);                         row.updatedAt = Date.now();
  audit('save', row.id);               db.put(row);
}                                      audit('save', row.id);
                                     }

A token-window detector sees this as two shorter exact clones with a gap between them, and that is what a default jscpd run reports, provided each side of the gap still clears --min-tokens and --min-lines.

From jscpd 5.2, --max-gap-lines N (config key maxGapLines) merges clones of the same file pair whose fragments follow each other in both files with at most N unmatched lines between them into one clone of kind similar. Its tokens value is the number of matched tokens and similarity is that number divided by the tokens of the longer merged span, so one inserted line in a 150-token block reads as roughly 0.9:

jscpd src/
# Found 2 clones.

jscpd --max-gap-lines 1 src/
# Clone found (javascript, similar (gap) ~0.91)
#  - save-account.js [1:1 - 12:2] (12 lines, 157 tokens)
#    save-user.js [1:1 - 11:2]
# Found 1 clones.

The merge only joins clones the exact run already found, so it cannot invent a match; it removes fragmentation. It is off at the default of 0, and the JSON report carries "kind": "similar" with the "similarity" value, while SARIF files these under jscpd/similar-code. A runnable pair lives in fixtures/type3-demo.

Edits spread through a whole function, with no single gap, still escape a token window. For JavaScript and TypeScript, --similarity RATIO compares whole functions by structure instead: each function, method or arrow function is summarized by the 4-grams of its syntax-tree node types, and two functions are reported as one similar clone when the weighted Jaccard index of those bags reaches RATIO. Names and values are not part of the summary, so a renamed copy scores 1.0, one inserted line about 0.9, and two inserted statements plus renames about 0.75:

jscpd --similarity 0.85 src/     # near-identical structure
jscpd --similarity 0.7 src/      # a couple of added or removed statements
# Clone found (javascript, similar (ast) ~0.75)
#  - credit-note.js [1:8 - 19:2] (19 lines, 126 tokens)
#    invoice.js [1:8 - 17:2]

Functions must clear --min-tokens and --min-lines on their own, and a pair already reported as an exact or merged clone is not repeated. The MCP server finds these pairs when a tool call asks for them, and its check_duplication tool takes the ratio as similarity for the functions of a snippet. Both options are off by default. A runnable pair lives in fixtures/type3-demo/similar-functions.

Type-4: semantic clones

Type-4 clones compute the same thing with different code, such as a for loop and a reduce call that both sum prices, two sorting routines, or one rule written in Rust for a backend and again in TypeScript for its frontend. No token run and no syntax tree connects the copies, so the passes above cannot see them.

From jscpd 5.3.3, the experimental --semantic mode looks for them with a code embedding model. It turns every function into a vector and reports two functions as a clone of kind semantic when each is the other's closest match and their vectors point the same way, within one language or across languages. The model runs inside jscpd after a one-time jscpd --semantic-download, or behind an embeddings API.

jscpd --semantic-download                        # once: the model, 548 MB
jscpd --semantic src/                            # pairs within one language and across languages
jscpd --semantic --semantic-scope cross .        # only pairs across languages

A similarity score says that two functions look alike to the model, and it cannot tell whether they behave alike, so review a pair before you merge the two functions. Semantic Clones covers the ways to use it and every option.

Reporting one kind

From jscpd 5.3.0, --kind (config key kind) keeps only the clones of the kinds it lists: exact, renamed, similar, or one of the two mechanisms behind similar, gap (--max-gap-lines) and ast (--similarity). From 5.3.3 it also takes semantic (--semantic). Statistics and --threshold are computed after the filter, so the percentage describes what is reported.

jscpd --ignore-identifiers --kind renamed src/                # only the renamed copies
jscpd --max-gap-lines 2 --similarity 0.7 --kind gap,ast src/  # only near-miss clones

The filter never turns a detector on. --kind ast without --similarity prints Warning: --kind ast: no such clones are found without --similarity, and an unknown kind is an error, so a typo cannot make a scan look clean. A runnable walk-through is in fixtures/type3-demo.

The tools of the MCP server take the same names in their kinds argument, and type1 to type4 too. There, asking for a kind turns its detector on, so a call that asks for renamed clones gets them even when the server runs without --ignore-identifiers.

Which type to look for

  • In CI, gate on Type-1. Exact clones are unambiguous and stable across runs, which is what a failing check needs. Use --min-tokens and --min-lines to set the size, and a baseline to fail only on new duplication.
  • When refactoring, scan for Type-2. Run --ignore-identifiers --ignore-literals on the area you are about to change; renamed copies are where a shared helper pays off most. Read the renamed results as leads to check, since some parametric similarity is idiomatic.
  • Treat adjacent Type-1 findings as a pair. Two clones between the same files a few lines apart usually are one near-miss clone with an edit in between. Rerun with --max-gap-lines 2 to see them as one and read the similarity value.
  • Hunt Type-3 in JavaScript and TypeScript with --similarity. Start at 0.85 to catch near-identical functions, lower it towards 0.7 to see heavier edits. The score ranks the pairs; it does not judge them.
  • Look for Type-4 when two halves repeat each other. Run --semantic on a backend and a frontend in different languages, on a library ported to another language, or on a code base where two teams solved one problem twice, and review each pair before you act on it.