Reference

Rust crates

The library crates behind the jscpd binary, what each one exposes, a compiling example for each, and their current versions.

The engine of jscpd is a Cargo workspace of library crates on crates.io, and the command-line tool is the jscpd crate (cargo install jscpd gives you the jscpd and cpd binaries). The CLI crate's own library target is internal and changes without notice; a Rust program that embeds detection depends on the engine crates below. From any other language, run the binary with the JSON reporter or talk to it over MCP; the Node.js API of v4 lives on the master-v4 branch.

CrateVersionWhat it holds
cpd-finder0.1.20The pipeline: walking files, tokenizing, detecting, statistics, git blame. The entry point
cpd-core0.1.20The data model, the rolling-hash detector, gap merging, function similarity, and the summary, health and history types
cpd-tokenizer0.1.19Tokenizers for the 224 formats, with oxc for JavaScript and TypeScript and per-block tokenizing for Vue, Svelte, Astro, Markdown and Razor. No file or network access
cpd-reporter0.1.21The 15 reporters and the rule ids they share
cpd-semantic0.1.1--semantic: function extraction, embeddings, calibrated thresholds and a vector cache, as a clone pass for the finder
basta0.3.0The dead-code engine behind --dead-code
jscpd5.4.0The CLI

The examples below compile against these versions with the Cargo.toml lines shown; cargo check against the published crates of jscpd 5.4.0 confirmed them.

cpd-finder

cpd_finder::orchestrate runs the whole pipeline. RunConfig mirrors the CLI flags (min_tokens, min_lines, max_gap_lines, similarity, mode, formats, ignore, cross_formats, kinds, workers, the Type-2 switches and more), run returns the clones, the statistics and the sources it read, and run_excluding leaves out folders. run fails only when a clone pass fails, and only --semantic adds one, so a run without passes can be unwrapped. The source_id of a fragment comes back as an absolute path; the CLI strips the scan root before it prints.

Cargo.toml
[dependencies]
cpd-finder = "0.1.20"
src/main.rs
use std::path::PathBuf;

use cpd_finder::orchestrate::{RunConfig, run};

fn main() {
    let config = RunConfig {
        paths: vec![PathBuf::from("src")],
        min_tokens: 50,
        min_lines: 5,
        ..Default::default()
    };
    let result = run(&config).expect("a run without passes cannot fail");
    println!("{} clones", result.statistics.total.clones);
    println!("{:.1}% duplication", result.statistics.total.percentage);
    for clone in &result.clones {
        println!(
            "{}:{} ~ {}:{} ({:?}, {} tokens)",
            clone.fragment_a.source_id,
            clone.fragment_a.start.line,
            clone.fragment_b.source_id,
            clone.fragment_b.start.line,
            clone.kind,
            clone.token_count,
        );
    }
}

On a folder with two copies of one file this prints 1 clones, 50.0% duplication and one line that ends in (Exact, 124 tokens). The other modules are walker (WalkConfig and file discovery that honors .gitignore), statistics::compute(sources, clones), blame (BlameMap) and pass (ClonePass, the trait a pass such as --semantic implements).

cpd-core

Types and algorithms with no I/O. cpd_core::models holds CpdClone (two Fragments, token_count, kind, similarity, is_new), CloneKind, Statistics and StatRow (the shape of the JSON report's statistics), SourceFile and Token. cpd_core::detect holds the rolling-hash detector (detect, detect_with_options, detect_prepared) and merge_gapped_clones behind --max-gap-lines; cpd_core::hash the Rabin-Karp window hash; cpd_core::similarity the function similarity of --similarity; summary, health, history and deadcode the types of the other modes.

Cargo.toml
[dependencies]
cpd-core = "0.1.20"
cpd-tokenizer = "0.1.19"
src/main.rs
use cpd_core::detect::detect;
use cpd_core::models::SourceFile;
use cpd_tokenizer::tokenizer::{Mode, tokenize};

fn main() {
    let source = "function area(w, h) {\n  return w * h;\n}\n";
    let files: Vec<SourceFile> = ["a.js", "b.js"]
        .iter()
        .map(|id| SourceFile {
            id: id.to_string(),
            format: "javascript".to_string(),
            tokens: tokenize("javascript", source, Mode::Mild),
            bytes: source.len() as u64,
        })
        .collect();
    let clones = detect(&files, 5);
    println!("{} clones", clones.len());
}

This prints 1 clones. detect takes the minimum tokens; detect_with_options adds the minimum lines and the path filters of --skip-local and --skip-isolated.

cpd-tokenizer

cpd_tokenizer::tokenizer::tokenize(format, source, mode) returns the tokens a reporter shows; tokenize_to_detection returns the hashed stream the detector consumes, with TokenizeOptions carrying the mode, case folding, ignore ranges and the Type-2 switches. formats knows the 224 formats and their extensions, functions extracts JavaScript and TypeScript functions for --similarity, and sfc, markdown, razor and embedded split host files into blocks. The crate reads no files and opens no sockets, and CI enforces that.

Cargo.toml
[dependencies]
cpd-tokenizer = "0.1.19"
src/main.rs
use cpd_tokenizer::tokenizer::{Mode, tokenize};

fn main() {
    for token in tokenize("javascript", "const total = price * 2;", Mode::Mild) {
        println!("{:?} {:?} at {}:{}", token.kind, token.value, token.start.line, token.start.column);
    }
}
Keyword "const" at 1:0
Identifier "total" at 1:6
Operator "=" at 1:12
Identifier "price" at 1:14
Operator "*" at 1:20
Literal "2" at 1:22
Punctuation ";" at 1:23

Mode::Mild drops whitespace, Mode::Weak drops comments as well, and Mode::Strict keeps every token.

cpd-reporter

create_reporter(name, &options) returns any of the 15 reporters by its CLI name, aliases included. ReporterOptions carries the output directory, threshold, blame, sarif_error_tokens and the tool version stamped into SARIF and HTML; ReportContext wraps the statistics, the duration and the optional summary and history. cpd_reporter::rules has the rule ids (jscpd/duplicate-code and the others) that SARIF, CodeClimate and the language server share. A format of your own implements the Reporter trait.

Cargo.toml
[dependencies]
cpd-finder = "0.1.20"
cpd-reporter = "0.1.21"
src/main.rs
use std::path::PathBuf;
use std::time::Duration;

use cpd_finder::orchestrate::{RunConfig, run};
use cpd_reporter::{ReportContext, ReporterOptions, create_reporter};

fn main() {
    let config = RunConfig { paths: vec![PathBuf::from("src")], ..Default::default() };
    let result = run(&config).unwrap();
    let options = ReporterOptions::new(PathBuf::from("report"));
    let reporter = create_reporter("json", &options).expect("json is a built-in reporter");
    let ctx = ReportContext::new(&result.statistics, Duration::ZERO);
    reporter.report(&result.clones, &ctx, &options.output_dir).unwrap();
}

This prints JSON report saved to report/jscpd-report.json and writes the same file as jscpd src -r json -o report.

cpd-semantic

The crate behind --semantic: function extraction for Python, Rust and, through tree-sitter, C, C++, C#, Go, Java, Kotlin, PHP, Ruby, Scala and Swift (JavaScript and TypeScript come from cpd-tokenizer), embeddings from CodeRankEmbed or jina-embeddings-v2-base-code run in-process with candle or taken from an OpenAI-compatible API, the thresholds calibrated for nine models, a vector cache and the pairing rule. It plugs into the finder as a ClonePass (SemanticPass), so the other crates build without its dependencies. Unlike them, it reads and writes the user cache directory and makes network calls for the model download and the API. The repository's docs/api.md has the snippet that adds the pass to a RunConfig, and docs.rs has the API.

basta

--dead-code, and the dead-code section of --dashboard and --health, run the basta crate: unused files, exports, symbols and imports across JavaScript, TypeScript and Python. It is a separate project with releases of its own; Dead code covers the jscpd side.

Versions and features

  • The engine crates version on their own, 0.1.x today, while the CLI is 5.4.0. One release publishes all of them from one workspace, so the versions in the table belong together.
  • Minimum supported Rust version 1.96, edition 2024, MIT license.
  • Parallelism comes from rayon: tokenizing and detection run per format group on a local thread pool, and RunConfig::workers (--workers) sets its size.
  • Cross-format detection (cross_formats, --cross-formats) and the embedded blocks of Vue, Svelte, Astro, Markdown and Razor files are part of the tokenizer and the finder, so a library user gets them without extra setup.
  • API documentation is on docs.rs: cpd-finder, cpd-core, cpd-tokenizer, cpd-reporter, cpd-semantic.
  • Coming from the v4 @jscpd/* npm packages? The migration guide maps each one to its crate.
  • Installation for cargo install jscpd and cargo binstall jscpd.
  • Agents for the MCP server built into the binary.
  • JSON reporter for the report a program in another language reads.