docling-project/docling
Get your documents ready for gen AI
Python★ 67,196+129 that day#6 on trendingC 60.9 health25.5% duplicated
C60.9/100
Duplication66
5.1% in python, latex, apex, bash, perlDead code83
2%Complexity44
60.3%Most complex files
- docling/backend/html_backend.py1198 cx
- docling/backend/msword_backend.py749 cx
- docling/backend/opendocument_backend.py445 cx
- tests/test_service_client_sdk_unit.py402 cx
- docling/backend/xml/uspto_backend.py361 cx
- docling/service_client/client.py329 cx
- tests/data/latex/sources/2501.00089/aastex631.cls326 cx
- docling/backend/xml/jats_backend.py299 cx
- docling/backend/msexcel_backend.py288 cx
- tests/test_backend_msword.py245 cx
Dead code 1.99% unused · 163 findings
- docling/experimental/pipeline/threaded_layout_vlm_pipeline.pyunused-file
- docling/models/stages/ocr/_nemotron_ocr_model.pyunused-file
- docling/document_extractor.pyunused-file
- docling/models/extraction/nuextract_transformers_model.pyunused-file
- docling/pipeline/extraction_vlm_pipeline.pyunused-file
- docling/models/extraction/transformers_extraction_model.pyunused-file
- docling/models/extraction/prompt_utils.pyunused-file
- docling/pipeline/base_extraction_pipeline.pyunused-file
- docling/models/stages/ocr/auto_ocr_model.py:EngineTypeunused-import
- docling/models/stages/ocr/auto_ocr_model.py:RapidOCRunused-import
1,115files
581,778lines
2,847,099tokens
4,272clones
148,359duplicated lines (25.5%)
527,855duplicated tokens (18.54%)
By format
| Format | Files | Lines | Clones | Duplicated lines | Duplication |
|---|---|---|---|---|---|
| json | 208 | 344,361 | 3,448 | 135,190 | 39.26% |
| python | 480 | 156,541 | 597 | 8,788 | 5.61% |
| apex | 4 | 9,540 | 78 | 1,110 | 11.64% |
| markdown | 267 | 31,456 | 77 | 1,566 | 4.98% |
| markup | 40 | 4,498 | 37 | 1,240 | 27.57% |
| yaml | 25 | 4,071 | 10 | 153 | 3.76% |
| latex | 24 | 19,200 | 9 | 50 | 0.26% |
| txt | 17 | 2,384 | 8 | 58 | 2.43% |
| toml | 5 | 2,284 | 6 | 192 | 8.41% |
| bash | 27 | 4,504 | 1 | 5 | 0.11% |
Largest clones
- 5168 lines · 17,552 tokens jsontests/data/xlsx/groundtruth/xlsx_01.xlsx.json:12–5179⇆tests/data/xlsx/groundtruth/xlsx_04_inflated.xlsx.json:12–5179
- 1309 lines · 4,563 tokens jsontests/data/ppt/groundtruth/legacy_sample.ppt.json:1021–2329⇆tests/data/pptx/groundtruth/powerpoint_sample.pptx.json:1021–2329
- 276 lines · 3,151 tokens apextests/data/latex/sources/2501.00089/aastex631.cls:3817–4092⇆tests/data/latex/sources/2501.00089/aastex631.cls:4124–4399
- 276 lines · 3,151 tokens apextests/data/latex/sources/2501.00089/aastex631.cls:3817–4092⇆tests/data/latex/sources/2501.00089/aastex631.cls:4440–4715
- 833 lines · 2,896 tokens jsontests/data/odf/groundtruth/odf_presentation_01.odp.json:2981–3813⇆tests/data/ppt/groundtruth/legacy_sample.ppt.json:1494–2326
- 984 lines · 2,811 tokens jsondocs/examples/data/normal_4pages.json:280–1263⇆tests/data/pdf/groundtruth/normal_4pages.json:305–1288
- 734 lines · 2,564 tokens jsontests/data/xbrl/groundtruth/grve_10q_htm.xml.json:1387–2120⇆tests/data/xbrl/groundtruth/grve_10q_htm.xml.json:2315–3048
- 730 lines · 2,536 tokens jsontests/data/xlsx/groundtruth/xlsx_03_chartsheet.xlsx.json:766–1495⇆tests/data/xlsx/groundtruth/xlsx_comments.xlsx.json:377–1106
- 678 lines · 2,282 tokens jsondocs/examples/data/normal_4pages.json:2685–3362⇆tests/data/pdf/groundtruth/normal_4pages.json:2683–3360
- 68 lines · 1,808 tokens markdowntests/data/xbrl/groundtruth/mlac-20251231.xml.md:87–154⇆tests/data/xbrl/groundtruth/mlac-20251231.xml.md:210–277
Trending appearances
| Day | Rank | Stars | Files | Lines | Clones | Duplication | Health | Commit |
|---|---|---|---|---|---|---|---|---|
| Sep 20, 2026 | #6 | 67,196 (+129) | 1,115 | 581,778 | 4,272 | 25.5% | C 60.9 | 890dd42 |
Want this for your own project? Install jscpd and run jscpd .