Discover / RAG & Knowledge

Docling

by DS4SDPython

Document conversion toolkit that parses PDFs and office files into structured formats.

Repositorystable

Maturity: stable because 2y old, v2.117.0 released 4d ago. Derived from release and commit history, not a rating.

Stars
64k
Forks
4.6k
Downloads / mo
5.5M
Last commit
2026-08-02
License
MIT
Open issues
952

Market and trust evidence

Edition not yet matched

No exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.

Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.

In practice

Written by AI from this repository’s README · high confidence

Feeding real world documents to an LLM requires layout aware parsing, table structure and OCR that generic loaders lack.

Use it when

Use it when a RAG or extraction pipeline must ingest mixed document formats and preserve reading order and tables.

Not the right pick when

Not a retrieval or indexing system on its own, and Python 3.9 is no longer supported.

Capabilities

  • parses PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, images and more
  • advanced PDF understanding covering layout, reading order and table structure
  • unified DoclingDocument representation with Markdown, HTML and JSON export
  • local execution for sensitive data and air-gapped environments
  • integrations with LangChain, LlamaIndex, Crew AI and Haystack
  • MCP server and API server deployment options

Requirements

  • Python 3.10 or higher

Cost: Free and open source

Install

Derived from the published package name in the repository, not from a model.

Video walkthroughs

Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.

What the repository ships

Ships CLAUDE.mdHas testsHas docsDocker imageCI configured

Detected from the actual files in the repository root.

Latest release v2.117.0

Published 2026-07-30

Feature

  • service datamodels: Chunking options and targets (#3857) (0877ac0)
  • vlm: Expose OpenAI logprobs as generated tokens (#3903) (c18bdf8)

Fix

  • tests: Increase tolerance for fuzzy test on bbox (#3912) (8f9f2c8)
  • ocr: Make the OCR render scale configurable instead of hardcoded (#3877) (e7f9e60)
  • pdf-outline: Use iterative walk to avoid RecursionError on deep outlines (#3855) (81a0149)
  • Skip image enrichment without pages (#3875) (00acb59)
  • odf: Skip a draw:object whose embedded part is missing (#3876) (2c3e55b)

Documentation

  • Add trivial (pass-through) chunker example (#3861) (ba8251e)

Tags

README

<p align="center">

<a href="https://github.com/docling-project/docling">

<img loading="lazy" alt="Docling" src="https://github.com/docling-project/docling/raw/main/docs/assets/docling_processing.png" width="100%"/>

</a>

</p>

Docling

<p align="center">

<a href="https://trendshift.io/repositories/17240" target="_blank"><img src="https://trendshift.io/api/badge/repositories/17240" alt="DS4SD%2Fdocling | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>

</p>

arXiv

Docs

PyPI version

PyPI - Python Version

uv

Ruff

Pydantic v2

prek

License MIT

PyPI Downloads

Docling Actor

Chat with Dosu

Discord

OpenSSF Best Practices

LF AI & Data

What is Docling ?

Docling simplifies document processing by parsing diverse formats — including advanced PDF understanding — and providing seamless integrations with the generative AI ecosystem.

Features

  • 🗂️ Parsing of [multiple document formats][supported_formats] including PDF, DOCX, PPTX, XLSX, HTML, EPUB, WAV, MP3, WebVTT, Box Notes, email formats (EML, MSG), images (PNG, TIFF, JPEG, ...), LaTeX, DocLang, plain text, and more
  • 📑 Advanced PDF understanding incl. page layout, reading order, table structure, code, formulas, image classification, and more
  • 🧬 A unified, expressive [DoclingDocument][docling_document] representation format
  • ↪️ Various [export formats][supported_formats] and options, including Markdown, HTML, WebVTT, DocLang, DocTags and lossless JSON
  • 📜 Support for several application-specific XML schemas including DocLang, USPTO patents, JATS articles, and XBRL financial reports.
  • 🔒 Local execution capabilities for sensitive data and air-gapped environments
  • 🤖 Plug-and-play [integrations][integrations] incl. LangChain, LlamaIndex, Crew AI & Haystack for agentic AI
  • 🔍 Extensive OCR support for scanned PDFs and images
  • 👓 Support for several Visual Language Models, such as (GraniteDocling)
  • 🎙️ Audio support with Automatic Speech Recognition (ASR) models
  • 🔌 Connect to any agent using the MCP server
  • 🌐 Run Docling as a service with the API server (docling-serve)
  • 💻 Simple and convenient CLI

What's new

  • 🎬 Parsing of video files (MP4, AVI, MOV, MKV, and WebM) with an ASR transcript and representative keyframes
  • 📄 Parsing of ODF (OpenDocument Format) files for text documents (.odt), spreadsheets (.ods), and presentations (.odp)
  • 💼 Parsing of XBRL (eXtensible Business Reporting Language) documents for financial reports
  • 📧 Parsing of email files (.eml, .msg)
  • 📚 Parsing of EPUB (Electronic Publication) files for e-books
  • 📝 Parsing of plain-text files (.txt, .text) and Markdown supersets (.qmd, .Rmd)
  • 📊 Chart understanding (Barchart, Piechart, LinePlot): convert them into tables or code and add detailed descriptions

Coming soon

  • 📝 Metadata extraction, including title, authors, references & language
  • 📝 Complex chemistry understanding (Molecular structures)

Quickstart

1. Install


pip install docling

Note: Python 3.9 support was dropped in docling version 2.70.0. Please use Python 3.10 or higher.

Works on macOS, Linux and Windows environments for both x86_64 and arm64 architectures.

More detailed installation instructions are available in the docs.

2. Convert a document (CLI)


docling https://arxiv.org/pdf/2206.01062

This generates a .md file in the current directory containing structured document content.

You can also use 🥚GraniteDocling and other VLMs via Docling CLI:


docling --pipeline vlm --vlm-model granite_docling https://arxiv.org/pdf/2206.01062

3. Python usage (recommended)


from docling.document_converter import DocumentConverter

source = "https://arxiv.org/pdf/2408.09869"  # a document via a local path or URL
converter = DocumentConverter()
result = converter.convert(source)
print(result.document.export_to_markdown())  # output: "## Docling Technical Report[...]"

More advanced usage and configuration options.

Documentation

Check out Docling's documentation for details on

installation, usage, concepts, recipes, extensions, and more.

Examples

Go hands-on with our examples,

demonstrating how to address different application use cases with Docling.

Integrations

To further accelerate your AI application development, check out Docling's native

integrations with popular frameworks

and tools.

Get help and support

Please feel free to connect with us using the discussion section.

Technical report

For more details on Docling's inner workings, check out the Docling Technical Report.

Contributing

Please read Contributing to Docling for details.

References

If you use Docling in your projects, please consider citing the following:


@techreport{Docling,
  author = {Deep Search Team},
  month = {8},
  title = {Docling Technical Report},
  url = {https://arxiv.org/abs/2408.09869},
  eprint = {2408.09869},
  doi = {10.48550/arXiv.2408.09869},
  version = {1.0.0},
  year = {2024}
}

License

The Docling codebase is under MIT license.

For individual model usage, please refer to the model licenses found in the original packages.

LF AI & Data

Docling is hosted as a project in the LF AI & Data Foundation.

IBM ❤️ Open Source AI

The project was started by the AI for knowledge team at IBM Research Zurich.

[supported_formats]: https://docling-project.github.io/docling/usage/supported_formats/

[docling_document]: https://docling-project.github.io/docling/concepts/docling_document/

[integrations]: https://docling-project.github.io/docling/integrations/

[extraction]: https://docling-project.github.io/docling/_generated/examples/extraction/

Related tools