Discover / RAG & Knowledge
Unstructured
by Unstructured-IOHTML
Preprocess and parse documents (PDF, HTML, docs) for LLM ingestion.
Maturity: experimental because latest release 0.24.1 is pre 1.0. Derived from release and commit history, not a rating.
- Stars
- 15k
- Forks
- 1.3k
- Downloads / mo
- 5.3M
- Last commit
- 2026-08-02
- License
- Apache-2.0
- Open issues
- 275
Market and trust evidence
Edition not yet matchedNo exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.
Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.
In practice
Written by AI from this repository’s README · medium confidenceFeeding real documents to an LLM means writing brittle parsers for every file type you receive.
Use it when
Use it when you need a document ingestion and partitioning step in front of an LLM or retrieval pipeline.
Not the right pick when
It is a preprocessing library, not a retrieval or serving system, so you still need somewhere to store the output.
Capabilities
- ingest and pre-process images and text documents such as PDFs, HTML and Word docs
- modular functions and connectors forming one ingestion system
- transform unstructured data into structured outputs for LLMs
- Unstructured Transform available to agents as an MCP server
- adaptable across different platforms
Cost: Free and open source
Install
Derived from the published package name in the repository, not from a model.
Video walkthroughs
STOP Fighting Messy PDFs! Unstructured.io is the RAG Preprocessing Tool Every AI Developer NEEDS
Build a Multi-Modal RAG Pipeline That Actually Works (Unstructured.io)
Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.
What the repository ships
Detected from the actual files in the repository root.
Latest release 0.24.1
Published 2026-07-11
What's Changed
- fix: sanitize v2 HTML output to prevent stored XSS (GHSA-v5mq-3xhg-98m9) by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4394
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.24.0...0.24.1
Tags
README
<h3 align="center">
<img
src="https://raw.githubusercontent.com/Unstructured-IO/unstructured/main/img/unstructured_logo.png"
height="200"
</h3>
<div align="center">
<a href="https://github.com/Unstructured-IO/unstructured/blob/main/LICENSE.md">https://pypi.python.org/pypi/unstructured/</a>
<a href="https://pypi.python.org/pypi/unstructured/">https://pypi.python.org/pypi/unstructured/</a>
<a href="https://GitHub.com/unstructured-io/unstructured/graphs/contributors">https://GitHub.com/unstructured-io/unstructured.js/graphs/contributors</a>
<a href="https://github.com/Unstructured-IO/unstructured/blob/main/CODE_OF_CONDUCT.md">code_of_conduct.md </a>
<a href="https://GitHub.com/unstructured-io/unstructured/releases">https://GitHub.com/unstructured-io/unstructured.js/releases</a>
<a href="https://pypi.python.org/pypi/unstructured/">https://github.com/Naereen/badges/</a>
<a
href="https://www.phorm.ai/query?projectId=34efc517-2201-4376-af43-40c4b9da3dc5">
<img src="https://img.shields.io/badge/Phorm-Ask_AI-%23F2777A.svg?&logo=data:image/svg+xml;base64,PHN2ZyB3aWR0aD0iNSIgaGVpZ2h0PSI0IiBmaWxsPSJub25lIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPgogIDxwYXRoIGQ9Ik00LjQzIDEuODgyYTEuNDQgMS40NCAwIDAgMS0uMDk4LjQyNmMtLjA1LjEyMy0uMTE1LjIzLS4xOTIuMzIyLS4wNzUuMDktLjE2LjE2NS0uMjU1LjIyNmExLjM1MyAxLjM1MyAwIDAgMS0uNTk1LjIxMmMtLjA5OS4wMTItLjE5Mi4wMTQtLjI3OS4wMDZsLTEuNTkzLS4xNHYtLjQwNmgxLjY1OGMuMDkuMDAxLjE3LS4xNjkuMjQ2LS4xOTFhLjYwMy42MDMgMCAwIDAgLjItLjEwNi41MjkuNTI5IDAgMCAwIC4xMzgtLjE3LjY1NC42NTQgMCAwIDAgLjA2NS0uMjRsLjAyOC0uMzJhLjkzLjkzIDAgMCAwLS4wMzYtLjI0OS41NjcuNTY3IDAgMCAwLS4xMDMtLjIuNTAyLjUwMiAwIDAgMC0uMTY4LS4xMzguNjA4LjYwOCAwIDAgMC0uMjQtLjA2N0wyLjQzNy43MjkgMS42MjUuNjcxYS4zMjIuMzIyIDAgMCAwLS4yMzIuMDU4LjM3NS4zNzUgMCAwIDAtLjExNi4yMzJsLS4xMTYgMS40NS0uMDU4LjY5Ny0uMDU4Ljc1NEwuNzA1IDRsLS4zNTctLjA3OUwuNjAyLjkwNkMuNjE3LjcyNi42NjMuNTc0LjczOS40NTRhLjk1OC45NTggMCAwIDEgLjI3NC0uMjg1Ljk3MS45NzEgMCAwIDEgLjMzNy0uMTRjLjExOS0uMDI2LjIyNy0uMDM0LjMyNS0uMDI2TDMuMjMyLjE2Yy4xNTkuMDE0LjMzNi4wMy40NTkuMDgyYTEuMTczIDEuMTczIDAgMCAxIC41NDUuNDQ3Yy4wNi4wOTQuMTA5LjE5Mi4xNDQuMjkzYTEuMzkyIDEuMzkyIDAgMCAxIC4wNzguNThsLS4wMjkuMzJaIiBmaWxsPSIjRjI3NzdBIi8+CiAgPHBhdGggZD0iTTQuMDgyIDIuMDA3YTEuNDU1IDEuNDU1IDAgMCAxLS4wOTguNDI3Yy0uMDUuMTI0LS4xMTQuMjMyLS4xOTIuMzI0YTEuMTMgMS4xMyAwIDAgMS0uMjU0LjIyNyAxLjM1MyAxLjM1MyAwIDAgMS0uNTk1LjIxNGMtLjEuMDEyLS4xOTMuMDE0LS4yOC4wMDZsLTEuNTYtLjEwOC4wMzQtLjQwNi4wMy0uMzQ4IDEuNTU5LjE1NGMuMDkgMCAuMTczLS4wMS4yNDgtLjAzM2EuNjAzLjYwMyAwIDAgMCAuMi0uMTA2LjUzMi41MzIgMCAwIDAgLjEzOS0uMTcyLjY2LjY2IDAgMCAwIC4wNjQtLjI0MWwuMDI5LS4zMjFhLjk0Ljk0IDAgMCAwLS4wMzYtLjI1LjU3LjU3IDAgMCAwLS4xMDMtLjIwMi41MDIuNTAyIDAgMCAwLS4xNjgtLjEzOC42MDUuNjA1IDAgMCAwLS4yNC0uMDY3TDEuMjczLjgyN2MtLjA5NC0uMDA4LS4xNjguMDEtLjIyMS4wNTUtLjA1My4wNDUtLjA4NC4xMTQtLjA5Mi4yMDZMLjcwNSA0IDAgMy45MzhsLjI1NS0yLjkxMUExLjAxIDEuMDEgMCAwIDEgLjM5My41NzIuOTYyLjk2MiAwIDAgMSAuNjY2LjI4NmEuOTcuOTcgMCAwIDEgLjMzOC0uMTRDMS4xMjIuMTIgMS4yMy4xMSAxLjMyOC4xMTlsMS41OTMuMTRjLjE2LjAxNC4zLjA0Ny40MjMuMWExLjE3IDEuMTcgMCAwIDEgLjU0NS40NDhjLjA2MS4wOTUuMTA5LjE5My4xNDQuMjk1YTEuNDA2IDEuNDA2IDAgMCAxIC4wNzcuNTgzbC0uMDI4LjMyMloiIGZpbGw9IndoaXRlIi8+CiAgPHBhdGggZD0iTTQuMDgyIDIuMDA3YTEuNDU1IDEuNDU1IDAgMCAxLS4wOTguNDI3Yy0uMDUuMTI0LS4xMTQuMjMyLS4xOTIuMzI0YTEuMTMgMS4xMyAwIDAgMS0uMjU0LjIyNyAxLjM1MyAxLjM1MyAwIDAgMS0uNTk1LjIxNGMtLjEuMDEyLS4xOTMuMDE0LS4yOC4wMDZsLTEuNTYtLjEwOC4wMzQtLjQwNi4wMy0uMzQ4IDEuNTU5LjE1NGMuMDkgMCAuMTczLS4wMS4yNDgtLjAzM2EuNjAzLjYwMyAwIDAgMCAuMi0uMTA2LjUzMi41MzIgMCAwIDAgLjEzOS0uMTcyLjY2LjY2IDAgMCAwIC4wNjQtLjI0MWwuMDI5LS4zMjFhLjk0Ljk0IDAgMCAwLS4wMzYtLjI1LjU3LjU3IDAgMCAwLS4xMDMtLjIwMi41MDIuNTAyIDAgMCAwLS4xNjgtLjEzOC42MDUuNjA1IDAgMCAwLS4yNC0uMDY3TDEuMjczLjgyN2MtLjA5NC0uMDA4LS4xNjguMDEtLjIyMS4wNTUtLjA1My4wNDUtLjA4NC4xMTQtLjA5Mi4yMDZMLjcwNSA0IDAgMy45MzhsLjI1NS0yLjkxMUExLjAxIDEuMDEgMCAwIDEgLjM5My41NzIuOTYyLjk2MiAwIDAgMSAuNjY2LjI4NmEuOTcuOTcgMCAwIDEgLjMzOC0uMTRDMS4xMjIuMTIgMS4yMy4xMSAxLjMyOC4xMTlsMS41OTMuMTRjLjE2LjAxNC4zLjA0Ny40MjMuMWExLjE3IDEuMTcgMCAwIDEgLjU0NS40NDhjLjA2MS4wOTUuMTA5LjE5My4xNDQuMjk1YTEuNDA2IDEuNDA2IDAgMCAxIC4wNzcuNTgzbC0uMDI4LjMyMloiIGZpbGw9IndoaXRlIi8+Cjwvc3ZnPgo=" />
</a>
</div>
<div>
<p align="center">
<a
href="https://short.unstructured.io/pzw05l7">
<img src="https://img.shields.io/badge/JOIN US ON SLACK-4A154B?style=for-the-badge&logo=slack&logoColor=white" />
</a>
<a href="https://www.linkedin.com/company/unstructuredio/">
<img src="https://img.shields.io/badge/LinkedIn-0077B5?style=for-the-badge&logo=linkedin&logoColor=white" />
</a>
</div>
<h2 align="center">
<p>Open-Source Pre-Processing Tools for Unstructured Data</p>
</h2>
The unstructured library provides open-source components for ingesting and pre-processing images and text documents, such as PDFs, HTML, Word docs, and many more. The use cases of unstructured revolve around streamlining and optimizing the data processing workflow for LLMs. unstructured modular functions and connectors form a cohesive system that simplifies data ingestion and pre-processing, making it adaptable to different platforms and efficient in transforming unstructured data into structured outputs.
Unstructured Transform MCP — Document Processing for your Agents
Unstructured Transform brings production-grade document processing to your agents as an MCP server. It gives them the ability to turn 60+ file types into structured data that's ready for your applications, vector databases, and any downstream processes by parsing, enriching, chunking, and embedding files directly inside their current session.
Setup Steps for Your Agent
- Pick your MCP client. Transform works with virtually any MCP-compatible host or agent framework — Claude Code, Cursor, Codex CLI and more.
- Add the Transform MCP server to your client's MCP configuration (via the CLI
mcp addcommand or the client's MCP settings/config file, depending on the tool).
- Authenticate once when your client prompts you. Sign in, and the Transform tools become available to your agent on its next message.
- Point your agent at a file. Drag and drop or reference a local file or URL. Transform handles 60+ formats (PDFs, emails, images, scanned files, and more).
- Describe what you need in plain language. Tell the agent your intent (e.g. "parse and chunk this contract for a vector store") and Transform partitions, enriches, chunks, and embeds the file, returning structured data ready to use.
15,000 free pages a month, 3 cents per page after!
📄 Full docs: https://docs.unstructured.io/transform/overview
Unstructured Pipelines
Ready to move your data processing pipeline to production, and take advantage of advanced features? Check out Unstructured Pipelines. In addition to better processing performance, take advantage of chunking, embedding, and image and table enrichment generation, all from a low code UI or an API. Request a demo from our sales team to learn more about how to get started.
:eight_pointed_black_star: Quick Start
There are several ways to use the unstructured library:
- Run the library in a container or
- Install the library
- For installation with
condaon Windows system, please refer to the documentation
Run the library in a container
The following instructions are intended to help you get up and running using Docker to interact with unstructured.
See here if you don't already have docker installed on your machine.
NOTE: we build multi-platform images to support both x86_64 and Apple silicon hardware. docker pull should download the corresponding image for your architecture, but you can specify with --platform (e.g. --platform linux/amd64) if needed.
We build Docker images for all pushes to main. We tag each image with the corresponding short commit hash (e.g. fbc7a69) and the application version (e.g. 0.5.5-dev1). We also tag the most recent image with latest. To leverage this, docker pull from our image repository.
docker pull downloads.unstructured.io/unstructured-io/unstructured:latest
Once pulled, you can create a container from this image and shell to it.
# create the container
docker run -dt --name unstructured downloads.unstructured.io/unstructured-io/unstructured:latest
# this will drop you into a bash shell where the Docker image is running
docker exec -it unstructured bash
You can also build your own Docker image. Note that the base image is wolfi-base, which is
updated regularly. If you are building the image locally, it is possible docker-build could
fail due to upstream changes in wolfi-base.
If you only plan on parsing one type of data you can speed up building the image by commenting out some
of the packages/requirements necessary for other data types. See Dockerfile to know which lines are necessary
for your use case.
make docker-build
# this will drop you into a bash shell where the Docker image is running
make docker-start-bash
Once in the running container, you can try things directly in Python interpreter's interactive mode.
# this will drop you into a python console so you can run the below partition functions
python3
>>> from unstructured.partition.pdf import partition_pdf
>>> elements = partition_pdf(filename="example-docs/layout-parser-paper-fast.pdf")
>>> from unstructured.partition.text import partition_text
>>> elements = partition_text(filename="example-docs/fake-text.txt")
Installing the library
Use the following instructions to get up and running with unstructured and test your
installation.
- Install the Python SDK to support all document types with
pip install "unstructured[all-docs]" - For plain text files, HTML, XML, JSON and Emails that do not require any extra dependencies, you can run
pip install unstructured - To process other doc types, you can install the extras required for those documents, such as
pip install "unstructured[docx,pptx]" - Install the following system dependencies if they are not already available on your system.
Depending on what document types you're parsing, you may not need all of these.
libmagic-dev(filetype detection)poppler-utils(images and PDFs)tesseract-ocr(images and PDFs, installtesseract-langfor additional language support)libreoffice(MS Office docs)pandocis bundled automatically via thepypandoc-binaryPython package (no system install needed)
- For suggestions on how to install on the Windows and to learn about dependencies for other features, see the
installation documentation here.
At this point, you should be able to run the following code:
from unstructured.partition.auto import partition
elements = partition(filename="example-docs/eml/fake-email.eml")
print("\n\n".join([str(el) for el in elements]))
Installation Instructions for Local Development
The following instructions are intended to help you get up and running with unstructured
locally if you are planning to contribute to the project.
This project uses uv for dependency management. Install it first:
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
Then install all dependencies (base, extras, dev, test, and lint groups):
make install
This runs uv sync --locked --all-extras --all-groups, which creates a virtual environment
and installs everything in one step. No need to manually create or activate a virtualenv.
To install only specific document-type extras:
uv sync --extra pdf
uv sync --extra csv --extra docx
To update the lock file after changing dependencies in pyproject.toml:
make lock
- Optional:
- To install extras for processing images and PDFs locally, run
uv sync --extra pdf --extra image. - For processing image files,
tesseractis required. See here for installation instructions. - For processing PDF files,
tesseractandpopplerare required. The pdf2image docs have instructions on installingpoppleracross various platforms.
Additionally, if you're planning to contribute to unstructured, we provide you an optional pre-commit configuration
file to ensure your code matches the formatting and linting standards used in unstructured.
If you'd prefer not to have code changes auto-tidied before every commit, you can use make check to see
whether any linting or formatting changes should be applied, and make tidy to apply them.
If using the optional pre-commit, you'll just need to install the hooks with pre-commit install since the
pre-commit package is
Truncated. Read the full README on GitHub ↗