Discover / AI Agents

Microsoft OmniParser GUI Agent

by microsoftJupyter Notebook

Screen parsing tool that converts UI screenshots into structured elements for vision agents.

Toolactive

Maturity: active because commit 14d ago, latest release v.2.0.1. Derived from release and commit history, not a rating.

Stars
25k
Forks
2.2k
Downloads / mo
Last commit
2026-07-20
License
CC-BY-4.0
Open issues
231

Market and trust evidence

Edition not yet matched

No exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.

Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.

In practice

Written by AI from this repository’s README · high confidence

Vision models guess at coordinates because a raw screenshot gives them no grounded, labelled interface regions.

Use it when

When building a computer use agent that needs interactable region detection and icon descriptions from screenshots.

Not the right pick when

Setup means cloning the repo, creating a conda environment and downloading weights, so it is not a packaged library.

Capabilities

  • Parses user interface screenshots into structured elements
  • Interactive region detection with a YOLOv9-E detector
  • Icon functional description captioning
  • Prediction of whether each screen element is interactable
  • Gradio demo and example notebook
  • OmniTool for controlling a Windows 11 VM

Requirements

  • A clone of the repository
  • A conda environment with Python 3.12
  • Detector and caption weights downloaded from Hugging Face

Cost: Free and open source

Video walkthroughs

Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.

What the repository ships

Has docsSecurity policy

Detected from the actual files in the repository root.

Latest release v.2.0.1

Published 2025-09-12

What's New in V2.0.1?

  • Security Updates
  • Dependencies Versioning Fixes
  • Documentation Improvements

Tags

README

OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent

<p align="center">

<img src="imgs/logo.png" alt="Logo">

</p>

<!-- <a href="https://trendshift.io/repositories/12975" target="_blank"><img src="https://trendshift.io/api/badge/repositories/12975" alt="microsoft%2FOmniParser | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a> -->

arXiv

License

📢 [Project Page] [V2 Blog Post] [Models V2] [Models V1.5] [HuggingFace Space Demo]

OmniParser is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.

News

  • [2026/7] We add a YOLOv9-E interactive region detector. Its inference-only weight is available in Hugging Face PR #37.
  • [2025/3] We support local logging of trajecotry so that you can use OmniParser+OmniTool to build training data pipeline for your favorate agent in your domain. [Documentation WIP]
  • [2025/3] We are gradually adding multi agents orchstration and improving user interface in OmniTool for better experience.
  • [2025/2] We release OmniParser V2 checkpoints. Watch Video
  • [2025/2] We introduce OmniTool: Control a Windows 11 VM with OmniParser + your vision model of choice. OmniTool supports out of the box the following large language models - OpenAI (4o/o1/o3-mini), DeepSeek (R1), Qwen (2.5VL) or Anthropic Computer Use. Watch Video
  • [2025/1] V2 is coming. We achieve new state of the art results 39.5% on the new grounding benchmark Screen Spot Pro with OmniParser v2 (will be released soon)! Read more details here.
  • [2024/11] We release an updated version, OmniParser V1.5 which features 1) more fine grained/small icon detection, 2) prediction of whether each screen element is interactable or not. Examples in the demo.ipynb.
  • [2024/10] OmniParser was the #1 trending model on huggingface model hub (starting 10/29/2024).
  • [2024/10] Feel free to checkout our demo on huggingface space! (stay tuned for OmniParser + Claude Computer Use)
  • [2024/10] Both Interactive Region Detection Model and Icon functional description model are released! Hugginface models
  • [2024/09] OmniParser achieves the best performance on Windows Agent Arena!

Install

First clone the repo, and then install environment:


cd OmniParser
conda create -n "omni" python==3.12
conda activate omni
pip install -r requirements.txt

Until Hugging Face PR #37 is merged, download the latest YOLOv9-E detector from the PR:


huggingface-cli download microsoft/OmniParser-v2.0 icon_detect_v3/model.pt \
  --revision refs/pr/37 --local-dir weights

OmniParser prefers this local weight. After the PR is merged, it will download the same weight automatically on first use. Download the caption weights into the weights folder:


   for f in icon_caption/{config.json,generation_config.json,model.safetensors}; do huggingface-cli download microsoft/OmniParser-v2.0 "$f" --local-dir weights; done
   mv weights/icon_caption weights/icon_caption_florence

<!-- ## [deprecated]

Then download the model ckpts files in: https://huggingface.co/microsoft/OmniParser, and put them under weights/, default folder structure is: weights/icon_detect, weights/icon_caption_florence, weights/icon_caption_blip2.

For v1:

convert the safetensor to .pt file.


python weights/convert_safetensor_to_pt.py

For v1.5:
download 'model_v1_5.pt' from https://huggingface.co/microsoft/OmniParser/tree/main/icon_detect_v1_5, make a new dir: weights/icon_detect_v1_5, and put it inside the folder. No weight conversion is needed.

Examples:

We put together a few simple examples in the demo.ipynb.

Gradio Demo

To run gradio demo, simply run:


python gradio_demo.py

Model Weights License

icon_detect_v3 is based on the MIT-licensed YOLOv9 implementation. Earlier Ultralytics-based icon detectors retain their original AGPL license. The caption models are under the MIT license.

📚 Citation

Our technical report can be found here.

If you find our work useful, please consider citing our work:


@misc{lu2024omniparserpurevisionbased,
      title={OmniParser for Pure Vision Based GUI Agent},
      author={Yadong Lu and Jianwei Yang and Yelong Shen and Ahmed Awadallah},
      year={2024},
      eprint={2408.00203},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2408.00203},
}

Related tools