Discover / Data & Research
AutoScraper
by alirezamikaPython
Smart automatic web scraper that learns scraping rules from example data.
Maturity: active because commit 5d ago, latest release v1.1.14. Derived from release and commit history, not a rating.
- Stars
- 7.8k
- Forks
- 800
- Downloads / mo
- 2.0k
- Last commit
- 2026-07-29
- License
- MIT
- Open issues
- 1
Market and trust evidence
Edition not yet matchedNo exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.
Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.
In practice
Written by AI from this repository’s README · high confidenceWriting and maintaining CSS or XPath selectors for every site is slow and breaks whenever the markup shifts.
Use it when
Use it when you can give a few sample values from a page and want a reusable scraper object for similar pages.
Not the right pick when
It learns from samples on a static page, so dynamic content that changes between loads needs the wanted list updated.
Capabilities
- learns scraping rules from a wanted list of sample values
- get_result_similar for related elements
- get_result_exact for the same elements in order
- accepts a URL or raw HTML content
- passes custom requests parameters such as proxies and headers
- saves and loads the built model
Requirements
- Python 3
Cost: Free and open source
Install
Derived from the published package name in the repository, not from a model.
Video walkthroughs
Automating Web Scrapping Using AutoScraper Library
The EASIEST way to do WEB SCRAPING with PYTHON (using autoscraper) - Crypto price checker
Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.
What the repository ships
Detected from the actual files in the repository root.
Latest release v1.1.14
Published 2022-07-17
No release notes provided.
Tags
README
AutoScraper: A Smart, Automatic, Fast and Lightweight Web Scraper for Python
This project is made for automatic web scraping to make scraping easy.
It gets a url or the html content of a web page and a list of sample data which we want to scrape from that page. This data can be text, url or any html tag value of that page. It learns the scraping rules and returns the similar elements. Then you can use this learned object with new urls to get similar content or the exact same element of those new pages.
Installation
It's compatible with python 3.
- Install latest version from git repository using pip:
$ pip install git+https://github.com/alirezamika/autoscraper.git
- Install from PyPI:
$ pip install autoscraper
- Install from source:
$ python setup.py install
How to use
Getting similar results
Say we want to fetch all related post titles in a stackoverflow page:
from autoscraper import AutoScraper
url = 'https://stackoverflow.com/questions/2081586/web-scraping-with-python'
# We can add one or multiple candidates here.
# You can also put urls here to retrieve urls.
wanted_list = ["What are metaclasses in Python?"]
scraper = AutoScraper()
result = scraper.build(url, wanted_list)
print(result)
Here's the output:
[
'How do I merge two dictionaries in a single expression in Python (taking union of dictionaries)?',
'How to call an external command?',
'What are metaclasses in Python?',
'Does Python have a ternary conditional operator?',
'How do you remove duplicates from a list whilst preserving order?',
'Convert bytes to a string',
'How to get line count of a large file cheaply in Python?',
"Does Python have a string 'contains' substring method?",
'Why is “1000000000000000 in range(1000000000000001)” so fast in Python 3?'
]
Now you can use the scraper object to get related topics of any stackoverflow page:
scraper.get_result_similar('https://stackoverflow.com/questions/606191/convert-bytes-to-a-string')
Getting exact result
Say we want to scrape live stock prices from yahoo finance:
from autoscraper import AutoScraper
url = 'https://finance.yahoo.com/quote/AAPL/'
wanted_list = ["124.81"]
scraper = AutoScraper()
# Here we can also pass html content via the html parameter instead of the url (html=html_content)
result = scraper.build(url, wanted_list)
print(result)
Note that you should update the wanted_list if you want to copy this code, as the content of the page dynamically changes.
You can also pass any custom requests module parameter. for example you may want to use proxies or custom headers:
proxies = {
"http": 'http://127.0.0.1:8001',
"https": 'https://127.0.0.1:8001',
}
result = scraper.build(url, wanted_list, request_args=dict(proxies=proxies))
Now we can get the price of any symbol:
scraper.get_result_exact('https://finance.yahoo.com/quote/MSFT/')
You may want to get other info as well. For example if you want to get market cap too, you can just append it to the wanted list. By using the get_result_exact method, it will retrieve the data as the same exact order in the wanted list.
Another example: Say we want to scrape the about text, number of stars and the link to issues of Github repo pages:
from autoscraper import AutoScraper
url = 'https://github.com/alirezamika/autoscraper'
wanted_list = ['A Smart, Automatic, Fast and Lightweight Web Scraper for Python', '6.2k', 'https://github.com/alirezamika/autoscraper/issues']
scraper = AutoScraper()
scraper.build(url, wanted_list)
Simple, right?
Saving the model
We can now save the built model to use it later. To save:
# Give it a file path
scraper.save('yahoo-finance')
And to load:
scraper.load('yahoo-finance')
Tutorials
- See this gist for more advanced usages.
- AutoScraper and Flask: Create an API From Any Website in Less Than 5 Minutes
Issues
Feel free to open an issue if you have any problem using the module.
Support the project
<a href="https://www.buymeacoffee.com/alirezam" target="_blank"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-black.png" alt="Buy Me A Coffee" height="45" width="163" ></a>