apify/crawlee

▲ 26 stars today★ 26,086⑂ 1,697

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs.

About apify/crawlee

apify/crawlee is an open-source project on GitHub, mainly written in TypeScript. Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. It currently holds 26,086 stars and 1,697 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #59 with 26 new stars today.

GitHub Repository Details

Repository apify/crawlee · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

https://github.com/apify/crawlee/blob/HEAD/Crawlee
A web scraping and browser automation library

https://github.com/apify/crawlee/blob/HEAD/apify%2Fcrawlee | Trendshift

https://github.com/apify/crawlee/blob/HEAD/NPM latest version https://github.com/apify/crawlee/blob/HEAD/Downloads https://github.com/apify/crawlee/blob/HEAD/Chat on discord https://github.com/apify/crawlee/blob/HEAD/Build Status

Crawlee covers your crawling and scraping end-to-end and helps you build reliable scrapers. Fast.

Your crawlers will appear human-like and fly under the radar of modern bot protections even with the default configuration. Crawlee gives you the tools to crawl the web for links, scrape data, and store it to disk or cloud while staying configurable to suit your project's needs.

Crawlee is available as the crawlee NPM package.

👉 View full documentation, guides and examples on the Crawlee project website 👈
Do you prefer 🐍 Python instead of JavaScript? 👉 Checkout Crawlee for Python 👈.

Installation

We recommend visiting the Introduction tutorial in Crawlee documentation for more information.

Crawlee requires Node.js 22.13 or higher.

With Crawlee CLI

The fastest way to try Crawlee out is to use the Crawlee CLI and choose the Getting started example. The CLI will install all the necessary dependencies and add boilerplate code for you to play with.

npx crawlee create my-crawler
cd my-crawler
npm start

Manual installation

If you prefer adding Crawlee into your own project, try the example below. Because it uses PlaywrightCrawler we also need to install Playwright. It's not bundled with Crawlee to reduce install size.
npm install crawlee playwright
import { PlaywrightCrawler, Dataset } from 'crawlee';

// PlaywrightCrawler crawls the web using a headless // browser controlled by the Playwright library. const crawler = new PlaywrightCrawler({ // Use the requestHandler to process each of the crawled pages. async requestHandler({ request, page, enqueueLinks, log }) { const title = await page.title(); log.info(Title of ${request.loadedUrl} is '${title}');

// Save results as JSON to ./storage/datasets/default await Dataset.pushData({ title, url: request.loadedUrl });

// Extract links from the current page // and add them to the crawling queue. await enqueueLinks(); }, // Uncomment this option to see the browser window. // headless: false, });

// Add first URL to the queue and start the crawl. await crawler.run(['https://crawlee.dev']);

By default, Crawlee stores data to ./storage in the current working directory. You can override this directory via Crawlee configuration. For details, see Configuration guide, Request storage and Result storage.

Installing pre-release versions

We provide automated beta builds for every merged code change in Crawlee. You can find them in the npm list of releases. If you want to test new features or bug fixes before we release them, feel free to install a beta build like this:

npm install crawlee@next

If you also use the Apify SDK, you need to specify dependency overrides in your package.json file so that you don't end up with multiple versions of Crawlee installed:

{
    "overrides": {
       "apify": {
           "@crawlee/core": "$crawlee",
           "@crawlee/types": "$crawlee",
           "@crawlee/utils": "$crawlee"
       }
    }
}

🛠 Features

👾 HTTP crawling

💻 Real browser crawling

Usage on the Apify platform

Crawlee is open-source and runs anywhere, but since it's developed by Apify, it's easy to set up on the Apify platform and run in the cloud. Visit the Apify SDK website to learn more about deploying Crawlee to the Apify platform.

Support

If you find any bug or issue with Crawlee, please submit an issue on GitHub. For questions, you can ask on Stack Overflow, in GitHub Discussions or you can join our Discord server.

Contributing

Your code contributions are welcome, and you'll be praised to eternity! If you have any ideas for improvements, either submit an issue or create a pull request. For contribution guidelines and the code of conduct, see CONTRIBUTING.md.

License

This project is licensed under the Apache License 2.0 - see the LICENSE.md file for details.

GitHub Stars & Activity

26,086Stars
1,697Forks
0Open issues
TypeScriptLanguage

GitHub Popularity

GitHub stars26,086
Forks1,697
Open issues0
Primary languageTypeScript
License-
Stars gained today26
Created-
Last pushed-

Trending History

Daily boardrank #59 · ▲ 26 stars

Related AI Projects

1

thedotmack / claude-mem

TypeScript★ 98,943⑂ 8,669▲ 838 stars
→
2

morluto / rea

TypeScript★ 41,818⑂ 6,532▲ 15,335 stars
→
3

Vincentwei1021 / video-shotcraft

TypeScript★ 11,025⑂ 984▲ 161 stars
→
4

thesysdev / openui

TypeScript★ 10,556⑂ 713▲ 363 stars
→
5

makecindy / cindy

TypeScript★ 2,970⑂ 455▲ 24 stars
→
6

alsk1992 / CloddsBot

TypeScript★ 2,944⑂ 347▲ 27 stars
→
7

PurpleDoubleD / locally-uncensored

TypeScript★ 2,116⑂ 341▲ 70 stars
→
8

obra / superpowers

Shell★ 296,844⑂ 26,508▲ 397 stars
→

More AI Rankings