Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
-
Updated
Sep 29, 2026 - HTML
Web scraping is the process of programmatically retrieving web pages and extracting structured data from them for analysis, storage, or reuse. It powers price monitoring, search indexing, market research, and training-data collection.
Scraping ranges from plain HTTP requests and HTML parsing to full browser automation for JavaScript-rendered and anti-bot-protected pages. The ecosystem includes request libraries, HTML parsers, headless browsers, and cloud extraction platforms.
Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
Extract any website's complete design system with one command. DTCG tokens, semantic+primitive+composite, MCP server for Claude Code/Cursor/Windsurf, multi-platform emitters (iOS SwiftUI, Android Compose, Flutter, WordPress), Tailwind v4, Figma variables, shadcn/ui, CSS health audit, WCAG remediation, Chrome extension. MIT, Playwright, Node 20+.
Learn Python for the next 30 (or so) Days.
NBA Stats API via Basketball Reference
Jekyll-based static site for The Programming Historian
`scrape_linkedin` is a python package that allows you to scrape personal LinkedIn profiles & company pages - turning the data into structured json.
Learn everything web scraping with David Teather Codes on YouTube
Scrape, standardize and share public meetings from local government websites
A fork of Dragnet that also extract author, headline, date, keywords from context, as well as built in metadata extraction all in one package
The repository and website hosting the peer review process for new Programming Historian lessons
Two methods to collect real Google SERP data—a free scraper for basic use and the enterprise-grade Bright Data API for high-volume demands.
Best MangaFox Downloader Script 2026: Batch Manga Grabber Tool
Scape top GitHub repositories and users based on keywords
Pareto-optimal models for cleaning the web — fast, encoder-based main-content extraction from HTML.
A library to read a YML file with Xpath or CSS Selectors and extract data from HTML pages using them
A high-performance personal fund tracker focused on providing real-time net value estimations. It features deep stock penetration, smart reverse-calculation, and robust multi-level caching for a seamless experience. 一款专注于提供基金实时净值估算的高性能追踪看板。支持底层重仓股穿透、智能净值反向推算,并内置防御级三级缓存架构。
Open source implementation of Sova - RAG-based Web search engine using power of LLMs. Using Langchain, Ollama, HuggingFace Embeddings and scraping google search results.
Exercises, data, and more for our 2017 summer workshop (funded by the Estes Fund and in partnership with Project Jupyter and Berkeley's D-Lab)
Materials to reproduce findings in our story, "Google’s Top Search Result? Surprise! It’s Google"
Building a Concurrent Web Scraper with Python and Selenium