Web scraping

Published entries across all sections carrying the “Web scraping” tag, newest first by publication date on this site.

3 entries

  1. APIsProducts

    CapSolver

    An AI and computer-vision powered automated CAPTCHA recognition service accelerating web data extraction pipelines.

    #CAPTCHA#Web scraping#OCR#API#Extension

  2. Developer toolsGitHub

    Open Lovable: clone any webpage into a React app

    An open-source sample app from the Firecrawl team that scrapes a webpage and hands it to an LLM to rebuild as a runnable React app, with conversational iteration, a choice of four models, and cloud sandbox previews—handy for replicating a frontend or spinning up a prototype fast.

    #Website cloning#Web scraping#Code generation#Prototyping#Open source

  3. ScrapingGitHub

    AnyCrawl: A Web Scraping Toolkit That Turns Pages into LLM-Ready Data

    A self-hostable, TypeScript web scraping toolkit that covers everything from static pages to JS-rendered pages across three engines, with bulk search-result collection, full-site crawling, and LLM structured extraction, all outputting LLM-friendly data.

    #Web scraping#AI crawler#Extraction#RAG#SERP API#Self-hosted#Open source