scrape permissions
Guide on using git-scraping with GitHub Actions to scrape data, focusing on the required permissions flag.
Guide on using git-scraping with GitHub Actions to scrape data, focusing on the required permissions flag.
Python scripts to scrape and convert SwissDelphiCenter Delphi tips from the Wayback Machine for GExperts import.
Release notes for shot-scraper 1.11, a CLI tool for website screenshots, video demos, and JavaScript scraping with minor improvements.
Release notes for shot-scraper 1.11, a CLI tool for website screenshots, video demos, and scraping with JavaScript improvements.
How to extract delete keys from explain.depesz.com plans posted via tools like curl, not web browsers.
Tutorial on building an AI agent with Pydantic AI to generate cold emails based on company websites.
Blogger temporarily takes site offline due to AI crawlers causing unsustainable traffic and loss of income from book sales.
A developer explains why they are giving up on building apps that rely on external APIs due to access issues, ethical concerns, and platform risks.
shot-scraper 1.9 CLI tool released, featuring a new -x option to extract page resources and accessibility command fixes.
A technical guide explaining how to use wget with recursive options to download entire websites for offline viewing, including a breakdown of key command-line flags.
Analyzing All The Places' open-source location data project, detailing the technical setup and process for downloading and examining millions of brand locations.
A technical tutorial on building a smart web scraping system that automatically escalates through four tiers of complexity until it succeeds.
A technical analysis of Claude Code's WebFetch and WebSearch tools, detailing their internal architecture and processing pipelines.
Discover an undocumented trick to get xkcd comics at double resolution using a simple URL modification and a Python script to check availability.
Discusses the trend of websites walling off content from AI bots, arguing it undermines open internet principles and may concentrate power.
A talk exploring adversarial web scraping, covering bot detection techniques and ethical methods to bypass them from both scraper and site operator perspectives.
A security researcher discovers goHardDrive exposed thousands of customer records via an insecure RMA status check form with no authentication.
Explores the ethics of LLM training data and proposes a technical method to poison AI crawlers using nofollow links.
Part two of building a personal recommendation system, covering data collection from Pocket and content extraction using the Jina Reader API.
A developer's frustration with aggressive LLM crawlers causing outages and consuming resources, detailing past abuse like crypto mining and Go module mirror issues.