Skip to main content

Building public-data crawlers

Build an ethical crawler in six steps with Playwright, http_utils, and APScheduler.

Difficulty
Intermediate
Lessons
6

Public data like NPS, DART, and HIRA is accessible to everyone, but automation comes with rules — robots.txt, rate limits, terms of service. Six steps to an ethical and sustainable crawler.

Who it's for

  • Developers who need more control than portal APIs offer
  • Anyone who has been blocked by a crawl target
  • Teams who want incremental collection, schedules, and observability

What you can do afterwards

  • Separate dynamic pages (Playwright) from static ones (BS4)
  • Apply robots.txt + rate limit + backoff
  • Schedule in KST with APScheduler
  • Combine public APIs, ministry CSVs, and web scraping
  • Incremental collection, dedup, checkpoints
  • Healthchecks and failure alerts

Flow

A safe collection pipeline

Permission boundary

Check ethical and legal constraints before choosing a collection tool.

Request control

Protect the source system with rate limits and schedules.

Data integrity

Make reruns safe through incremental processing and deduplication.

Operational visibility

Detect failures and stale data with metrics and alerts.

The goal of a crawler: refresh our DB without loss while not trespassing on the source site. The flow above strengthens both axes in turn.

Steps

  1. Crawler ethics and legal boundaries — robots.txt · terms · personal data
  2. Static vs dynamic — BS4 + Playwright — pick the right tool
  3. Rate limiting · retries · backoff — exponential + jitter
  4. APScheduler + KST — idempotency · replace_existing=True · double-trigger defence
  5. Incremental collection · deduplication — checkpoints · unique keys · change detection
  6. Observability · alerts — success rate · latency · Slack · PagerDuty

Prerequisites — complete python-data-pipeline.

Lessons

  1. 1

    Crawler ethics and legal boundaries

  2. 2

    Static vs dynamic — BS4 + Playwright

  3. 3

    Rate limit · retries · backoff

  4. 4

    APScheduler + KST schedules

  5. 5

    Incremental collection · deduplication

  6. 6

    Observability · alerts

Other courses

All courses →