Scheduling Scrapy with Airflow: retries, backfills, alerts
· Surendra Tamang
Cron is fine for one spider. Once you have ten sources feeding the same database, the gaps show: one failed source means rerunning everything, nobody can see what ran last night, and backfilling “last Tuesday for source B” means editing a crontab by hand.
This is how I move Scrapy projects to Apache Airflow without rewriting the spiders.
Make every spider rerunnable first
Airflow can only retry and backfill safely if a run for a given date is idempotent: running it twice gives the same result as running it once. Two things make that true:
- The spider takes the logical date as an argument, instead of reading “today” from the clock:
class PricesSpider(scrapy.Spider): name = "prices"
def __init__(self, run_date=None, **kwargs): super().__init__(**kwargs) self.run_date = run_date # "2025-10-07", set by Airflow- The pipeline upserts on a natural key (the site’s own ID) instead of inserting. Reruns then update rows instead of duplicating them. The schema for this is in From a scraper to a monitored data pipeline.
If your spiders already work like that, the Airflow part is small.
One DAG, one task per source
This example uses Airflow 3’s airflow.sdk imports; on Airflow 2.x the same ideas apply with airflow.decorators and airflow.operators.bash.
from datetime import datetime, timedelta
from airflow.providers.standard.operators.bash import BashOperatorfrom airflow.sdk import dag, task
SOURCES = ["retailer_a", "retailer_b", "retailer_c"]
@dag( schedule="15 2 * * *", start_date=datetime(2025, 10, 1), catchup=False, default_args={ "retries": 3, "retry_delay": timedelta(minutes=10), "retry_exponential_backoff": True, },)def prices(): for source in SOURCES: crawl = BashOperator( task_id=f"crawl_{source}", bash_command=f"cd /opt/scrapers && scrapy crawl {source} -a run_date={{{{ ds }}}}", pool="residential_proxies", )
@task(task_id=f"check_{source}") def check(source=source, ds=None): rows = count_rows(source, ds) # your query against the target table if rows == 0: raise ValueError(f"{source}: zero rows for {ds}")
crawl >> check()
prices()What each part gives you:
- One task per source. Retailer B failing doesn’t block A and C, and a retry reruns only B.
{{ ds }}: Airflow passes the run’s logical date, so a rerun of Tuesday scrapes as Tuesday.- Retries with exponential backoff: temporary blocks and proxy hiccups often clear on their own after a few minutes.
- A check task after each crawl: a spider that exits cleanly with zero rows now fails loudly, which is the most common silent failure in scraping.
Pools: limit concurrency by proxy, not by server
Scrapy already controls concurrency inside one spider. The problem is five spiders starting at 02:15 and all hitting the same residential proxy pool at once.
An Airflow pool caps how many tasks can use a shared resource at the same time:
airflow pools set residential_proxies 2 "Paid residential proxy pool"Now at most two proxy-heavy crawls run together, and the rest queue. Sources on cheap datacenter IPs can use a different pool with more slots.
Backfills without editing anything
A site was down for three days, or you fixed a parser bug and want to rebuild last week. With idempotent spiders, that’s one command:
# Airflow 3airflow backfill create --dag-id prices --from-date 2025-09-29 --to-date 2025-10-05
# Airflow 2.xairflow dags backfill prices -s 2025-09-29 -e 2025-10-05Each day runs with its own ds, and the upserts make the reruns safe.
Alerts that fire on bad data
Airflow can call a function whenever a task fails. Send that to Slack or email:
def notify(context): ti = context["task_instance"] post_to_slack(f"{ti.dag_id}.{ti.task_id} failed for {context['ds']}: {ti.log_url}")
# add to default_args: "on_failure_callback": notifyBecause the check tasks raise on zero or low row counts, this one callback covers both crashes and silent failures. It stays quiet when everything is fine, so people keep reading it.
When Airflow is overkill
Stay on cron if you have one or two spiders, no downstream dependencies, and someone checks the output anyway. Airflow is a service you have to run and upgrade. Move when retries per source, backfills, dependencies or run history start costing you real time.
FAQ
Should the spider run inside the Airflow worker? For small setups, yes (a BashOperator is fine). For heavier crawls, run each spider in its own container with KubernetesPodOperator or DockerOperator, so a memory-hungry browser crawl can’t take down the scheduler.
What about Scrapyd or Zyte’s Scrapy Cloud? They’re good for running spiders. Airflow adds the orchestration around them: dependencies, checks, backfills and one place to see everything. You can trigger Scrapyd jobs from Airflow too.
How do I load into a warehouse after all sources finish? Add a final task downstream of every check task. It only runs when all sources pass, so your warehouse never gets a half-loaded day.
Airflow, Prefect or Dagster? All three work. Airflow has the largest ecosystem and is the one most job postings ask for, which matters if clients will maintain it after you.