Projects
Selected work. Client names withheld under NDA; details available on request.
-
Scraper fleet infrastructure
- Problem
- Millions of pages a month across sites protected by Cloudflare, DataDome and similar vendors. Generic scrapers were blocked within hours.
- Built
- Distributed Scrapy and Playwright crawlers with proxy rotation, TLS and browser fingerprint management, retry and backoff policies, and a monitoring layer that alerts on block rate, zero-row runs and schema drift.
- Stack
- Python, Scrapy, Playwright, residential proxies, PostgreSQL, Docker, Airflow
- Outcome
- Runs unattended on schedule; failures surface as alerts instead of silent gaps in the data.
-
E-commerce price and competitor monitoring
- Problem
- A retail team checked competitor prices and listings by hand every morning.
- Built
- Scheduled collection of competitor prices, availability and listings, cleaned and deduplicated, with change history so every price move is recorded.
- Stack
- Python, Playwright, PostgreSQL, scheduled jobs, CSV and API delivery
- Outcome
- Daily structured dataset delivered before the team starts work; manual checks removed.
-
Data extraction pipelines with LLM parsing
- Problem
- Source data spread across websites, PDFs and APIs with no consistent structure.
- Built
- One pipeline that merges the sources, uses Claude / OpenAI for the unstructured parts under a strict output schema, validates every field, and loads into the client database, API or warehouse.
- Stack
- Python, Claude and OpenAI APIs, PDF and OCR tooling, PostgreSQL, Snowflake
- Outcome
- Validated records instead of free text; the model fills fields but cannot invent them.
-
Local business data at scale
- Problem
- Lead and market datasets needed from Google Maps and business directories across many regions.
- Built
- Geo-batched scraping jobs into PostgreSQL with deduplication and enrichment, exported as ready-to-use lead and market datasets.
- Stack
- Python, PostgreSQL, proxies, geo-batching, enrichment APIs
- Outcome
- Clean, deduplicated business records at regional scale.
Open source
- Graphify
Turns code, docs and papers into a queryable knowledge graph with community detection.