Web scraping & data pipelines · Python · AI extraction
Web scraping that keeps running
when others get blocked.
Cloudflare, DataDome and site changes break most scrapers within weeks. I build scrapers and data pipelines that survive both — then keep them running as a monitored data feed your team can rely on.
Top Rated on Upwork · Apify Ambassador · Small data engineering team
5.0★
Upwork, 50+ projects
Since 2015
scraping & data eng
Millions
pages / month
Surendra Tamang
Pokhara, NP
How I work
01 · Scope and risk check
We confirm the target sites, the exact fields, and the delivery format. I flag ToS, personal-data and rate-limit risks before any code is written.
02 · Paid phase-1 sample
A small, fixed-price run on real targets. You see output quality and the true cost per record before committing to the full build.
03 · Built to keep running
Anti-bot handling, proxy rotation, scheduling, deduplication with change history, and alerts when a run returns nothing.
04 · Handoff, not lock-in
Clear code, a short doc and a walkthrough video. The code and the accounts are yours.
What I build
Scraping that survives anti-bot systems
Cloudflare, DataDome, Kasada, Akamai, PerimeterX. Fingerprint and session handling, residential proxies, Playwright and Scrapy at scale.
Data pipelines
Websites, PDFs and APIs merged into one dataset. Incremental loads, data-quality checks, delivered to PostgreSQL, Snowflake, Sheets or an API.
AI extraction with validation
Claude and OpenAI for the unstructured parts, under a strict schema with accuracy checks. The model fills fields; it cannot invent them.
Monitored data feeds
Competitor prices, listings or records tracked on schedule with change history and alerting. A feed your team relies on, not a script someone babysits.
Client results
“Another successful job with Surendra! I really recommend his work, he always delivers in time and with really high quality! Great communication and great quality!”
“Great to work with, really professional. Quality of code was beyond expectation.”
“Surendra did a great job on a LAMP REST API deployment to Google Cloud Platform. We'd be happy to work with this freelancer again.”
From my Upwork profile.
Selected work
Scraper fleet infrastructure
Distributed crawls across millions of pages a month. Proxy rotation, fingerprint management, retries, and alerting on block rate and zero-row runs.
LLM extraction pipelines
Websites, PDFs and APIs merged into validated records. Claude and OpenAI under a strict output schema, loaded into the client database or warehouse.
Ways to work together
Consultation
30 minutes on your target site or pipeline: is it scrapable, what blocks it, what it costs to run. You leave with an approach and an estimate.
Book a consultation →Fixed-scope build
Defined targets and output, delivered with docs and monitoring. Starts with a paid phase-1 sample. Larger builds are handled with my small team.
Monitored data feed
I run and maintain the pipeline; you receive clean data on schedule. Sites change every few weeks, so fixes are included.
- Basic $300/mo scheduled runs, failure alerts, up to 2 site-change fixes
- Standard $600/mo + quality checks, change history, monthly health report
- Priority $1,000+/mo + 24h fixes, one new source or field a month, proxy costs managed
Every engagement includes a compliance check up front (terms of service, personal data, rate limits) and monitoring with the delivery. I use Claude Code daily; the engineering discipline is still mine.
Questions
Can you tell if a site can be scraped before I pay for a build?
Yes. I check the site's protection (Cloudflare, DataDome, Akamai), login needs and page structure, and tell you what's realistic and roughly what it costs to run.
My scraper worked before and now it gets blocked. Can you help?
Usually, yes. Send the code or logs; the cause is typically fingerprinting, rate limits, proxy quality or a site change, and each has a known fix.
Why a paid phase-1?
You see real data from your real targets before committing to the full build. If the data is not what you need, you have spent a fraction of the budget.
What happens after delivery?
You own the code and can run it yourself. Or I keep it running as a monitored data feed, so a site change is my problem, not yours.
Do you only work with Python?
Mostly Python (Scrapy, Playwright, FastAPI, Airflow), but I can advise on any stack and integrate with Node, Apify, or your existing systems.
Can you extract data from PDFs or messy pages with AI?
Yes. LLM extraction under a strict schema, checked against a hand-verified sample so you know the per-field accuracy before relying on it.
What about legality?
I focus on publicly available data and flag terms-of-service, login-wall and personal-data risks up front. For legal decisions, confirm with your own counsel.
Field notes
- From a scraper to a monitored data pipeline
· 5 min read
- LLM extraction that never invents a value
· 4 min read
- What web scraping really costs per 1,000 pages
· 4 min read
A site that keeps blocking you, or a data feed you need built and maintained?
Send the target and the output you need. I'll tell you honestly if it's doable, the risks, and a phase-1 price.
Start a project