Surendra Tamang

Web scraping & data pipelines · Python · AI extraction

Web scraping that keeps running
when others get blocked.

Cloudflare, DataDome and site changes break most scrapers within weeks. I build scrapers and data pipelines that survive both — then keep them running as a monitored data feed your team can rely on.

Top Rated on Upwork · Apify Ambassador · Small data engineering team

5.0★

Upwork, 50+ projects

Since 2015

scraping & data eng

Millions

pages / month

Surendra Tamang

Surendra Tamang
Pokhara, NP

How I work

01 · Scope and risk check

We confirm the target sites, the exact fields, and the delivery format. I flag ToS, personal-data and rate-limit risks before any code is written.

02 · Paid phase-1 sample

A small, fixed-price run on real targets. You see output quality and the true cost per record before committing to the full build.

03 · Built to keep running

Anti-bot handling, proxy rotation, scheduling, deduplication with change history, and alerts when a run returns nothing.

04 · Handoff, not lock-in

Clear code, a short doc and a walkthrough video. The code and the accounts are yours.

What I build

Scraping that survives anti-bot systems

Cloudflare, DataDome, Kasada, Akamai, PerimeterX. Fingerprint and session handling, residential proxies, Playwright and Scrapy at scale.

Data pipelines

Websites, PDFs and APIs merged into one dataset. Incremental loads, data-quality checks, delivered to PostgreSQL, Snowflake, Sheets or an API.

AI extraction with validation

Claude and OpenAI for the unstructured parts, under a strict schema with accuracy checks. The model fills fields; it cannot invent them.

Monitored data feeds

Competitor prices, listings or records tracked on schedule with change history and alerting. A feed your team relies on, not a script someone babysits.

Client results

“Another successful job with Surendra! I really recommend his work, he always delivers in time and with really high quality! Great communication and great quality!”
Pharmacy data scraping · Upwork client · 5.0
“Great to work with, really professional. Quality of code was beyond expectation.”
Job posting crawler · Upwork client · 5.0
“Surendra did a great job on a LAMP REST API deployment to Google Cloud Platform. We'd be happy to work with this freelancer again.”
GCP REST API deployment · Upwork enterprise client · 5.0

From my Upwork profile.

Selected work

Scraper fleet infrastructure

Distributed crawls across millions of pages a month. Proxy rotation, fingerprint management, retries, and alerting on block rate and zero-row runs.

LLM extraction pipelines

Websites, PDFs and APIs merged into validated records. Claude and OpenAI under a strict output schema, loaded into the client database or warehouse.

All projects →

Ways to work together

Consultation

30 minutes on your target site or pipeline: is it scrapable, what blocks it, what it costs to run. You leave with an approach and an estimate.

Book a consultation →

Fixed-scope build

Defined targets and output, delivered with docs and monitoring. Starts with a paid phase-1 sample. Larger builds are handled with my small team.

Monitored data feed

I run and maintain the pipeline; you receive clean data on schedule. Sites change every few weeks, so fixes are included.

  • Basic $300/mo scheduled runs, failure alerts, up to 2 site-change fixes
  • Standard $600/mo + quality checks, change history, monthly health report
  • Priority $1,000+/mo + 24h fixes, one new source or field a month, proxy costs managed

Every engagement includes a compliance check up front (terms of service, personal data, rate limits) and monitoring with the delivery. I use Claude Code daily; the engineering discipline is still mine.

Questions

Can you tell if a site can be scraped before I pay for a build?

Yes. I check the site's protection (Cloudflare, DataDome, Akamai), login needs and page structure, and tell you what's realistic and roughly what it costs to run.

My scraper worked before and now it gets blocked. Can you help?

Usually, yes. Send the code or logs; the cause is typically fingerprinting, rate limits, proxy quality or a site change, and each has a known fix.

Why a paid phase-1?

You see real data from your real targets before committing to the full build. If the data is not what you need, you have spent a fraction of the budget.

What happens after delivery?

You own the code and can run it yourself. Or I keep it running as a monitored data feed, so a site change is my problem, not yours.

Do you only work with Python?

Mostly Python (Scrapy, Playwright, FastAPI, Airflow), but I can advise on any stack and integrate with Node, Apify, or your existing systems.

Can you extract data from PDFs or messy pages with AI?

Yes. LLM extraction under a strict schema, checked against a hand-verified sample so you know the per-field accuracy before relying on it.

What about legality?

I focus on publicly available data and flag terms-of-service, login-wall and personal-data risks up front. For legal decisions, confirm with your own counsel.

A site that keeps blocking you, or a data feed you need built and maintained?

Send the target and the output you need. I'll tell you honestly if it's doable, the risks, and a phase-1 price.

Start a project