Skip to content
ProxyHub

Web scraping with proxies: best practices, limits and legal aspects

Published · 3 min read

Monitoring prices, availability or search results is everyday work for e-commerce teams and analysts. Doing it well means getting reliable data without overloading websites and staying within the rules. Proxies are just one piece of the puzzle.

Before you start: is there an easier way?

Many websites offer APIs, product feeds or data exports. They're more stable than scraping, don't break with every redesign and make clear what you may do with the data. Scraping makes sense when there's no alternative or it doesn't cover what you need.

This isn't legal advice, but a few points apply almost everywhere:

  • Personal data: in Europe GDPR applies to data published online too. Names, emails, photos and people's profiles require a legal basis and specific safeguards. If you can, avoid collecting them at all.
  • Terms of use: many websites regulate scraping explicitly. Breaching them can have contractual consequences.
  • Copyright and database rights: copying and republishing content or whole databases may infringe the owner's rights.
  • Protected areas: don't collect data behind logins or access controls without permission.
If in doubt, talk to a lawyer before starting a large-scale data collection project.

When you need proxies

  • Local content: prices, availability and search results change by country. With an Italian, Dutch, British or French IP you see what a local user sees.
  • Mobile version: with a mobile proxy you see content as a smartphone user on a carrier network.
  • Spreading load: many requests from a single IP get rate-limited even when the overall pace is reasonable.

Which proxy type to choose depends on volume and target sites: see Mobile vs residential vs datacenter proxies.

Technical best practices

  1. Respect robots.txt and any crawl-rate hints.
  2. Go slow: a few requests per second, with varied pauses. Pacing stands out more than the IP.
  3. Cache pages that rarely change and download only what you need.
  4. Handle errors: on 429 (too many requests) or 503, wait, increasing the delay each attempt.
  5. Rotate sensibly: on a timer or when a site starts limiting you, not on every request. See IP rotation.
  6. Use realistic, consistent headers (user agent, language).
  7. Work during the site's off-peak hours when possible.

Python example

A skeleton with pauses, retries and IP rotation when the site responds with a limit:

Python
import random, time, requests

PROXY = "http://USER:PASSWORD@HOST:PORT"
ROTATE = "ROTATION_LINK"
session = requests.Session()
session.proxies = {"http": PROXY, "https": PROXY}

def fetch(url, tries=4):
    for attempt in range(tries):
        r = session.get(url, timeout=30)
        if r.status_code == 200:
            return r.text
        if r.status_code in (429, 503):
            requests.get(ROTATE, timeout=30)   # new IP
            time.sleep(30 + 2 ** attempt)      # growing delay
            continue
        r.raise_for_status()
    raise RuntimeError(f"Too many attempts: {url}")

for url in ["https://example.com/a", "https://example.com/b"]:
    html = fetch(url)
    time.sleep(random.uniform(2, 5))       # human pace

Proxy setup in Python and Node.js is covered in the Python and Node.js guides.

With an automated browser

Sites that load content with JavaScript need an automated browser (Playwright, Puppeteer, Selenium). The same rules apply, with extra attention to consistency: the browser's language, timezone and operating system must match the proxy, including the TCP/IP fingerprint.

Frequently asked questions

Is web scraping legal?

Collecting public data isn't illegal in itself, but it depends on what data you collect, how you use it and the site's terms. Personal data is subject to GDPR. For significant projects, get legal advice.

How many requests can I make through a mobile proxy?

With ProxyHub traffic is unlimited, but the real limit is set by the website: keep a reasonable pace and spread load over time.

Is a mobile proxy worth it for scraping?

Yes, if you need reliable mobile or local content and sites with strict filters. For large volumes on unfiltered sites a datacenter proxy may be enough.

Try a 5G mobile proxy

Dedicated smartphone, unlimited traffic, real carrier IPs. Live in minutes.

See plans

Read next

All articles