All systems operational•IP pool status
Coronium Mobile Proxies
Developer proxy configuration

Scrapy Proxy Configuration: Authentication, TLS and Crawl Limits

Scrapy’s HTTP proxy middleware reads a proxy URL from request metadata or environment variables. For a whole crawl, configure every request that needs that route, including middleware-generated requests such as robots.txt. Start with the tested diagnostic spider to verify routing before expanding the crawl.

Coronium Technical TeamSources checked 7 min read

Before running the code

  • Use request.meta["proxy"] for the proxy route; website HTTP authentication is a different setting.
  • Configure TLS certificate verification explicitly for the tested Scrapy version.
  • Keep destination concurrency under control when multiple proxy endpoints serve the same crawl.

Check the middleware and download handler together

This guide was reviewed on October 11, 2026 and tested with Scrapy 2.19.0 on Python 3.12. The release history records the current changes. Install the tested version in an isolated environment:

python -m pip install Scrapy==2.19.0

HttpProxyMiddleware accepts an HTTP proxy URL in a request’s proxy metadata. This overrides proxy environment variables and ignores no_proxy for that request. A URL may contain encoded Basic-auth credentials; use urllib.parse.quote(value, safe='') on each username and password when constructing it.

The installed default HTTP/1.1 download handler supports HTTP proxies. Download handler documentation lists protocol support separately: SOCKS support belongs to the optional Httpx handler, and the HTTP11 handler has limits for HTTPS proxy endpoints. Do not infer support for an HTTPS proxy from support for an HTTPS destination through an HTTP proxy.

Route the diagnostic request and robots.txt through the same endpoint

Set PROXY_URL in the runner’s environment. A placeholder has this form: http://user:password@proxy.example:8080. Keep the real value out of source control and logs.

Save the following as scrapy_proxy_check.py:

import ipaddress
import os
from urllib.parse import urlsplit
import scrapy
from scrapy.exceptions import CloseSpider


class ProxyRoute:
    def __init__(self):
        self.proxy_url = os.environ['PROXY_URL']
        endpoint = urlsplit(self.proxy_url)
        if endpoint.scheme != 'http' or not endpoint.hostname or not endpoint.port:
            raise ValueError('Set PROXY_URL to an HTTP proxy with host and port')

    def process_request(self, request):
        request.meta['proxy'] = self.proxy_url


class ProxyCheck(scrapy.Spider):
    name = 'proxy_check'
    custom_settings = {
        'DOWNLOADER_MIDDLEWARES': {f'{__name__}.ProxyRoute': 50},
        'HTTPPROXY_ENABLED': True,
        'DOWNLOAD_VERIFY_CERTIFICATES': True,
        'ROBOTSTXT_OBEY': True,
        'CONCURRENT_REQUESTS': 1,
        'DOWNLOAD_TIMEOUT': 20,
        'DOWNLOAD_MAXSIZE': 65536,
        'RETRY_ENABLED': False,
        'REDIRECT_ENABLED': False,
        'LOG_LEVEL': 'WARNING',
    }

    async def start(self):
        target = getattr(self, 'target', 'https://api.ipify.org?format=json')
        yield scrapy.Request(target, callback=self.parse,
                             errback=self.failed,
                             meta={'handle_httpstatus_all': True})

    def parse(self, response):
        if response.status != 200:
            raise CloseSpider(f'http_{response.status}')
        try:
            address = str(ipaddress.ip_address(response.json()['ip']))
        except (ValueError, KeyError, TypeError):
            raise CloseSpider('invalid_ip_response') from None
        yield {'ip': address}

    def failed(self, failure):
        raise CloseSpider('request_failed')

Run it with:

scrapy runspider scrapy_proxy_check.py -O exit.json

A successful response produces one item containing the observed IP. Inspect exit.json and the crawl outcome: a Scrapy process ending does not by itself mean the spider collected a valid item. The default target is ipify; use -a target=https://your-test-endpoint.example/ip for a controlled equivalent.

The small ProxyRoute middleware sets the same endpoint on every request in this diagnostic spider. Its priority runs before robots processing. Scrapy’s robots middleware source creates its own request, so setting metadata only on the spider’s initial request would leave that extra request to other proxy configuration. This example keeps robots checks enabled and applies the route to those requests too.

The asynchronous start() method follows the current Spider API. The diagnostic disables retries and redirects so failures remain visible instead of turning into an unexplained sequence of requests.

Enable certificate verification explicitly

The settings reference lists DOWNLOAD_VERIFY_CERTIFICATES as false by default in the reviewed release. The example sets it to true. This makes certificate failures visible rather than treating an unverified TLS peer as trusted.

A proxy password and a destination certificate solve different problems. If a managed environment uses an approved private certificate authority, configure the selected handler’s trust path and test it deliberately. Disabling certificate verification is not an authentication fix.

The example also sets a twenty-second download timeout and a small response-size limit. These values fit an IP diagnostic. Choose limits appropriate to the documents your production spider is authorized to fetch, and verify that an alternative download handler honors the settings you rely on.

Separate proxy credentials, website credentials and crawl state

Scrapy’s proxy middleware implementation extracts credentials from the configured URL and prepares Proxy-Authorization. Destination HTTP authentication uses different settings, including the HTTP auth middleware and its domain restrictions. Do not put a proxy password into a destination Authorization header.

Treat request dumps, exception details and persisted crawl state as potentially sensitive. Even when middleware removes user information from a working URL, your original environment or custom configuration still contains it. Keep endpoint identifiers separate from passwords in metrics.

When you rotate between approved endpoints, assign the endpoint where requests are constructed or through one controlled middleware. The diagnostic’s ProxyRoute intentionally overwrites each request with one fixed endpoint. Change that policy deliberately before using per-request routes; otherwise a later setting may replace the one you thought you selected.

Cookies can also outlive a single request. Preserve the endpoint and cookie state required by an authorized session, and avoid assuming a different exit address creates a new account context.

Budget concurrency for the destination

AutoThrottle adjusts delays by download slot. The default slot assignment follows the request URL’s domain, and its target concurrency is an average goal rather than a hard guarantee. A pool of proxy addresses does not remove the destination’s rate limits.

Begin with low concurrency and an explicit request budget. Add AutoThrottle for adaptive pacing, while retaining the relevant concurrency and delay limits. If you customize slots to match proxies, review what happens when several slots request the same destination at once.

Retry only failures that your job can safely repeat. Preserve status codes and service instructions. A 407 requires proxy-authentication diagnosis; a 403 needs destination-response inspection. A 429 should trigger backoff according to the service’s response, not unlimited address changes. The IP restriction guide helps separate those cases.

Use the current optional handler for SOCKS requirements

The Httpx download handler requires Scrapy’s httpx extra and an asyncio-compatible setup. The documentation provides separate mappings for both http and https destination schemes. It also notes that a separate connection pool is created for each proxy URL, which can increase resource use with large rotating pools.

If SOCKS is a requirement, follow that handler’s installation and protocol instructions, then run a small routing and authentication test. Do not paste a SOCKS URL into an older default-handler recipe and assume it works. Our executed example uses the default handler with an HTTP proxy; it does not claim a SOCKS or HTTP/2 fixture test.

What was tested, and how this applies to AI pipelines

The published spider passed a local HTTP fixture test for encoded credentials, proxy routing of robots.txt, conflicting environment variables, separation from destination authentication and a 403 response producing no item without retries. Those are reproducible code checks, not measurements of a proxy provider or permission to collect a website’s data.

In an AI-assisted data pipeline, Scrapy controls the fetch stage. An LLM extraction call made later by a different client has its own network configuration. Store source URL, retrieval time and response status alongside extracted data so downstream summaries do not hide failed fetches.

Use Python Requests for a smaller synchronous HTTP job, or Playwright when authorized browser rendering is required. Pick the client for the request the workflow actually needs.

Sources and review scope

Sources reviewed October 11, 2026. This guide reviews primary documentation. Code examples illustrate configuration and error handling. Any local fixture checks are described in the article; they are not live-provider performance benchmarks or guarantees of destination access.

Frequently asked questions

Configure another runtime

Use the settings supported by the process making the request, then verify routing and authentication.

Related workflows

Python Requests authentication

Start with a smaller request when a crawl framework is unnecessary.

Playwright sessions

Configure authorized browser-rendered work.

Proxy restriction diagnosis

Interpret returned errors before changing endpoints.