
Zyte API is a managed web scraping API. An application sends a target URL and selects the output it needs, such as an HTTP response, browser-rendered HTML, a screenshot, or structured data. Zyte then manages some combination of network access, proxy selection, browser infrastructure, sessions, and extraction before returning the result.
That makes Zyte API different from a raw proxy. A raw proxy supplies a network route for a scraper you operate. Zyte API can manage substantially more of the retrieval stack.
Quick Answer
Zyte API is a managed web scraping service that can retrieve HTTP responses, render JavaScript, run browser actions, maintain sessions, and extract structured data. Developers call one authenticated endpoint and request the output they need. It is best suited to teams that want managed retrieval infrastructure; teams that already operate their own collectors may prefer direct proxy access and greater network-level control.
Key Takeaways
- Zyte API is more than a proxy endpoint. Its documented capabilities include HTTP fetching, rendered-browser requests, screenshots, browser actions, network capture, sessions, automatic extraction, and a CDP browser for Playwright or Puppeteer.
- The client chooses an operating mode. HTTP requests, browser requests, CDP, automatic extraction, and proxy mode provide different levels of control and different outputs.
- The main HTTP API uses one endpoint. Clients send an authenticated POST request to https://api.zyte.com/v1/extract with the target URL and requested fields.
- Geolocation and session controls are available, but they are not unlimited. The public API documentation describes country-level geolocation and different session semantics across the HTTP API, proxy mode, and CDP.
- Smart Proxy Manager is a legacy product. Zyte moved remaining Smart Proxy Manager traffic to Zyte API Proxy Mode; current integrations should normally start with Zyte API documentation.
- Zyte manages access, not the complete data product. Customers still need to define sources, validate results, store records, monitor quality, and ensure their collection and downstream use are permitted.
- The Python example below was syntax-checked but not run with live Zyte credentials. The practice page and selectors were checked separately in a direct browser, but that does not validate Zyte authentication, routing, rendering, or performance.
Methodology: We reviewed Zyte's public product and developer documentation on October 4, 2026. This is a documentation-based guide, not a controlled benchmark of Zyte, Proxidize, or another provider. We checked the practice page and selectors separately with a direct Playwright 1.61.0 browser, which found 10 complete quote records, but did not run the request through Zyte. Features, pricing, account requirements, and rate limits can change, so verify them against the linked first-party documentation before implementation.
Disclosure: Proxidize publishes this guide and sells residential and mobile proxy infrastructure. Zyte API is a managed scraping and extraction service, while Proxidize supplies raw proxy access for customers running their own collection software. The two products overlap around web access but do not replace the same engineering layer.
Zyte API at a Glance
| Category | Current documented offering |
|---|---|
| Product type | Managed web scraping, browser, and extraction API |
| Main endpoint | `POST https://api.zyte.com/v1/extract` |
| Authentication | HTTP Basic authentication using the Zyte API key as the username and an empty password |
| Retrieval modes | HTTP requests, browser requests, CDP headless browser, and proxy mode |
| Main outputs | Base64-encoded HTTP body, browser-rendered HTML, screenshots, captured network responses, and structured extraction results |
| Browser control | Predefined actions through browser requests or direct Playwright/Puppeteer control through CDP |
| Location control | Automatic location selection or documented country-level geolocation override |
| Sessions | Client-managed and server-managed sessions for supported HTTP/browser workflows; a CDP connection keeps its own state while connected |
| Scrapy integration | `scrapy-zyte-api` |
| Best fit | Teams that want to outsource access, browser, and optional extraction infrastructure |
| Customer still owns | Source selection, crawl policy, validation, scheduling, storage, monitoring, business logic, and compliance |
| Main tradeoff | Less direct control over the underlying exit network and retrieval decisions than an in-house collector using raw proxies |
What Is Zyte API?
Zyte API is an intermediary between a web-data application and a target website. Instead of configuring a browser, selecting a proxy, deciding when to retry, and returning the retrieved content through separate systems, a developer sends Zyte a structured request.
A basic request looks like this:
The client does not normally select an individual exit IP. According to the Zyte API feature documentation, Zyte can select geolocation and IP type automatically or accept supported overrides. The service can also escalate from an HTTP-oriented request to browser infrastructure when the selected mode or requested output requires it.
This abstraction is useful when a team wants to focus on the returned content rather than operating every access component. It also means the client has less direct visibility into and control over the underlying network than it would with raw proxy credentials.
How Does Zyte API Work?
- Select an output. Request httpResponseBody, browserHtml, a screenshot, network capture, browser actions, or a documented extraction type. Use the smallest mode that can produce a valid result; a server-rendered page may not need a browser.
- Send an authenticated request. The main HTTP API accepts JSON at https://api.zyte.com/v1/extract. The API key is the Basic-auth username and the password is empty.
- Let Zyte select the retrieval path. The requested fields and target determine whether the job uses HTTP fetching, browser infrastructure, extraction, and related access features.
- Receive a JSON response. HTTP bodies are Base64-encoded; browser HTML is returned as a string. Other fields depend on the requested feature.
- Validate before accepting the result. Check the requested URL and context, required output field, page identity, action outcome, location or language evidence, required data fields, freshness, and permission to use the data.
A successful API response is not automatically a valid dataset record. The page can still be stale, incomplete, localized incorrectly, or different from the expected entity.
Zyte API Request Modes Compared
Zyte's current usage documentation describes several ways to use the product. They should not be treated as interchangeable labels.
| Mode | Client sends | Zyte returns or exposes | Best fit | Main limitation |
|---|---|---|---|---|
| HTTP request | URL plus HTTP-oriented options | Base64-encoded response body and optional response information | Server-rendered pages, APIs, files, and requests needing method/body control | Does not execute page JavaScript |
| Browser request | URL plus browser output/actions | Rendered HTML, screenshot, action results, or selected network captures | JavaScript pages and predefined interactions | Initial request has less method/header flexibility than HTTP mode |
| Automatic extraction | URL plus a documented data type or schema | Structured fields | Products, articles, jobs, page content, and other supported extraction tasks | Output quality and available fields must be tested for the target |
| CDP headless browser | Playwright, Puppeteer, or another CDP client | Live browser connection | Complex, branching, or stateful browser code | More client responsibility; different session and pricing model |
| Proxy mode | Normal request through `api.zyte.com:8011` | Target response through a proxy-compatible interface | Migrating older proxy integrations or tools expecting a proxy endpoint | Does not expose the full HTTP API feature set and is not recommended for browser automation |
HTTP requests
Set httpResponseBody to true when the page does not require browser rendering. The returned body is Base64-encoded, so the client must decode it before parsing text, JSON, or binary content.
HTTP mode is also the appropriate starting point when the client needs to control the initial request method, body, or supported request headers. The HTTP request documentation covers those fields and the decoding step.
Browser requests
Set browserHtml to true when the target needs JavaScript rendering. Zyte returns the rendered DOM as a string. Browser mode can also produce screenshots, execute documented actions, and capture matching background responses.
Zyte documents an important limitation: a browser request does not provide the same control over the initial method, body, and headers as an HTTP request. If the workflow needs a live browser with more direct control, use CDP instead.
Automatic extraction
Automatic extraction asks Zyte to return structured fields rather than only page markup. The current API reference describes types for products, product lists, articles, job postings, page content, forum threads, navigation, SERPs, and custom attributes.
Treat automatic extraction as a data source that needs validation, not as guaranteed ground truth. Test required fields, variants, localization, missing values, and schema consistency across representative pages.
CDP headless browser
Zyte exposes a remote browser through the Chrome DevTools Protocol. A developer can connect an existing Playwright or Puppeteer script and continue using browser-library methods for navigation, cookies, network inspection, screenshots, and interactions.
CDP gives the client more browser control than a normal Zyte browser request. It also shifts more responsibility back to the client. Zyte's current CDP documentation says CDP requires a subscription or a PAYG account with a spending limit plus business verification; it is not available through the initial free credit alone.
Proxy mode
Proxy mode exposes api.zyte.com:8011 for software that expects proxy configuration rather than a JSON API. The Zyte API key is used as the proxy username with an empty password.
This does not turn Zyte into an ordinary raw-proxy list. Zyte still manages the destination-side access layer. Proxy mode is primarily useful for compatibility and migration, and Zyte states that it is not optimized for use with browser-automation tools.
Zyte API vs Smart Proxy Manager and Crawlera
Crawlera was renamed Zyte Smart Proxy Manager. Smart Proxy Manager is now a legacy product rather than the main service a new integration should target.
Zyte's Smart Proxy Manager sunset documentation says remaining Smart Proxy Manager traffic was moved to Zyte API Proxy Mode through a compatibility layer in December 2025. Old hostnames, headers, library names, and tutorials may therefore still appear in search results and existing codebases.
For a new project:
- Start with the current Zyte API usage documentation.
- Prefer the HTTP API when you need its richer output and feature set.
- Use scrapy-zyte-api for a current Scrapy integration.
- Use proxy mode when a compatible proxy interface is specifically required.
- Do not use an old Crawlera tutorial as the source of truth for a new implementation.
How to Authenticate With Zyte API
Use the API key as the Basic-auth username and an empty password:
In cURL, the trailing colon represents the empty password:
Avoid passing a real key directly in a shared command, process log, shell history, screenshot, or article. For production workloads, store it in the organization's approved secret-management system and restrict access to the service that needs it.
How to Use Zyte API With Python
This example requests browser-rendered HTML from the JavaScript version of Quotes to Scrape, extracts quote records, validates the required fields, and writes the result to JSON.
It demonstrates the Zyte request and response shape without targeting a commercial website. The code was syntax-checked locally, but it was not sent through a live Zyte account for this article.
1. Create an environment and install the packages
On Windows PowerShell, activate the environment with:
2. Set the API key
On macOS or Linux:
On Windows PowerShell:
3. Save the complete script
Save this as zyte_api_example.py:
4. Run the script
A successful run should create zyte_quotes.json. The record shape should resemble this, but the following is illustrative rather than measured output from our environment:
The validation matters. Saving an empty file after a technically successful API response would make the example look operational while hiding a rendering or selector failure.
How to Request a Normal HTTP Response in Python
When JavaScript rendering is unnecessary, request httpResponseBody instead of browserHtml. The response field is Base64-encoded:
Do not try to parse the Base64 string as if it were HTML. Decode it first. For arbitrary binary content or unknown character encoding, keep the value as bytes or determine the encoding from the returned response information before decoding.
How Geolocation Works in Zyte API
Zyte can select a location automatically based on the target. A client can also set the geolocation field to a supported country code:
The current public FAQ describes explicit geolocation at country granularity. Do not claim city-level targeting for Zyte API unless a newer first-party source documents it for the exact product and mode being discussed.
Some locations are classified as extended geolocations and can affect request cost. Geolocation also does not independently prove that a page used the expected language, currency, store, or business location. Validate the returned content rather than accepting the request field as evidence of the observed result.
How Sessions Work
Sessions are useful when multiple requests must share state, such as an IP address, cookie jar, or network context.
Zyte documents two session models for supported HTTP and browser requests:
- Client-managed sessions: the client supplies a session identifier and controls which requests belong together.
- Server-managed sessions: the client supplies a session context and prerequisites, and Zyte creates or reuses a suitable session for that context.
The semantics differ in CDP. One CDP connection already has a browser, IP, and cookie jar for the life of that connection. A separate CDP connection cannot automatically reuse the state of a previous one.
Do not assume that a generic “session” setting behaves identically across HTTP requests, browser requests, proxy mode, and CDP. Use the documentation for the selected mode and test continuity with the exact workflow.
How to Use Zyte API With Scrapy
Zyte maintains a dedicated integration called scrapy-zyte-api. Its transparent mode can map ordinary Scrapy requests onto Zyte API while still allowing per-request Zyte options.
Install it in the Scrapy environment:
A current settings pattern is:
Keep the key outside the repository. Confirm the configuration against the current scrapy-zyte-api documentation before deployment because Scrapy integration settings and version requirements can change.
The broader Python web scraping libraries guide explains where Scrapy fits relative to Requests, parsers, and browser automation. If a spider needs a local Playwright browser instead of a managed browser API, the Scrapy Playwright guide covers that architecture separately.
How Zyte API Pricing Works
Zyte API does not have one universal price per request. The current Zyte API pricing documentation says the base cost depends on the target website and whether the request uses HTTP or a browser. Actions, screenshots, network capture, automatic extraction, custom attributes, extended geolocation, device residential IP use, and CDP can introduce separate cost rules.
Zyte states that only responses it classifies as successful are billed. That does not mean every billed response is a valid record for the customer's dataset. A correctly returned page can still fail a business-level validation check.
Do not compare Zyte's price per successful response directly with a raw proxy's price per gigabyte without including the work each purchase covers. A fair comparison uses cost per accepted record or completed observation and includes browsers, compute, retries, extraction, validation, and engineering time.
This article intentionally does not reproduce a full plan table. Pricing is a separate search intent and should be checked immediately before a purchasing decision.
Benefits and Limitations of Zyte API
| Area | Benefit | Tradeoff to evaluate |
|---|---|---|
| Combined infrastructure | HTTP access, browsers, proxy selection, sessions, actions, capture, and extraction can sit behind one service | The customer has less direct control over the underlying network and access decisions |
| Selective escalation | Simple pages can use HTTP while harder pages use browser or extraction features | The modes have different methods, headers, session behavior, outputs, and prices |
| Browser options | Zyte can drive a browser request, or the customer can control a Zyte-hosted browser through CDP | CDP shifts browser logic and more access responsibility back to the customer |
| Scrapy integration | `scrapy-zyte-api` reduces request plumbing for Scrapy projects | The integration adds provider-specific settings and semantics |
| Structured output | Automatic extraction can reduce selector maintenance | Missing, stale, or misinterpreted fields still require validation |
| Geolocation | Zyte can choose a location or accept a supported country override | Current public documentation does not establish city- or ISP-level targeting for this product |
| Pricing | Failed and rate-limited requests are not billed under the documented model | Target tier, request mode, actions, extraction, screenshots, location, and network use can change cost |
| Portability | One API can replace several internal components | Zyte-specific request fields and schemas can make a future migration more involved than changing a proxy endpoint |
Zyte API vs a Raw Proxy
Zyte API and raw proxies belong to different layers of a web-data system.
| Responsibility | Zyte API | Raw proxy service such as Proxidize |
|---|---|---|
| Network route and exit selection | Managed by the provider | Provider supplies endpoints; customer selects documented targeting/session options |
| HTTP client | API request to Zyte; Zyte retrieves the target | Customer operates it |
| Browser infrastructure | Available as a managed request or CDP browser | Customer operates it |
| JavaScript rendering | Available | Customer operates it |
| Browser actions | Available through supported actions or CDP | Customer writes and runs them |
| Parsing and extraction | Optional structured extraction | Customer operates it |
| Retry/access strategy | Largely managed within the service | Customer operates it |
| Output | HTML, browser output, screenshot, capture, or structured data | Target response received by the customer's client |
| Billing unit | Successful requests and selected features, subject to current pricing rules | Usually bandwidth or proxy allocation, depending on product |
| Best fit | Teams wanting managed retrieval and optional extraction | Teams with their own scraper, browser, or agent stack |
The raw proxies vs. scraping APIs guide provides a fuller cost and architecture comparison.
Where Proxidize Fits
Do not place Proxidize in front of a Zyte API request expecting to control the IP that reaches the target website. A proxy between your application and Zyte changes how your application reaches Zyte; it does not replace the network route Zyte selects between its infrastructure and the target.
Proxidize is relevant when the team operates its own collector:
Proxidize Residential Proxies fit custom collectors that need global country, city, and ISP targeting with rotating or sticky sessions. Proxidize Mobile Proxies fit workflows that specifically need US mobile-network exits and city or carrier controls.
The customer still owns browser behavior, crawling, retries, parsing, validation, and storage. This requires more engineering than a managed scraping API but gives the team more direct control over the collector and network configuration.
Choose Zyte when managed retrieval, rendering, extraction, or hosted browser infrastructure removes work your team does not want to operate. Choose raw proxies when an existing Requests, Scrapy, Playwright, Puppeteer, or agent stack already handles collection and needs standard endpoints plus more direct targeting and session controls.
A hybrid architecture can route straightforward collection through an in-house stack and send a smaller difficult segment to a managed API. Keep the routing rules explicit and measure accepted-output cost for each path.
How to Evaluate Zyte API
Test the exact workload instead of one convenient URL.
- Select representative server-rendered, JavaScript-heavy, paginated, localized, and failure-prone pages.
- Define the required output fields and what makes a record valid.
- Start with HTTP mode and add browser infrastructure only where it improves accepted output.
- Test requested geolocations against visible page context.
- Test sessions using the precise mode and sequence the application will use.
- Compare raw HTML, browser HTML, and automatic extraction on the same allowed sources.
- Measure API success, valid-record rate, latency, retries, and cost separately.
- Test invalid authentication, rate limits, timeouts, malformed requests, missing fields, and target changes.
- Review the provider's terms, data-handling documentation, retention, incident response, and subprocessors as part of procurement.
- Record versions, dates, request fields, sample composition, and limitations with the results.
One successful request does not establish performance across different websites, page types, regions, or volumes.
Common Zyte API Problems
| Symptom | Likely cause | What to check |
|---|---|---|
| `401` response | Missing or invalid API key | Confirm the correct Zyte API key is used as the Basic-auth username with an empty password |
| Request rejected with `400` | Invalid field combination, value, or unsupported request | Read the returned problem JSON and compare the request with the current API reference |
| `429` response | Account or request rate limit | Bound concurrency, respect retry timing, and check current account limits |
| `httpResponseBody` looks unreadable | It is still Base64-encoded | Decode it before parsing the target response |
| JavaScript content is missing | HTTP mode returned the pre-rendered response | Test `browserHtml` or an appropriate browser/CDP workflow |
| Browser request cannot send the desired method or headers | Browser mode limits the initial request controls | Use HTTP mode or CDP when the workflow needs those controls |
| Proxy mode lacks a feature | Proxy mode does not expose the full HTTP API | Use `/v1/extract` for browser output, actions, screenshots, or richer response data |
| Session state is not retained | Session model or identifier does not match the selected mode | Review client-managed, server-managed, proxy-mode, and CDP session differences |
| Location looks wrong | Requested country did not produce the expected page context | Verify the returned language, currency, store, or page evidence; location is only one context signal |
| Spend is higher than expected | Browser mode, target tier, actions, extraction, screenshots, extended location, or network use changed the cost | Inspect usage by request type and reproduce the workload in the current cost estimator |
| API call succeeds but no records are saved | Selector, extraction schema, or content validation failed | Save reviewable response evidence and treat empty output as a failed observation |
Retries should be bounded and limited to failures that may actually be transient. Retrying invalid authentication, unsupported fields, or a deterministic validation failure only consumes time and may increase cost.
Is Zyte API Safe to Use?
Using a managed API reduces some operational work, but it does not remove security or governance responsibilities.
- Keep API keys out of source control and client-side applications.
- Restrict which services and people can read production credentials.
- Avoid placing secrets, personal data, or unnecessary sensitive values in request URLs or logs.
- Validate target URLs before allowing untrusted users or agents to submit them.
- Protect internal networks from server-side request forgery in any surrounding collection service.
- Set spending alerts and limits appropriate to the environment.
- Retain only the response data and evidence the workflow needs.
- Review Zyte's current security, privacy, permissions, and subprocessor information before procurement.
A scraping API also does not grant a right to collect or use data. Collect only information your organization is authorized to access, follow applicable laws and relevant contractual terms, respect privacy requirements, and use reasonable request rates. This article provides technical information, not legal advice.
Final Verdict: Is Zyte API Worth Using?
Zyte API is worth evaluating when a team wants to reduce the work involved in retrieving and rendering web pages. Its main advantage is the range of access layers behind one API: lightweight HTTP requests, managed browser output, browser actions, structured extraction, Scrapy integration, and a remote CDP browser.
The tradeoff is abstraction. Zyte decides more of the network and retrieval path, its modes have different capabilities, and the final cost depends on the target and requested features. Customers must still validate the content and operate the data pipeline around it.
It is strongest for developers, Scrapy teams, and data pipelines that want managed HTTP and browser access. It is less suitable when an official data source already answers the question, city- or network-level proxy targeting is essential, or the team already operates an efficient collector and needs only raw network access.
If you are comparing other managed APIs, workflow platforms, or a raw-proxy architecture, see the best Zyte alternatives in 2026 for a use-case-based shortlist.
Choose Zyte for managed retrieval and extraction. Choose raw proxies when your team wants to operate the scraper and control the network layer more directly. Test both decisions against accepted records—not marketing feature counts or successful HTTP responses alone.
Frequently asked questions
Zyte API is a managed web scraping API that can retrieve HTTP responses, render pages in a browser, take screenshots, perform browser actions, maintain supported sessions, and return structured extraction results. The customer sends an authenticated request and then validates and uses the response.
It includes managed proxy routing and offers a compatibility proxy mode, but its main HTTP API is broader than an ordinary raw proxy. It can return rendered HTML, screenshots, network captures, and structured data rather than only forwarding arbitrary client traffic through a selected exit.
Crawlera was renamed Zyte Smart Proxy Manager. Smart Proxy Manager is now end-of-life, and Zyte migrated remaining traffic to Zyte API Proxy Mode through a compatibility layer. New integrations should use current Zyte API documentation.
Use HTTP Basic authentication with the Zyte API key as the username and an empty password. Store the key in an environment variable or secret manager rather than hardcoding it.
Yes. A Python HTTP client can send a JSON POST request to https://api.zyte.com/v1/extract. Zyte also documents a Python client, and Scrapy users can install scrapy-zyte-api.
Browser requests can execute JavaScript and return the rendered DOM. For more direct control, Zyte also exposes a CDP browser that works with Playwright, Puppeteer, and other CDP-compatible clients.
Yes, through Zyte's CDP headless-browser interface. Current documentation says CDP requires a subscription or a PAYG account with a spending limit plus business verification and is not available through the initial free credit alone.
Yes. Zyte can select a location automatically or accept a supported country-level geolocation override. Some extended geolocations can affect pricing. Verify the actual page context because IP country is not the only localization signal.
Request cost depends on the target website, HTTP-versus-browser tier, and added features. Actions, screenshots, extraction, custom attributes, extended geolocation, device residential IP use, network capture, and CDP can have separate cost rules. Check the live estimator for the representative workload.
They solve different layers. Zyte API is a managed retrieval and optional extraction service. Residential proxies provide network access for a scraper or browser the customer operates. The better choice depends on which infrastructure the team wants to manage.
Normally, no. Zyte manages the target-side access route. Putting another proxy between your application and Zyte does not give you control over the exit Zyte uses to reach the target.