Skip to main content
Tech Tutorials & Programming

Published Apr 11, 2025 · Updated Oct 1, 2026

How to Scrape Google Search Results: APIs, Python, and Proxies

Compare safe ways to collect Google search results, use a structured Python API example, design valid rank observations, and learn where proxies fit.

How to Scrape Google Search Results: APIs, Python, and Proxies

To collect Google search results programmatically, first choose the data source that matches the question. Use Google Search Console for performance data about a verified property, a managed SERP API for structured public-result observations, or an in-house collector only when direct automated access is appropriate and permitted. Proxies can supply location-specific network routes for a custom collector, but they do not grant access rights or guarantee a particular SERP.

Quick Answer

The most maintainable way to scrape Google search results in 2026 is usually a managed SERP API that returns structured JSON. Search Console is better for clicks, impressions, CTR, and average position on your own properties. A custom browser collector offers more control but also requires retrieval, parser, proxy, session, compliance, and change-management work.

Key Takeaways

  • “Google search data” can mean first-party Search Console metrics, a live public SERP observation, or a finished rank-tracking report. These are different datasets.
  • A rank observation needs query, engine, location, language, device, result type, depth, method, and timestamp—not only a URL and position.
  • Managed SERP APIs return structured data and operate retrieval; raw proxies provide only the network route for your own collector.
  • Google’s Custom Search JSON API is closed to new customers, and Google says existing customers must transition by January 1, 2027.
  • Direct Google page markup changes frequently. Treat parser failures, challenge pages, and empty results as failures, not valid ranking data.
  • A proxy IP in a city does not prove that every search-localization signal matches that city.
  • Automated collection must comply with applicable law, Google’s current terms and technical rules, and any agreement governing the data or service.

What Does “Scrape Google Search Results” Mean?

The phrase covers several different jobs:

  1. Owned-site performance: retrieve queries, pages, clicks, impressions, CTR, and average position for a property you can access.
  2. Live SERP observation: record which results and features appear for a configured query, location, language, device, and time.
  3. Rank tracking: schedule repeated observations, match domains or URLs, preserve history, calculate changes, and report them.
  4. Market or feature research: observe ads, local packs, answer boxes, news, shopping results, AI features, or competitor domains.

The retrieval method should follow the job. Search Console is authoritative for a verified property’s recorded performance, but it does not reproduce a single live public result page. A SERP API can return an on-demand observation, but it does not replace Search Console’s click and impression data.

Choose the Right Collection Method

MethodBest forWhat you receiveMain limitation
Google Search Console APIPerformance of properties you can accessClicks, impressions, CTR, average position, query/page dimensionsNot a complete live public SERP
Existing Google Custom Search JSON API accountProgrammable Search Engine resultsStructured web or image resultsClosed to new customers; transition required by January 1, 2027
Managed SERP APIOn-demand structured result observationsParsed organic and rich-result dataVendor pricing, schemas, and throughput limits
Rank-tracking platformProjects, schedules, charts, and reportsFinished monitoring workflowLess low-level control
In-house HTTP/browser collectorCustom permitted observationsRaw response and your parsed outputHighest engineering and policy burden
Manual checkSmall spot checksHuman observationNot scalable or reproducible by itself

Google Search Console API

The Search Console API is the correct starting point when the question is “How did my site perform in Google Search?” It requires suitable permission for the property. Google documents dimensions such as query, page, country, and device and metrics including clicks, impressions, CTR, and average position.

Search Console data has reporting limits and aggregation behavior. Average position is not the same as a manually observed rank, and Google notes that the Search Analytics API may expose top rows rather than every possible row. Use it for what Google records about your property—not as a public SERP scraper.

Google Custom Search JSON API

Google’s Custom Search JSON API documentation says the service is closed to new customers. Existing customers can use it until January 1, 2027 and must transition to another solution. It retrieves results from a configured Programmable Search Engine, so it should not be presented as a new general-purpose full-Google SERP option.

Managed SERP API

A managed SERP API accepts a query and context parameters, operates retrieval and parsing, and returns a structured response. It is usually the shortest path for developers building a custom search-monitoring product without maintaining browsers and selectors.

Our SerpApi review explains one prominent service’s pricing, hourly throughput, schema, and tradeoffs. The broader rank tracker API guide compares data APIs with finished tracking workflows.

In-house collector

An in-house collector gives the team control over the client, proxy, browser state, parser, and evidence. It also makes the team responsible for reliability, policy review, access-denied handling, change detection, and all infrastructure costs.

Do not choose it because a short selector demo looks free. The operating cost includes browser compute, bandwidth, proxy use where needed, retries, parser maintenance, monitoring, and invalid data.

What a Valid SERP Observation Contains

A position without context cannot be reproduced.

bash

For each result, store at least:

  • result type, such as organic, ad, local, news, shopping, or answer feature;
  • position within that result type;
  • title and destination URL;
  • displayed domain or source when available;
  • matched canonical domain or tracked entity;
  • query and observation context;
  • source response ID, hash, or permitted snapshot;
  • parser and schema version;
  • validation status.

Do not flatten every feature into one universal “rank” without documenting the rule. Organic position 3, a local-pack position, and an AI-feature citation represent different surfaces.

Python Example Using a Managed SERP API

This example calls SerpApi’s documented Google endpoint and writes normalized organic results to JSON. SerpApi is not an official Google API. The example is included because it provides a current structured interface and avoids publishing another brittle Google CSS-selector script.

This code was syntax-checked on October 1, 2026. It was not run with a live SerpApi account, so the article does not claim an observed response or performance result.

Install Requests

bash

Set the API key

bash

Complete script

Save this as collect_google_results.py:

python

Run it:

bash

The validation is intentionally strict. If the expected list or required result fields disappear, the script fails rather than writing incomplete rows as valid rankings.

Why the Old Selenium Selector Pattern Is Fragile

A common tutorial launches Chrome, sleeps for three seconds, and selects hard-coded classes such as a result-card class copied from one page. That pattern has several problems:

  • Google markup and result features change;
  • consent, localization, or access-denied pages may replace the result;
  • a fixed sleep does not prove that required content loaded;
  • visual order across mixed feature types is not captured by one selector;
  • a successful process can export zero or partial records;
  • browser and driver instructions age quickly;
  • anti-detection arguments do not create permission or data quality.

If an authorized use case truly requires a browser, wait for defined evidence, validate the page type, save failure artifacts, and maintain parser fixtures. Use current Playwright or Selenium browser management rather than a hard-coded historical ChromeDriver download.

If You Operate a Direct Collector

Separate the responsibilities:

ComponentResponsibility
SchedulerDecides which query/context combinations run and when
HTTP client or browserRetrieves the configured result
Proxy, when neededSupplies the request’s network exit and location
ParserConverts the received page or response into fields
ValidatorRejects challenges, missing fields, wrong context, and malformed data
StoragePreserves observations and evidence
Rank logicMatches tracked sites and compares compatible observations
MonitoringMeasures validity, latency, failures, and cost

Use bounded concurrency and a per-host schedule. Treat access-denied responses and CAPTCHAs as stop-and-investigate conditions rather than an instruction to escalate evasion. Do not retry a policy block as if it were a transient network error.

Where Proxies Fit

Raw proxies matter only when you operate the collector:

bash

With a managed SERP API, the provider operates its own retrieval infrastructure; you do not normally insert a Proxidize endpoint between your application and the provider’s API.

A proxy can change network-derived properties such as the observed public IP and its approximate location. It does not guarantee the requested city SERP because Google can use domain, language, device, account state, cookies, settings, and other context.

Residential proxies

Proxidize Residential Proxies are the main fit for teams operating their own permitted global collector. They support country, city, and ISP targeting, rotating or sticky sessions, and HTTP, HTTPS, and SOCKS5 connections across 195+ countries.

Mobile proxies

Proxidize Mobile Proxies provide a US mobile-network route and city or carrier targeting. A mobile IP does not itself create a mobile SERP; set the browser or request device context separately.

Rotating and sticky behavior

Use separate rotating sessions for independent observations. Use a sticky session when one observation includes dependent page loads or interactions. Recheck the exit and location within the same client when those properties are material.

For provider and architecture decisions, see best proxies for SEO monitoring and IP rotation.

Location, Language, and Device Consistency

Use a test matrix rather than changing one setting at a time without recording it.

DimensionExample valueWhy it matters
Query`coffee shops`Defines the information need
Search domain`google.com`Affects market and interface behavior
CountryUSBroad market signal
City/locationAustin, TexasLocal-result context
LanguageEnglishInterface and result-language preference
DeviceDesktopResult layout and feature availability
DepthFirst result pageDefines observation scope
MethodManaged API or named collector versionSupports reproducibility
TimeUTC timestampCaptures volatility

Do not compare a mobile observation in London with a desktop observation in New York and label the difference a rank change. Normalize the dimensions first.

Organic Results Are Only One SERP Type

Modern search pages can include organic links, ads, local packs, maps, news, images, video, shopping, answer features, related questions, and AI-generated surfaces. A collector should preserve the result type and its own position rather than force every object into an organic list.

For each feature, define:

  • whether it is in scope;
  • its identity fields;
  • how positions are counted;
  • whether citations or links are extracted;
  • how absence differs from a collection failure;
  • how schema changes are reviewed.

Testing a Google SERP Collection System

Use a stable evaluation set containing:

  • branded and nonbranded queries;
  • local and nonlocal queries;
  • ambiguous queries;
  • queries that trigger different features;
  • known no-result or sparse-result cases;
  • tracked and untracked domains;
  • multiple cities, languages, and devices relevant to the business.

Measure:

  • valid observation rate;
  • requested-to-observed context match;
  • result and feature coverage;
  • parser accuracy against a reviewed sample;
  • missing or duplicate records;
  • median and p95 latency;
  • retry rate;
  • cost per valid observation;
  • change-detection false-positive rate.

An API response count or HTTP success rate is not enough.

Common Problems

SymptomLikely causeWhat to do
Empty organic listQuery has no organic results, parser changed, or wrong page returnedInspect metadata and raw evidence; classify the case
Wrong local resultsLocation parameters or application context do not matchValidate requested and observed context
Rank jumps unexpectedlyDevice, location, feature mix, or parser changedCompare only normalized observations
CAPTCHA or access-denied pageSource rejected or challenged the requestStop, record the failure, reduce load, and review the method
Results differ from a manual browserPersonalization, time, settings, method, or location differsReproduce every relevant dimension
Costs exceed forecastEmpty successful searches, retries, browser assets, or unused plan allowanceTrack cost per accepted observation
Proxy IP changed mid-observationRotation rule or peer availabilityUse a suitable sticky session and verify before/after
Search Console differs from observed rankAggregated historical metric compared with one live observationTreat the datasets as different measurements

Compliance and Data Governance

Review Google’s current terms, service-specific rules, applicable law, data licenses, privacy obligations, and the intended downstream use. Avoid collecting private or authenticated data without authorization. Do not use proxies, browser changes, or CAPTCHA services as a justification to ignore a denial of access.

Minimize retained data, protect API and proxy credentials, define retention periods, and keep provenance for business-critical observations. For regulated or high-risk uses, obtain qualified legal review.

Choose the Data Source Before the Proxy

Start by deciding whether you need owned-site performance, an on-demand public result observation, or a finished monitoring workflow. Search Console, a managed SERP API, and a rank tracker often solve those needs with less operational work than direct collection. Use raw proxies only when your team intentionally owns the collector and the network route is a real requirement.

For the wider operational model, see Proxidize’s SEO monitoring guide or compare raw proxies and scraping APIs.

Frequently asked questions

Search-result data can be collected through several methods, but the appropriate method depends on the data, permissions, Google’s current terms, and applicable law. Search Console, managed SERP APIs, rank trackers, and direct collectors have different scopes and responsibilities.

Google offers Search Console APIs for verified-property performance, but that is not a full live public-SERP API. The Custom Search JSON API is closed to new customers, and Google says existing customers must transition by January 1, 2027.

A managed SERP API is usually the simplest current method because it returns structured data and operates retrieval and parsing. The example in this guide uses Requests and a SerpApi endpoint.

No. SerpApi is a third-party managed search-results API. Review its current documentation, pricing, terms, and response schema independently.

Not when using Search Console or a managed SERP API that handles retrieval. Proxies are relevant when your team operates its own permitted collector and needs location-specific or separate network routes.

For an in-house global collector, residential proxies are often the practical starting point because of geographic coverage. US mobile proxies fit mobile-network observations. Test location accuracy, valid-result rate, latency, sessions, and cost rather than choosing by label alone.

No. IP location is one input. Domain, language, device, cookies, account state, settings, and other signals can affect the result. Validate the actual observation context.

Store feature type and feature-specific position separately. Document whether your product reports organic position, absolute visual order, local-pack position, or another metric. Do not mix them without a defined rule.

It can answer many owned-site performance questions and should be used where appropriate. It does not show a complete configured live SERP or every competitor result, so some workflows use both datasets.

Base the schedule on business need, result volatility, source rules, and cost. Important volatile queries may justify more frequent checks; stable long-tail queries can be sampled less often.

Ready to launch?

Proxies built for real operations.

For teams that depend on stability, not luck.