
To collect Google search results programmatically, first choose the data source that matches the question. Use Google Search Console for performance data about a verified property, a managed SERP API for structured public-result observations, or an in-house collector only when direct automated access is appropriate and permitted. Proxies can supply location-specific network routes for a custom collector, but they do not grant access rights or guarantee a particular SERP.
Quick Answer
The most maintainable way to scrape Google search results in 2026 is usually a managed SERP API that returns structured JSON. Search Console is better for clicks, impressions, CTR, and average position on your own properties. A custom browser collector offers more control but also requires retrieval, parser, proxy, session, compliance, and change-management work.
Key Takeaways
- “Google search data” can mean first-party Search Console metrics, a live public SERP observation, or a finished rank-tracking report. These are different datasets.
- A rank observation needs query, engine, location, language, device, result type, depth, method, and timestamp—not only a URL and position.
- Managed SERP APIs return structured data and operate retrieval; raw proxies provide only the network route for your own collector.
- Google’s Custom Search JSON API is closed to new customers, and Google says existing customers must transition by January 1, 2027.
- Direct Google page markup changes frequently. Treat parser failures, challenge pages, and empty results as failures, not valid ranking data.
- A proxy IP in a city does not prove that every search-localization signal matches that city.
- Automated collection must comply with applicable law, Google’s current terms and technical rules, and any agreement governing the data or service.
What Does “Scrape Google Search Results” Mean?
The phrase covers several different jobs:
- Owned-site performance: retrieve queries, pages, clicks, impressions, CTR, and average position for a property you can access.
- Live SERP observation: record which results and features appear for a configured query, location, language, device, and time.
- Rank tracking: schedule repeated observations, match domains or URLs, preserve history, calculate changes, and report them.
- Market or feature research: observe ads, local packs, answer boxes, news, shopping results, AI features, or competitor domains.
The retrieval method should follow the job. Search Console is authoritative for a verified property’s recorded performance, but it does not reproduce a single live public result page. A SERP API can return an on-demand observation, but it does not replace Search Console’s click and impression data.
Choose the Right Collection Method
| Method | Best for | What you receive | Main limitation |
|---|---|---|---|
| Google Search Console API | Performance of properties you can access | Clicks, impressions, CTR, average position, query/page dimensions | Not a complete live public SERP |
| Existing Google Custom Search JSON API account | Programmable Search Engine results | Structured web or image results | Closed to new customers; transition required by January 1, 2027 |
| Managed SERP API | On-demand structured result observations | Parsed organic and rich-result data | Vendor pricing, schemas, and throughput limits |
| Rank-tracking platform | Projects, schedules, charts, and reports | Finished monitoring workflow | Less low-level control |
| In-house HTTP/browser collector | Custom permitted observations | Raw response and your parsed output | Highest engineering and policy burden |
| Manual check | Small spot checks | Human observation | Not scalable or reproducible by itself |
Google Search Console API
The Search Console API is the correct starting point when the question is “How did my site perform in Google Search?” It requires suitable permission for the property. Google documents dimensions such as query, page, country, and device and metrics including clicks, impressions, CTR, and average position.
Search Console data has reporting limits and aggregation behavior. Average position is not the same as a manually observed rank, and Google notes that the Search Analytics API may expose top rows rather than every possible row. Use it for what Google records about your property—not as a public SERP scraper.
Google Custom Search JSON API
Google’s Custom Search JSON API documentation says the service is closed to new customers. Existing customers can use it until January 1, 2027 and must transition to another solution. It retrieves results from a configured Programmable Search Engine, so it should not be presented as a new general-purpose full-Google SERP option.
Managed SERP API
A managed SERP API accepts a query and context parameters, operates retrieval and parsing, and returns a structured response. It is usually the shortest path for developers building a custom search-monitoring product without maintaining browsers and selectors.
Our SerpApi review explains one prominent service’s pricing, hourly throughput, schema, and tradeoffs. The broader rank tracker API guide compares data APIs with finished tracking workflows.
In-house collector
An in-house collector gives the team control over the client, proxy, browser state, parser, and evidence. It also makes the team responsible for reliability, policy review, access-denied handling, change detection, and all infrastructure costs.
Do not choose it because a short selector demo looks free. The operating cost includes browser compute, bandwidth, proxy use where needed, retries, parser maintenance, monitoring, and invalid data.
What a Valid SERP Observation Contains
A position without context cannot be reproduced.
For each result, store at least:
- result type, such as organic, ad, local, news, shopping, or answer feature;
- position within that result type;
- title and destination URL;
- displayed domain or source when available;
- matched canonical domain or tracked entity;
- query and observation context;
- source response ID, hash, or permitted snapshot;
- parser and schema version;
- validation status.
Do not flatten every feature into one universal “rank” without documenting the rule. Organic position 3, a local-pack position, and an AI-feature citation represent different surfaces.
Python Example Using a Managed SERP API
This example calls SerpApi’s documented Google endpoint and writes normalized organic results to JSON. SerpApi is not an official Google API. The example is included because it provides a current structured interface and avoids publishing another brittle Google CSS-selector script.
This code was syntax-checked on October 1, 2026. It was not run with a live SerpApi account, so the article does not claim an observed response or performance result.
Install Requests
Set the API key
Complete script
Save this as collect_google_results.py:
Run it:
The validation is intentionally strict. If the expected list or required result fields disappear, the script fails rather than writing incomplete rows as valid rankings.
Why the Old Selenium Selector Pattern Is Fragile
A common tutorial launches Chrome, sleeps for three seconds, and selects hard-coded classes such as a result-card class copied from one page. That pattern has several problems:
- Google markup and result features change;
- consent, localization, or access-denied pages may replace the result;
- a fixed sleep does not prove that required content loaded;
- visual order across mixed feature types is not captured by one selector;
- a successful process can export zero or partial records;
- browser and driver instructions age quickly;
- anti-detection arguments do not create permission or data quality.
If an authorized use case truly requires a browser, wait for defined evidence, validate the page type, save failure artifacts, and maintain parser fixtures. Use current Playwright or Selenium browser management rather than a hard-coded historical ChromeDriver download.
If You Operate a Direct Collector
Separate the responsibilities:
| Component | Responsibility |
|---|---|
| Scheduler | Decides which query/context combinations run and when |
| HTTP client or browser | Retrieves the configured result |
| Proxy, when needed | Supplies the request’s network exit and location |
| Parser | Converts the received page or response into fields |
| Validator | Rejects challenges, missing fields, wrong context, and malformed data |
| Storage | Preserves observations and evidence |
| Rank logic | Matches tracked sites and compares compatible observations |
| Monitoring | Measures validity, latency, failures, and cost |
Use bounded concurrency and a per-host schedule. Treat access-denied responses and CAPTCHAs as stop-and-investigate conditions rather than an instruction to escalate evasion. Do not retry a policy block as if it were a transient network error.
Where Proxies Fit
Raw proxies matter only when you operate the collector:
With a managed SERP API, the provider operates its own retrieval infrastructure; you do not normally insert a Proxidize endpoint between your application and the provider’s API.
A proxy can change network-derived properties such as the observed public IP and its approximate location. It does not guarantee the requested city SERP because Google can use domain, language, device, account state, cookies, settings, and other context.
Residential proxies
Proxidize Residential Proxies are the main fit for teams operating their own permitted global collector. They support country, city, and ISP targeting, rotating or sticky sessions, and HTTP, HTTPS, and SOCKS5 connections across 195+ countries.
Mobile proxies
Proxidize Mobile Proxies provide a US mobile-network route and city or carrier targeting. A mobile IP does not itself create a mobile SERP; set the browser or request device context separately.
Rotating and sticky behavior
Use separate rotating sessions for independent observations. Use a sticky session when one observation includes dependent page loads or interactions. Recheck the exit and location within the same client when those properties are material.
For provider and architecture decisions, see best proxies for SEO monitoring and IP rotation.
Location, Language, and Device Consistency
Use a test matrix rather than changing one setting at a time without recording it.
| Dimension | Example value | Why it matters |
|---|---|---|
| Query | `coffee shops` | Defines the information need |
| Search domain | `google.com` | Affects market and interface behavior |
| Country | US | Broad market signal |
| City/location | Austin, Texas | Local-result context |
| Language | English | Interface and result-language preference |
| Device | Desktop | Result layout and feature availability |
| Depth | First result page | Defines observation scope |
| Method | Managed API or named collector version | Supports reproducibility |
| Time | UTC timestamp | Captures volatility |
Do not compare a mobile observation in London with a desktop observation in New York and label the difference a rank change. Normalize the dimensions first.
Organic Results Are Only One SERP Type
Modern search pages can include organic links, ads, local packs, maps, news, images, video, shopping, answer features, related questions, and AI-generated surfaces. A collector should preserve the result type and its own position rather than force every object into an organic list.
For each feature, define:
- whether it is in scope;
- its identity fields;
- how positions are counted;
- whether citations or links are extracted;
- how absence differs from a collection failure;
- how schema changes are reviewed.
Testing a Google SERP Collection System
Use a stable evaluation set containing:
- branded and nonbranded queries;
- local and nonlocal queries;
- ambiguous queries;
- queries that trigger different features;
- known no-result or sparse-result cases;
- tracked and untracked domains;
- multiple cities, languages, and devices relevant to the business.
Measure:
- valid observation rate;
- requested-to-observed context match;
- result and feature coverage;
- parser accuracy against a reviewed sample;
- missing or duplicate records;
- median and p95 latency;
- retry rate;
- cost per valid observation;
- change-detection false-positive rate.
An API response count or HTTP success rate is not enough.
Common Problems
| Symptom | Likely cause | What to do |
|---|---|---|
| Empty organic list | Query has no organic results, parser changed, or wrong page returned | Inspect metadata and raw evidence; classify the case |
| Wrong local results | Location parameters or application context do not match | Validate requested and observed context |
| Rank jumps unexpectedly | Device, location, feature mix, or parser changed | Compare only normalized observations |
| CAPTCHA or access-denied page | Source rejected or challenged the request | Stop, record the failure, reduce load, and review the method |
| Results differ from a manual browser | Personalization, time, settings, method, or location differs | Reproduce every relevant dimension |
| Costs exceed forecast | Empty successful searches, retries, browser assets, or unused plan allowance | Track cost per accepted observation |
| Proxy IP changed mid-observation | Rotation rule or peer availability | Use a suitable sticky session and verify before/after |
| Search Console differs from observed rank | Aggregated historical metric compared with one live observation | Treat the datasets as different measurements |
Compliance and Data Governance
Review Google’s current terms, service-specific rules, applicable law, data licenses, privacy obligations, and the intended downstream use. Avoid collecting private or authenticated data without authorization. Do not use proxies, browser changes, or CAPTCHA services as a justification to ignore a denial of access.
Minimize retained data, protect API and proxy credentials, define retention periods, and keep provenance for business-critical observations. For regulated or high-risk uses, obtain qualified legal review.
Choose the Data Source Before the Proxy
Start by deciding whether you need owned-site performance, an on-demand public result observation, or a finished monitoring workflow. Search Console, a managed SERP API, and a rank tracker often solve those needs with less operational work than direct collection. Use raw proxies only when your team intentionally owns the collector and the network route is a real requirement.
For the wider operational model, see Proxidize’s SEO monitoring guide or compare raw proxies and scraping APIs.
Frequently asked questions
Search-result data can be collected through several methods, but the appropriate method depends on the data, permissions, Google’s current terms, and applicable law. Search Console, managed SERP APIs, rank trackers, and direct collectors have different scopes and responsibilities.
Google offers Search Console APIs for verified-property performance, but that is not a full live public-SERP API. The Custom Search JSON API is closed to new customers, and Google says existing customers must transition by January 1, 2027.
A managed SERP API is usually the simplest current method because it returns structured data and operates retrieval and parsing. The example in this guide uses Requests and a SerpApi endpoint.
No. SerpApi is a third-party managed search-results API. Review its current documentation, pricing, terms, and response schema independently.
Not when using Search Console or a managed SERP API that handles retrieval. Proxies are relevant when your team operates its own permitted collector and needs location-specific or separate network routes.
For an in-house global collector, residential proxies are often the practical starting point because of geographic coverage. US mobile proxies fit mobile-network observations. Test location accuracy, valid-result rate, latency, sessions, and cost rather than choosing by label alone.
No. IP location is one input. Domain, language, device, cookies, account state, settings, and other signals can affect the result. Validate the actual observation context.
Store feature type and feature-specific position separately. Document whether your product reports organic position, absolute visual order, local-pack position, or another metric. Do not mix them without a defined rule.
It can answer many owned-site performance questions and should be used where appropriate. It does not show a complete configured live SERP or every competitor result, so some workflows use both datasets.
Base the schedule on business need, result volatility, source rules, and cost. Important volatile queries may justify more frequent checks; stable long-tail queries can be sampled less often.