Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Python can turn a keyword list into structured research data by querying SEO APIs in batches. The important part is not just retrieving a volume or difficulty number: store each metric with its provider, market, definition, and retrieval date. Those details determine what the number means and whether it can be compared with another snapshot.

For example, start with a small list such as wireless headphones, best noise cancelling headphones, and headphones for travel. Set the country and language explicitly before requesting data; the same phrase can have different demand and search results in different markets.

What a bulk keyword-research workflow should return

A useful output is one row per keyword and market, not a spreadsheet of unqualified scores. Preserve the request context and metric definitions alongside the values so later analysis does not treat different providers or dates as interchangeable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Field What to record
Keyword The exact query submitted, preserving spelling and punctuation.
Provider and endpoint The service, endpoint name, and API version used to retrieve the result.
Market Country or location, language, and search engine or database when the endpoint exposes them.
Volume The returned value plus its window and source, such as average monthly volume over a stated period or a latest-month value.
Difficulty The score, scale, and provider’s definition.
Other signals Intent, CPC, competition, SERP features, and historical values when returned and needed.
Snapshot context Retrieval timestamp and, where supplied, the SERP’s own last-update date.

Do not infer a missing field or silently fill it with another provider’s value. If an endpoint does not return a field, keep it absent or mark it as unavailable in your own data model.

Choose an API by the fields and workflow you need

Endpoint limits are not provider-wide limits. Check the specific endpoint and current documentation before sizing batches or building around a field.

Provider and documented option Relevant capabilities and qualifications
Ahrefs API Overview Country-scoped requests can return estimated volume, latest-month volume, keyword difficulty, SERP features, device shares, and the SERP last-update date. The feature list includes ai_overview. The documentation includes a Python requests example: Ahrefs API Overview.
Ahrefs Keywords Explorer interface The Help Center says one search can accept up to 10,000 keywords pasted in or uploaded as .txt or .csv. It describes advanced metrics as consuming one credit per keyword; this is a UI capability and credit detail, not an API batch limit. See Ahrefs Keywords Explorer bulk search.
Semrush Keyword Reports v4 Can return volume, difficulty, intent, CPC, competition, trends, and SERP features, including AI Overview. Semrush labels v4 Early Access and warns endpoints, response formats, and pricing may change before general availability. See Semrush Keyword Reports v4.
Semrush v3 Batch Keyword Overview Its regional-database batch overview documents up to 100 keywords and returns volume, CPC, competition, and number of results. Semrush says older v3 methods are deprecated and not recommended for new integrations, although existing use continues temporarily. See Semrush v3 API documentation.
DataForSEO Labs Keyword Overview Accepts batches up to 700 and can return CPC, paid competition, volume, intent, SERP, backlink, and clickstream data. Its Bulk Keyword Difficulty endpoint accepts up to 1,000 and returns a proprietary 0–100 score relative to Google’s top ten. See DataForSEO Labs documentation.
DataForSEO bulk endpoints The provider’s guide lists up to 1,000 keywords per request for Google Ads Search Volume, Bulk Clickstream Search Volume, Labs Bulk Difficulty, and Search Intent, and demonstrates a Python-oriented flow. Google Ads data and the provider’s proprietary database metrics are distinct sources. See DataForSEO bulk workflow guide.
DataForSEO Historical Keyword Data Accepts up to 700 keywords per request and documents historical data reaching back to the beginning of 2019. Use it separately from current overview data when a trend series is required. See DataForSEO Historical Keyword Data.

These are vendor-documented features, not a controlled comparison of accuracy or speed. DataForSEO also says its Google Keyword Database draws on sources including Google Ads and Google SERPs, and that updates occur gradually in the latter part of each month as Google’s Ads update cycle progresses; this is a stated pattern, not a promise that every term refreshes at once. See DataForSEO Google Keyword Database.

Send a Python request and preserve the response

Use the provider’s official endpoint, authentication scheme, request fields, and response schema. Ahrefs publishes a Python requests example in its API documentation, while DataForSEO’s bulk guide describes a Python-oriented request flow. Because endpoints and schemas differ, there is no single request body that works across these services.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Prepare the input. Deduplicate the keyword list without changing the original phrases. Set country or location and language explicitly using the exact parameters supported by the chosen endpoint.
  2. Load credentials securely. Keep API keys outside the script and source control, for example in environment variables or a secrets manager. Use only the authentication method documented by the provider.
  3. Batch to the endpoint’s limit. The documented maximums vary: for example, DataForSEO Labs Overview documents 700 per batch while its Bulk Keyword Difficulty endpoint documents 1,000. Respect any lower account, task, or rate limits that apply.
  4. Check each response. Handle HTTP failures, provider-level errors, partial results, and rate limits explicitly. Retry transient failures with backoff rather than immediately resubmitting a whole batch without checking its status.
  5. Store raw and normalized data. Retain the response payload with request parameters and a retrieval timestamp, then map returned fields into your normalized schema. Raw data makes later audits and parser updates possible.
  6. Validate before analysis. Confirm that every requested keyword has a corresponding result or recorded error, and that the result’s market and endpoint match the request.

For repeatable jobs, log the request’s endpoint and version, batch identifiers if provided, and the provider’s status or error details. Avoid embedding credentials in notebooks, logs, or saved request examples.

Interpret search volume before sorting keywords

Search volume is an estimate, not a universal count of searches. Its meaning depends on the provider, source, market, and time window. Ahrefs documents average monthly volume over the latest known 12 months as well as a separate latest-month field. Those fields answer different questions: one smooths a longer period, while the other reflects the most recent month available. See Ahrefs API Overview.

DataForSEO distinguishes Google Ads data from metrics calculated from its own keyword and SERP databases. A pipeline should preserve which source produced the value rather than combining all volume fields into a single column without qualification. Its database documentation describes gradual updates in the latter part of each month; do not assume that every keyword was refreshed simultaneously.

  • Keep the exact volume field name and any documented window.
  • Record provider, source, country or location, and retrieval time with the number.
  • Do not compare two volume values as if they share a methodology simply because both are labeled monthly volume.
  • Use historical endpoints when you need a series; a current overview value alone is not a trend.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Interpret keyword difficulty as a provider-specific estimate

Difficulty scores are estimates, and matching 0–100 scales do not make them equivalent. Ahrefs defines KD as an estimate of the difficulty of ranking in Google’s top ten, based on referring domains to top-ten organic pages; it says the score does not include on-page SEO factors. See Ahrefs KD definition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DataForSEO describes its own Bulk Keyword Difficulty score as a proprietary 0–100 estimate relative to the current Google top ten. Semrush documents a keyword-difficulty report estimating difficulty in Google’s top ten, but that does not establish equivalence with another provider’s algorithm. See DataForSEO Labs documentation and Semrush v3 API documentation.

Use difficulty as one prioritization signal alongside intent, your site’s relevance and authority, and the actual search results. Do not average vendor scores or apply a threshold learned from one provider to another without a separately validated reason.

Detect AI Overviews as a SERP feature

Where the selected provider’s result schema includes SERP features, check the returned feature list for the provider’s AI Overview identifier. Ahrefs lists ai_overview among its SERP feature values, and Semrush v4 includes AI Overview in its feature list. The observation belongs to that provider’s data snapshot; store it with the keyword, market, retrieval timestamp, and SERP update information when available.

An AI Overview feature flag does not say that your site appears in the Overview, nor does it predict a click-through-rate change. It records a SERP characteristic, not a ranking or traffic outcome. Semrush’s v4 field availability also sits within its Early Access interface, whose endpoints, formats, and pricing may change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a useful comparison without claiming a winner

Choose an implementation based on the exact fields and operational constraints your project needs. The published capabilities above do not establish which provider has the most accurate volume or difficulty values.

  • Metric provenance: Identify whether volume comes from an advertising dataset, a provider database, or another disclosed source, and document how difficulty is defined.
  • Market coverage: Check that the endpoint supports your target country or location, language, and search engine/database.
  • Batching: Compare limits per endpoint, not just a headline number for the product. Plan for partial responses, throttling, and the cost or credit model where applicable.
  • Analysis fields: Confirm whether intent, history, backlinks, CPC, and SERP features are returned in the endpoint you will actually call.
  • Freshness: Look for update behavior and SERP timestamps. A provider’s general database cadence does not prove every keyword’s data has the same age.
  • Integration stability: Account for API version, deprecation status, schema-change risk, and how much response normalization your code must maintain.

For a new integration, prefer a currently supported endpoint and treat early-access or deprecated interfaces as a maintenance risk. Recheck the provider’s documentation for live limits, fields, and availability when implementing, since those details can change.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.