Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable, currently documented way to automate Yandex text search is the Yandex Search API—not a script that downloads the consumer SERP page. The API accepts REST, gRPC, or the Yandex AI Studio SDK requests, returns XML or HTML, and supports controls for language, region, ranking, grouping, filtering, and pagination. This guide shows complete Python and Node.js REST clients, explains authentication and response decoding, and covers the limits and failure modes you need to handle in production.

People commonly use “scrape” to mean any automated extraction. Here, that means retrieving search results programmatically. Directly requesting and parsing Yandex’s public results HTML is a separate technique with different terms and risks. The old Yandex.XML license says it became void on November 1, 2024 and describes restrictions on other automated request methods; treat it as a legacy warning, not as current authorization. Check the current Search API terms, access requirements, limits, and pricing before deployment.

Choose the supported interface before writing code

Yandex documents three ways to submit a text-search request:

  • REST: ordinary HTTPS requests, a good fit for Python requests, Node.js fetch, queues, and serverless functions.
  • gRPC: useful when your service already standardizes on generated protobuf clients and long-lived connections.
  • Yandex AI Studio SDK: a supported client-library option when its language and release fit your stack.

The examples below use REST because it is easy to inspect and portable. The REST request uses CamelCase field names; gRPC uses snake_case equivalents. A synchronous response contains the result document in Base64-encoded rawData, so decoding is a required step before XML or HTML parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set up authentication and scope

Credentials and role

Every request must be authenticated. For a user or federated account, send an IAM token as a Bearer token and include the folder ID. A service account may authenticate with an IAM token or an API key in the Authorization header and can use its own folder. The account needs the search-api.webSearch.user role.

Keep tokens, API keys, and folder IDs in environment variables or a secret manager. Do not commit them to source control, put them in browser code, or log the complete Authorization header.

Environment variables

export YANDEX_IAM_TOKEN='your-iam-token'
export YANDEX_FOLDER_ID='your-folder-id'
# For a service account using an API key instead:
# export YANDEX_API_KEY='your-api-key'

Use either YANDEX_IAM_TOKEN or YANDEX_API_KEY, not both. The exact API endpoint and current account setup are defined in Yandex AI Studio’s Search API documentation; use the endpoint shown for your cloud region and account.

Understand the request parameters

The documented REST fields are CamelCase. A practical request normally includes:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Field Purpose Important qualification
searchType Chooses the search corpus and language context. Documented types include Russian, Turkish, international, Kazakh, Belarusian, and Uzbek.
queryText The text to search. Maximum documented length is 400 characters.
familyMode Applies family-content filtering. Choose deliberately for your audience.
page Selects a result page. Pagination is not an unlimited or guaranteed stable snapshot.
fixTypoMode Controls spelling correction. Record the setting when reproducibility matters.
sortMode, sortOrder Controls ranking/sort behavior. Do not assume a custom order equals the consumer SERP order.
groupMode, groupsOnPage, docsInGroup Controls grouping and the number of documents shown. Valid ranges differ between XML and HTML.
region Targets a geographic region. Supported only with Russian and Turkish search types.
l10n Sets localization of returned content. Keep it aligned with search type and user locale.
responseFormat Selects XML or HTML output. XML is the default; HTML can contain ads, quick responses, and other page elements.
resultsWithin Restricts the time window for results. Use only when your product needs recency filtering.

The API documents a maximum of 250 results per query. State your search type, language, region, family mode, and sorting settings in logs so another operator can reproduce the same intent.

Python: complete REST client

This example sends a synchronous XML request, decodes rawData, writes the document to disk, and extracts a few fields defensively. The endpoint path shown in your Yandex Search API project should replace API_ENDPOINT.

import base64
import json
import os
import requests
import xml.etree.ElementTree as ET

API_ENDPOINT = os.environ.get("YANDEX_SEARCH_API_ENDPOINT", "API_ENDPOINT")

def auth_headers():
    token = os.getenv("YANDEX_IAM_TOKEN")
    api_key = os.getenv("YANDEX_API_KEY")
    if token:
        return {"Authorization": f"Bearer {token}", "Content-Type": "application/json"}
    if api_key:
        return {"Authorization": f"Api-Key {api_key}", "Content-Type": "application/json"}
    raise RuntimeError("Set YANDEX_IAM_TOKEN or YANDEX_API_KEY")

def search_yandex(query, page=0):
    if len(query) > 400:
        raise ValueError("queryText is limited to 400 characters")
    payload = {
        "queryText": query,
        "searchType": "SEARCH_TYPE_RU",
        "familyMode": "FAMILY_MODE_MODERATE",
        "page": page,
        "fixTypoMode": "FIX_TYPO_MODE_ON",
        "sortMode": "SORT_MODE_BY_RELEVANCE",
        "sortOrder": "SORT_ORDER_DESC",
        "groupMode": "GROUP_MODE_FLAT",
        "groupsOnPage": 10,
        "docsInGroup": 1,
        "region": "225",
        "l10n": "ru",
        "responseFormat": "FORMAT_XML"
    }
    folder_id = os.getenv("YANDEX_FOLDER_ID")
    if folder_id:
        payload["folderId"] = folder_id
    response = requests.post(API_ENDPOINT, headers=auth_headers(), json=payload, timeout=60)
    response.raise_for_status()
    envelope = response.json()
    raw = envelope.get("rawData")
    if not raw:
        raise RuntimeError(f"Response has no rawData: {envelope.keys()}")
    document = base64.b64decode(raw).decode("utf-8")
    return document

xml_text = search_yandex("пример запроса")
with open("yandex-results.xml", "w", encoding="utf-8") as f:
    f.write(xml_text)

root = ET.fromstring(xml_text)
for item in root.findall(".//doc"):
    url = item.findtext("url") or ""
    title = item.findtext("title") or ""
    print(title, url)

Namespace names and element paths can differ by response version and settings. Inspect one saved response before writing a production parser. The documentation warns that fields may be absent and that response content may change without prior notice; treat every field as optional.

Requesting HTML instead

Change responseFormat to the documented HTML value, decode rawData in the same way, and parse the resulting document with an HTML parser such as Beautiful Soup. HTML is appropriate when you need page-level elements such as ads or quick responses, but those elements make selectors more fragile than a structured XML extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Node.js: complete REST client

This Node.js 18+ example uses the built-in fetch. It accepts either an IAM token or service-account API key, decodes the Base64 payload, and saves XML.

import { writeFile } from "node:fs/promises";

const endpoint = process.env.YANDEX_SEARCH_API_ENDPOINT || "API_ENDPOINT";
const token = process.env.YANDEX_IAM_TOKEN;
const apiKey = process.env.YANDEX_API_KEY;
const folderId = process.env.YANDEX_FOLDER_ID;

function headers() {
  if (token) return { Authorization: `Bearer ${token}`, "Content-Type": "application/json" };
  if (apiKey) return { Authorization: `Api-Key ${apiKey}`, "Content-Type": "application/json" };
  throw new Error("Set YANDEX_IAM_TOKEN or YANDEX_API_KEY");
}

async function searchYandex(query, page = 0) {
  if ([...query].length > 400) throw new Error("queryText is limited to 400 characters");
  const body = {
    queryText: query,
    searchType: "SEARCH_TYPE_RU",
    familyMode: "FAMILY_MODE_MODERATE",
    page,
    fixTypoMode: "FIX_TYPO_MODE_ON",
    sortMode: "SORT_MODE_BY_RELEVANCE",
    sortOrder: "SORT_ORDER_DESC",
    groupMode: "GROUP_MODE_FLAT",
    groupsOnPage: 10,
    docsInGroup: 1,
    region: "225",
    l10n: "ru",
    responseFormat: "FORMAT_XML",
    ...(folderId ? { folderId } : {})
  };
  const response = await fetch(endpoint, {
    method: "POST",
    headers: headers(),
    body: JSON.stringify(body),
    signal: AbortSignal.timeout(60000)
  });
  if (!response.ok) throw new Error(`Yandex HTTP ${response.status}: ${await response.text()}`);
  const envelope = await response.json();
  if (!envelope.rawData) throw new Error("Response has no rawData");
  return Buffer.from(envelope.rawData, "base64").toString("utf8");
}

const xml = await searchYandex("пример запроса");
await writeFile("yandex-results.xml", xml, "utf8");
console.log(xml);

For HTML, request the HTML response format and pass the decoded string to an HTML parser. Avoid brittle regular expressions for either format.

Equivalent cURL request for diagnostics

cURL is useful for separating credential or endpoint problems from application code. Substitute the endpoint and the authentication scheme documented for your account:

curl -sS -X POST "$YANDEX_SEARCH_API_ENDPOINT" 
  -H "Authorization: Bearer $YANDEX_IAM_TOKEN" 
  -H "Content-Type: application/json" 
  -d '{"queryText":"пример запроса","searchType":"SEARCH_TYPE_RU","familyMode":"FAMILY_MODE_MODERATE","page":0,"responseFormat":"FORMAT_XML","folderId":"'"$YANDEX_FOLDER_ID"'"}'

Pagination, synchronous mode, and deferred mode

Pagination

Increment page only while the response contains usable results, stop at the documented 250-result ceiling, and deduplicate URLs because ranking and grouping can change between calls. Store the exact query and settings with each page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Deferred processing

The API also supports deferred mode. Instead of the completed document, the initial response returns an operation object. Persist its ID, poll or track that operation, and read the result only after done becomes true. Use a bounded retry schedule and a deadline; do not poll indefinitely. Design the worker so a restart can resume an operation from its stored ID.

Parsing safely and preserving meaning

  • Check HTTP status before decoding the envelope.
  • Validate that rawData exists and is valid Base64.
  • Decode as UTF-8 for XML, the documented default.
  • Allow missing titles, snippets, URLs, groups, and metadata.
  • Keep the original response for audits, then normalize into your own schema.
  • Expect content and field structure to change without prior notice; monitor parser failures rather than silently dropping records.

Do not interpret a missing field as proof that a result has no such property. It may simply be omitted for that response or version.

Geography, language, and reproducibility

searchType, l10n, and region materially affect results. The documented search types cover Russian, Turkish, international, Kazakh, Belarusian, and Uzbek contexts. Region selection is supported only for Russian and Turkish search types. A Russian query with a Russian region is not interchangeable with an international query, even when the words are identical. Include these settings in test fixtures and report them alongside exported results.

Performance, reliability, and cost controls

  • Reuse HTTP connections where your client supports it and set explicit connect/read timeouts.
  • Use deferred mode for work that can leave the request path; use synchronous mode for interactive, bounded requests.
  • Retry only transient failures with exponential backoff and a maximum attempt count. Do not retry authentication or malformed-request errors blindly.
  • Cache identical query-and-settings combinations when freshness requirements allow it.
  • Throttle concurrency to the limits and quotas assigned to your account; the reviewed documentation does not establish a universal rate number.
  • Record status, latency, page, response format, and operation ID without logging secrets.
  • Plan for the 250-result maximum and for ranking changes; this is not a promise of an immutable search snapshot.

Current pricing, quotas, and access conditions can change. Confirm them in the Yandex account and API documentation before budgeting a production crawler.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Fix
401 or 403 Expired token, wrong authorization prefix, missing role, or unavailable folder. Issue a fresh credential, use the documented Bearer or Api-Key header, verify search-api.webSearch.user, and check folder access.
Invalid argument Wrong CamelCase field, unsupported enum, overlong query, or incompatible region/search type. Validate names and enum values, keep queryText within 400 characters, and use region only with Russian or Turkish search types.
JSON parses but results are empty rawData was not decoded, the page is beyond available results, or filters are restrictive. Decode Base64, inspect the saved document, try page zero, and review family/group settings.
XML parser error HTML was requested, the payload was truncated, or the response includes an unexpected structure. Confirm responseFormat, verify the decoded bytes, and choose an HTML parser for HTML output.
Deferred job never completes Polling too aggressively, losing the operation ID, or hitting a transient service error. Persist the ID, poll with backoff, enforce a deadline, and record the final operation status.
Parser breaks after an API update Optional fields or response structure changed. Use defensive access, schema fixtures, alerts, and a parser that tolerates absent fields.

Direct SERP HTML scraping: why it is different

A browser or HTTP client aimed at the consumer Yandex results page is not the documented Search API integration. It can encounter consent screens, JavaScript rendering, bot checks, changing markup, and terms that differ by service and location. Yandex Webmaster’s Allow/Disallow guidance concerns how site owners instruct crawlers to access their own sites; it is not permission to automate requests to Yandex Search. If you have a legitimate requirement for consumer-page automation, obtain current written permission and verify the applicable terms first.

Or skip the browser setup

If your actual need is a clean visual capture of a Yandex results page rather than structured ranking data, ScreenshotNeo makes one API call. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in X-Page-Verdict and X-Billed headers. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://yandex.com -o shot.webp

See the ScreenshotNeo API documentation for options such as viewport and device presets, full-page lazy-image loading, CSS-selector element capture, dark mode, custom JavaScript and CSS, click-before-capture, waits, request blocking, headers and cookies, timezone and geolocation, PDF output, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and usage reporting. Every feature is on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free.

Python equivalent:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://yandex.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js equivalent:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://yandex.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Create a free ScreenshotNeo account to use the 1,000 monthly screenshots without adding a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Production checklist

  • Use the documented Search API interface and confirm current terms.
  • Assign search-api.webSearch.user and test authentication with a minimal query.
  • Keep secrets outside source code and logs.
  • Choose XML or HTML intentionally and decode synchronous rawData.
  • Record search type, locale, region, family mode, sorting, grouping, and page.
  • Cap collection at 250 results per query and deduplicate across pages.
  • Handle absent fields and changing response structures.
  • Use bounded retries, timeouts, caching, and deferred operations where appropriate.

Frequently Asked Questions

Can I use the same parser for XML and HTML responses?

No. Decode both from rawData, then use an XML parser for XML and an HTML parser for HTML; their structures and optional elements differ.

What is the maximum query length?

The Yandex Search API documentation sets queryText at a maximum of 400 characters.

Does page 20 guarantee the same results tomorrow?

No. Pagination is bounded and search content or ranking can change, so store the query settings and retrieval time with each result set.

Is Yandex.XML still the API to build against?

No. Its license page states that document became void on November 1, 2024. Use the currently documented Search API and verify its present terms instead.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.