Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not build a headless-browser scraper for Tumblr posts unless Tumblr has given you express prior written permission. Tumblr’s Terms of Service prohibit scraping without that permission and restrict automated access to Tumblr’s published interfaces or access allowed by robots.txt or another robot-exclusion mechanism. Tumblr’s developer agreement separately prohibits page scraping when it is used to create application capabilities beyond those provided by the API or Firehose. For a permitted project, use Tumblr’s documented API at https://api.tumblr.com; if the API does not expose what you need, ask Tumblr for written authorization or a supported data arrangement rather than trying to evade the restriction with browser automation.

What Tumblr allows

Tumblr’s “Limitations on Automated Use” clause says, in part, that without express prior written permission you may not access or search the service by means other than Tumblr’s currently available, published interfaces (or an applicable robot-exclusion permission), and you may not scrape the service or its content. The developer agreement adds a specific warning for app builders: downloading and parsing whole Tumblr pages to create capabilities beyond the API or Firehose is “page scraping” and is prohibited.

Those are separate constraints. A robots.txt entry does not automatically cancel the Terms of Service prohibition on scraping, and having a technical ability to render a page does not establish permission. This article therefore does not provide a browser-scraper recipe. It shows the supported API workflow, how to handle its authentication and limits, and what to do when the API cannot satisfy a legitimate requirement.

Use the official Tumblr API for permitted collection

Tumblr’s API is hosted at https://api.tumblr.com. Blog resources use paths such as /v2/blog/{blog-identifier}/...; post retrieval is performed through the blog posts route. Register an application to obtain a consumer/API key, then check the authentication requirement shown for the exact method you plan to call.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Access level What it means What to verify before coding
No authentication The method can be called without credentials. Confirm that the specific endpoint and requested fields are documented as public.
API key The request identifies your registered application. Keep the key server-side and send it exactly as the endpoint documentation specifies.
OAuth-signed or user-authorized The method requires a signed request and, often, a user grant. Implement the required OAuth flow; do not substitute a browser session or copied cookie.

Identify the blog and endpoint

Use the blog identifier accepted by Tumblr’s route (commonly a hostname such as staff.tumblr.com). Do not assume that every post, field, or blog is public. Read the current method documentation, note its authentication level, and request only the fields and volume your use case needs.

cURL example

This request retrieves posts through the published route. Replace the environment values with an application key and a blog you are permitted to access.

BLOG_ID='staff.tumblr.com'
TUMBLR_API_KEY='your-registered-key'

curl --fail --get 'https://api.tumblr.com/v2/blog/'"$BLOG_ID"'/posts' 
  --data-urlencode 'api_key='"$TUMBLR_API_KEY" 
  --data-urlencode 'limit=20'

The response is JSON. Save the complete response, including HTTP status and headers, so you can audit failures, rate-limit events, and the exact request parameters used. Use only pagination parameters documented for the endpoint; do not invent a second request pattern to work around a limit.

Python example

The following uses the requests package, fails loudly on HTTP errors, and writes the returned JSON to standard output for further processing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import os
import requests

blog_id = os.environ['TUMBLR_BLOG_ID']
api_key = os.environ['TUMBLR_API_KEY']
url = f'https://api.tumblr.com/v2/blog/{blog_id}/posts'

response = requests.get(
    url,
    params={'api_key': api_key, 'limit': 20},
    timeout=30,
)
response.raise_for_status()
print(response.text)

Set TUMBLR_BLOG_ID and TUMBLR_API_KEY in the process environment rather than committing credentials to source control. If the endpoint requires OAuth, this API-key-only example is not sufficient; implement the method’s documented authorization flow.

Node.js example

Node 18 or later includes fetch. This example checks the status before parsing JSON.

const blogId = process.env.TUMBLR_BLOG_ID;
const apiKey = process.env.TUMBLR_API_KEY;

if (!blogId || !apiKey) {
  throw new Error('Set TUMBLR_BLOG_ID and TUMBLR_API_KEY');
}

const query = new URLSearchParams({
  api_key: apiKey,
  limit: '20'
});
const endpoint = `https://api.tumblr.com/v2/blog/${blogId}/posts?${query}`;
const response = await fetch(endpoint);

if (!response.ok) {
  throw new Error(`Tumblr API returned ${response.status}: ${await response.text()}`);
}

const payload = await response.json();
console.log(JSON.stringify(payload, null, 2));

Official JavaScript client

Tumblr also publishes tumblr.js, an official JavaScript API client. Its documentation notes that most methods require at least an API key, while fully signed methods additionally require OAuth tokens. It is an API client, not a headless-browser scraping library. Use it when its method wrappers match your project, and still follow the authentication and rate-limit rules of the underlying API.

Build a collector that stays within the interface

Request only what you need

Choose the narrowest documented method and fields that answer your use case. Keep the blog identifier, endpoint name, authentication mode, request timestamp, HTTP status, and response metadata with each job. This makes it possible to explain where a record came from without retaining unnecessary page data.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Paginate deliberately

Use the endpoint’s documented pagination controls and stop when the response indicates there is no more data or when your defined collection window is complete. Persist a checkpoint after each successful page. A checkpoint lets you resume after a crash without replaying an entire collection and without increasing request volume unnecessarily.

Deduplicate and preserve provenance

Store Tumblr’s post identifier and blog identifier as your primary provenance fields, along with the retrieval time and the API response version or shape your parser handled. Treat edited or deleted posts as state changes instead of silently creating a second record. Keep raw responses only when your permission and retention policy allow it.

Throttle and back off

Use a bounded worker queue rather than launching one request per post or blog at once. On a 429 response, honor any server-provided retry information, pause, and retry with increasing delays. Never rotate IP addresses, keys, or accounts to evade a limit.

Published limits and the 429 response

Tumblr’s Help Center listed default limits of 1,000 API calls per hour and 5,000 calls per day per consumer key in 2022. Tumblr says limits can change and may be adjusted for an app; the API documentation also describes IP-level and feature-specific limits. Treat the figures as a planning baseline, not a permanent guarantee, and recheck the current documentation before deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Signal Likely meaning Safe response
HTTP 429, “Limit Exceeded” A request limit was reached. Stop sending new work, back off, record the event, and resume at a lower rate.
Repeated 429s after waiting The hourly, daily, IP, or feature-specific budget may still be exhausted. Inspect response headers and current documentation; request an app adjustment if appropriate.
Many small requests The design is spending calls inefficiently. Use the largest documented page size that fits your need and checkpoint results.

When the API does not provide the data you need

Do not switch to a browser merely because a web page visibly contains a field that the API does not return. The developer agreement specifically addresses that “page scraping” use case. Ask Tumblr for express written permission and describe the blogs, fields, volume, retention period, and purpose. If Tumblr offers a supported feed, Firehose arrangement, or API capability, use that instead.

Keep the permission record with your project documentation. It should identify who granted it, what access is covered, any geographic or time limits, and whether redistribution is allowed. Until those terms are clear, limit your implementation to the published interface.

Why a headless browser is the wrong workaround

A headless browser can execute JavaScript, accept cookies, and render the same HTML a visitor sees, but those technical features do not change Tumblr’s contractual rules. Automating login, copying session cookies, bypassing a challenge, hiding your identity, or spreading requests across proxies would add security and authorization risks while still leaving the scraping prohibition unresolved. For that reason, a compliant implementation should not contain a browser step for Tumblr post extraction unless Tumblr has expressly authorized it and supplied a supported procedure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual requirement is a visual capture of a page you are authorized to view—not extraction of Tumblr post data—ScreenshotNeo can return a screenshot or PDF through one GET request. It is not a substitute for Tumblr’s API and does not grant permission to collect Tumblr content. Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the complete parameter reference in the ScreenshotNeo documentation. The following calls use the public Tumblr home URL only as an example; replace it with a page you are allowed to capture.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://www.tumblr.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://www.tumblr.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.tumblr.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);

ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, click-before-capture, selector hiding, selector/delay/network-idle waits, request and resource blocking, custom headers/cookies/user agents, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, an OpenAPI specification, and compatible parameter names used by other screenshot APIs. Every plan includes every feature. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Troubleshooting

Symptom Cause to check Fix
401 or an authentication error The key is missing, invalid, or the method requires OAuth. Verify the registered key and endpoint requirement; implement OAuth for signed methods.
403 or an access-denied response The blog or content is restricted, or your requested action is not authorized. Confirm permission and scope. Do not retry through a browser or copied session.
429 Limit Exceeded An hourly, daily, IP, or feature limit was reached. Back off, reduce concurrency, checkpoint progress, and recheck current limits.
Empty or incomplete results The selected method, blog identifier, fields, or pagination state is wrong. Test one documented request, inspect the complete JSON, and follow that endpoint’s pagination and field rules.
Parser breaks after an API change The response shape changed or an undocumented field was assumed. Validate required fields, log the response version/shape, and update against current API documentation.
ScreenshotNeo returns a non-clean result The page failed, timed out, triggered a bot check, or was served from cache. Read X-Page-Verdict and X-Billed, then adjust waits, headers, blocking, or viewport settings in the documented request.

FAQ

Frequently Asked Questions

Does a robots.txt allowance by itself authorize Tumblr scraping?

No. Tumblr’s Terms of Service also prohibit scraping without express prior written permission. Treat robots.txt as one condition, not a replacement for Tumblr’s terms or written authorization.

Can ScreenshotNeo extract Tumblr post text or JSON?

No. ScreenshotNeo captures an authorized page as an image or PDF. Use Tumblr’s API for structured post data and follow the endpoint’s authentication and usage rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I do if my project needs more API volume?

Design around the documented limits first, then contact Tumblr about an app adjustment or supported arrangement. Do not evade limits with extra keys, accounts, proxies, or browser automation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.