Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use an official API, RSS/Atom feed, JSON feed, or sitemap when the publisher provides one. If you must read permitted HTML, a responsible Python scraper should check robots.txt, identify itself, make slow bounded requests, parse stable fields, validate and deduplicate records, and preserve the retrieval time and licensing context. The example below collects article links and headlines from one allowed listing page; adapt its selectors only after inspecting that publisher’s permitted markup.

Start by defining the news data you actually need

“Scrape a news website” can mean one page, a section archive, a feed of new links, or a large historical crawl. Scope the job before writing code:

  • Publisher and sections: name the specific site and paths, such as /world/ or /technology/.
  • URL boundary: decide whether article pages, author pages, tags, search results, and external links are in scope.
  • Fields: choose a schema such as canonical URL, headline, publication time, update time, byline, section, summary, article body, publisher, retrieval time, parser version, and license metadata.
  • Output: JSON and CSV are convenient for a small run; a database is safer for recurring jobs.
  • Limits: set a maximum page count, a delay, and a stop condition for repeated failures.

Begin with one permitted page and one output record. That exposes access and selector problems before they become a crawl.

Check permission, robots.txt, and reuse rights

Fetch the publisher’s current robots file before requesting pages. Python’s urllib.robotparser.RobotFileParser can answer whether a named user agent may fetch a URL; its crawl_delay(), request_rate(), and site_maps() methods expose additional directives when present. For a site at https://example-news-site.test/news, the file is https://example-news-site.test/robots.txt.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots rules manage crawler access and traffic. They do not grant copyright, privacy, database-rights, or terms-of-service permission, and they are not a way to remove a page from search results. Read the publisher’s terms, licensing notices, privacy policy, and any API agreement. Do not bypass authentication, a paywall, CAPTCHA, bot check, rate limit, or an explicit prohibition. If access or reuse rights are unclear, stop and ask the publisher or use its licensed API.

Prefer a structured source

Look for an official API first, then RSS or Atom, a JSON feed, and a sitemap. These sources are generally less fragile than article HTML and document authentication, quotas, update semantics, and reuse terms. A sitemap is useful for discovering URLs; it is not automatically permission to download every page.

Install the small Python stack

For static HTML, Requests retrieves pages and Beautiful Soup parses them:

python -m pip install requests beautifulsoup4

Use a supported Python 3 release and a virtual environment for repeatable deployments:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4

Minimal, polite scraper for a news listing

This complete example checks robots rules, identifies the client, sets a finite timeout, fails on an HTTP error, parses only article cards, normalizes text, resolves relative links, and records when the page was retrieved.

from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
from datetime import datetime, timezone
import json
import requests
from bs4 import BeautifulSoup

URL = "https://example-news-site.test/news"
UA = "ExampleResearchBot/1.0 (+contact@example.org)"
TIMEOUT = 15

parsed = urlparse(URL)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
robots = RobotFileParser(robots_url)
robots.read()
if not robots.can_fetch(UA, URL):
    raise RuntimeError(f"robots.txt does not allow {URL}")

delay = robots.crawl_delay(UA)
if delay:
    print(f"Honor crawl delay: {delay} seconds")

response = requests.get(URL, headers={"User-Agent": UA}, timeout=TIMEOUT)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
retrieved_at = datetime.now(timezone.utc).isoformat()

articles = []
for card in soup.select("article"):
    link = card.select_one("a[href]")
    headline = card.select_one("h1, h2, h3")
    if not link or not headline:
        continue
    href = urljoin(URL, link["href"])
    articles.append({
        "url": href,
        "headline": headline.get_text(" ", strip=True),
        "publisher": parsed.netloc,
        "retrieved_at": retrieved_at,
    })

# Deduplicate by URL while preserving order.
unique = {item["url"]: item for item in articles}
with open("articles.json", "w", encoding="utf-8") as output:
    json.dump(list(unique.values()), output, ensure_ascii=False, indent=2)
print(f"Saved {len(unique)} records")

The article and heading selectors are illustrative, not universal. Inspect the target site’s allowed HTML and make selectors configurable. A site may use a div card, a link-only listing, or a framework-generated attribute instead.

Extract article metadata without brittle assumptions

For each article page, prefer the canonical URL, headline, publication and update times, byline, section, summary or deck, and the article-body container. Semantic elements such as <article>, <h1>, <time>, and <nav> are useful starting points. When supplied by the publisher, JSON-LD often contains headline, datePublished, dateModified, author, and mainEntityOfPage. Treat JSON-LD as input to validate, not as a guarantee that every field is complete.

Keep selectors in a configuration object or separate file. Normalize whitespace with get_text(" ", strip=True), parse timestamps with their stated timezone, and preserve the original URL as well as a canonicalized URL. Never assume a visible date is the publication date: label it as published, updated, or unknown according to the markup.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate before saving

  • Reject or quarantine a record without a canonical URL or headline.
  • Resolve relative links with urljoin and discard links outside your approved host/path boundary.
  • Deduplicate by canonical URL, not by headline.
  • Store the retrieval timestamp separately from the publisher’s publication time.
  • Record publisher, byline, parser version, source URL, and license metadata when available.
  • Log status code, response time, page count, and skipped-record reasons.

Pagination, retries, caching, and scale

Bound pagination with a maximum page number, a maximum number of records, or a date cutoff. Stop when the next link repeats, disappears, or leaves the approved path. Request one page at a time where possible. Use a modest delay, finite timeouts, and retries only for transient failures such as a 502 or 503; exponential backoff prevents a failing server from receiving a burst of traffic. Do not retry a 401, 403, CAPTCHA page, or explicit prohibition.

Cache successful responses by URL and relevant request headers. Conditional requests using the publisher’s cache headers can reduce traffic, but honor the publisher’s instructions. For a recurring job, persist the last successful cursor or timestamp and make the run idempotent: rerunning a page should update a record rather than create a duplicate.

Requests plus Beautiful Soup are appropriate for static pages and small jobs. Scrapy adds crawl orchestration, queues, throttling, retries, and pagination for larger permitted crawls. Use browser automation only when content is genuinely rendered client-side and the publisher’s rules allow that access; a browser does not make an otherwise prohibited crawl lawful.

When JavaScript or an API changes the approach

Use the API or feed first

An official API usually gives stable field names, documented quotas, authentication, and explicit reuse terms. Prefer it over reverse-engineering a private endpoint. RSS/Atom and JSON feeds can provide headlines, dates, summaries, and links with far less HTML maintenance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recognize client-rendered pages

If the initial HTML contains no article cards but the browser displays them after scripts run, Requests will see only the shell. Confirm that the data is not already present in a public feed or documented endpoint. If browser rendering is necessary and permitted, configure a browser tool with the same low rate, bounded scope, and logging; do not use it to evade access controls.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Symptom Likely cause Fix
robots.txt denies the URL Your user agent or path is disallowed. Stop, narrow the scope, or request permission; do not rotate identities to evade the rule.
403, 429, or CAPTCHA Access control or rate limiting. Stop retries, lower the rate, use the official API/feed, or contact the publisher.
Timeout or 5xx Transient server or network failure. Use a finite timeout, bounded exponential backoff, and a run-level failure limit.
Zero records Wrong selector or JavaScript-rendered content. Inspect permitted HTML, check a feed/API, and update configurable selectors.
Broken or duplicate URLs Relative links, tracking parameters, or repeated cards. Resolve with urljoin, apply a host/path allowlist, canonicalize, and deduplicate.
Dates are inconsistent Mixed time zones or publication/update labels. Preserve the source value, parse its offset, and store published and modified fields separately.
Parser silently degrades Template changed. Validate required fields, quarantine failures, alert on a sudden count drop, and review selectors.

Or skip the browser setup

If your goal is a clean visual capture of a news page rather than structured article data, ScreenshotNeo provides a single-call website screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo documentation for all options. A direct request looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It also supports full-page and element captures, device or custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting, and an OpenAPI specification. Every feature is on every plan: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000, with yearly billing giving two months free. Sign up free for ScreenshotNeo.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Operational and legal checklist

  • Confirm the exact URL and user agent against the publisher’s current robots file.
  • Read terms, copyright/licensing, privacy, and database-rights notices.
  • Prefer documented APIs and feeds; honor quotas and attribution requirements.
  • Use timeouts, status checks, bounded pagination, caching, backoff, and low concurrency.
  • Store source, publisher, byline, publication time, retrieval time, parser version, and license metadata.
  • Recheck permissions and selectors whenever the publisher changes its template or policy.

Frequently Asked Questions

Can I scrape any news site if robots.txt allows it?

No. Robots.txt addresses crawler access and traffic; you must also evaluate terms of service, copyright or licensing, privacy, database rights, and any explicit access restriction.

Should I save the full article text?

Only when the publisher’s license and your purpose permit it. Otherwise, limit collection to fields you need, such as links and metadata, and retain the applicable licensing information.

When should I use Scrapy instead of Requests?

Use Requests and Beautiful Soup for a small static-page job. Choose Scrapy when you need queueing, throttling, retries, structured pagination, and a larger permitted crawl.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.