Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape articles responsibly, first look for an official API, RSS feed, sitemap, or permission process. Then define a small scope, check the site’s terms and robots.txt, fetch pages slowly, parse the returned HTML, validate the fields you extracted, and keep collection separate from publishing or redistribution. Publicly visible text is not automatically free to copy or republish.

This guide shows a complete Python workflow for known article URLs, explains when to use Scrapy for a bounded crawl, and covers JavaScript-rendered pages, failures, privacy, copyright, and operational safeguards.

1. Define exactly what you need

A scraper is easier to control when its scope is written down before the first request. Record:

  • The target host and permitted URL pattern, such as https://example.com/news/.
  • The fields required: title, author, publication date, article body, canonical URL, and perhaps section or tags.
  • The number of pages and how discovery stops. A list of 20 supplied URLs is safer than an unrestricted site crawl.
  • Your purpose, storage location, retention period, and who will receive the output.
  • Whether you need article text at all, or only metadata, links, and factual values.

“Scraping” usually means collecting data from a page; “crawling” describes following links to discover more pages. A project can scrape a few known URLs without crawling, or combine bounded crawling with extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

2. Check for an authorized data source first

Before writing selectors, look for a documented API, RSS or Atom feed, sitemap, downloadable dataset, licensing page, or a publisher contact for research access. The Carpentries web-scraping lesson recommends checking whether structured access exists and asking the organization when appropriate.

An API or feed is generally more stable than parsing presentation HTML. It may also specify fields, limits, authentication, and permitted uses. If an official route supplies the same information, prefer it over automated page collection.

3. Review terms, privacy rules, and robots.txt

Terms and intended use

Read the target site’s terms of service and privacy policy. They may restrict automated collection, bulk downloading, commercial use, or redistribution. Collection, analysis, storage, and republication are separate actions: permission to view a page does not necessarily authorize all four.

Reuters Connect’s terms updated September 2024, for example, prohibit scraping and automated collection of platform content without prior written consent and require compliance with exclusionary protocols. Treat this as an example of why each target’s live terms must be checked, not as a universal rule for every site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the correct robots.txt

Request the root-level file for the same host, protocol, and port, such as https://example.com/robots.txt. Google explains in its robots.txt specification that a file applies only to that origin; a file on blog.example.com does not automatically govern www.example.com.

Match the rules to your user-agent and the paths you plan to request. Robots.txt is an exclusion signal for crawlers, not a blanket license or a complete statement of copyright and contract rights. Review terms and obtain authorization independently. If the site signals that automated requests are unwanted, stop and use an authorized route.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

4. Install a small, auditable Python scraper

Prerequisites

Use Python 3, then install the HTTP client and parser:

python -m pip install requests beautifulsoup4 lxml

The following script handles a supplied list of article URLs, identifies common metadata, extracts paragraph text, uses a descriptive user-agent, applies a delay, and writes JSON. It does not discover links, bypass access controls, or retry indefinitely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import time
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

URLS = [
    "https://example.com/article-one",
    "https://example.com/article-two",
]

HEADERS = {
    "User-Agent": "ResearchArticleCollector/1.0 (contact: data@example.org)",
    "Accept": "text/html,application/xhtml+xml",
}


def meta(soup, *names):
    for name in names:
        tag = soup.find("meta", attrs={"name": name}) or soup.find(
            "meta", attrs={"property": name}
        )
        if tag and tag.get("content"):
            return tag["content"].strip()
    return None


def extract_article(url, session):
    response = session.get(url, headers=HEADERS, timeout=30)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "lxml")

    title_tag = soup.find("h1") or soup.find("title")
    title = title_tag.get_text(" ", strip=True) if title_tag else None

    body = soup.select_one("article")
    if body is None:
        body = soup.select_one("main")
    paragraphs = body.select("p") if body else soup.find_all("p")
    text = "nn".join(
        p.get_text(" ", strip=True) for p in paragraphs
        if p.get_text(" ", strip=True)
    )

    return {
        "url": response.url,
        "host": urlparse(response.url).netloc,
        "title": title,
        "author": meta(soup, "author", "article:author"),
        "published": meta(
            soup, "article:published_time", "date", "pubdate"
        ),
        "text": text,
        "retrieved_at": datetime.now(timezone.utc).isoformat(),
    }


def main():
    records = []
    with requests.Session() as session:
        for index, url in enumerate(URLS):
            try:
                records.append(extract_article(url, session))
            except requests.RequestException as exc:
                records.append({"url": url, "error": str(exc)})
            if index < len(URLS) - 1:
                time.sleep(2)  # choose a rate the site can reasonably tolerate

    with open("articles.json", "w", encoding="utf-8") as output:
        json.dump(records, output, ensure_ascii=False, indent=2)


if __name__ == "__main__":
    main()

Replace the sample URLs, user-agent contact address, and selectors with values appropriate to the authorized site. Do not assume that every page uses an article element. Inspect a few returned documents and adjust selectors deliberately.

What the parser is doing

  • raise_for_status() turns HTTP 4xx and 5xx responses into explicit errors instead of silently saving an error page.
  • Metadata is read from common meta names, while the title falls back from h1 to title.
  • The body selector prefers article, then main, and finally all paragraphs. That fallback is convenient but can include navigation or related links, so validate it.
  • The final record preserves the resolved URL and retrieval timestamp, which helps audit redirects and changes.

5. Validate before collecting more

Run the script against a small sample and compare each record with the rendered page. Check that:

  • The title is the article title, not the site name.
  • The author and date come from the article, not a sidebar or modified-date widget.
  • The text excludes navigation, comments, cookie notices, newsletter forms, and “related stories.”
  • Paragraph order, Unicode characters, links, and headings are retained as your downstream task requires.
  • Pages with templates, paywalls, language variants, or corrections are represented accurately.

Layouts change. Keep a fixture of a few permitted pages and rerun extraction tests after selector changes. Log URL, status code, elapsed time, and parser warnings, but avoid logging personal data unnecessarily.

6. When a bounded crawl is justified: Scrapy

For many authorized article URLs, Scrapy provides scheduling, duplicate filtering, item pipelines, and throttling. Its downloader middleware documentation states that RobotsTxtMiddleware filters requests forbidden by the robots.txt exclusion standard.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Enable compliance in the project settings:

ROBOTSTXT_OBEY = True
USER_AGENT = "ResearchArticleCollector/1.0 (contact: data@example.org)"
DOWNLOAD_DELAY = 2
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1

Write an allowlisted spider that starts from known article URLs or a documented section, accepts only the target host and path, and stops after the required depth or item count. Do not turn a content extraction task into an unrestricted site crawl. Keep the crawl bounded even when robots.txt permits more paths.

7. Handle pages whose HTML lacks the article

Confirm the problem

Save the raw response and search it for a distinctive headline or paragraph. If the text is absent, the page may render it with JavaScript, require a session, show a consent gate, or return a bot-check page. Do not immediately add browser automation or attempt to defeat a challenge.

Preferred order of solutions

  1. Use the publisher’s API, feed, sitemap, export, or permissioned endpoint.
  2. Ask the publisher for authorized access or a research data extract.
  3. If browser rendering is explicitly allowed, use a normal browser session with conservative limits and no challenge bypass.
  4. Record that the content was unavailable rather than fabricating an incomplete article.

Browser automation adds JavaScript execution, cookies, memory use, and more failure modes. It is not automatically required just because a page is dynamic.

8. “Or skip the browser setup”: ScreenshotNeo

When your authorized workflow needs a rendered visual record of a page, ScreenshotNeo can return a PNG, JPEG, WebP, or PDF from one request. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is not a replacement for permission to collect or republish article text. Use it when a screenshot or rendered-page artifact is an authorized part of your project, and still review the target’s rules.

See the ScreenshotNeo documentation for all options. A minimal call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python:

import requests
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For article capture, relevant options include full-page screenshots with lazy images loaded, a CSS-selector element capture, custom JavaScript or CSS, waiting for a selector, delay, or network idle, custom cookies and headers, timezone and geolocation, ad or tracker blocking, image resizing, a chosen viewport or device preset, dark mode, and PDF page ranges. Async jobs with signed webhooks, bulk capture of up to 100 URLs per call, caching with a chosen TTL, signed public-image links, and a usage API are available when your authorized workload needs them. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

The Free plan includes 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to try the 1,000 monthly shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Protect the site and people represented in the data

  • Identify your crawler where appropriate and provide a monitored contact address.
  • Use a modest per-domain rate, delays, connection timeouts, and a small concurrency. Collect during off-peak hours when practical; the GSA guidance published July 7, 2021 emphasizes transparency and minimizing impact.
  • Cache responses and avoid requesting the same URL repeatedly. Honor clear stop signals, server errors, and rate-limit responses.
  • Do not bypass logins, paywalls, CAPTCHAs, access controls, or technical restrictions.
  • Minimize personal data. Define retention, access controls, deletion procedures, and whether downstream users may receive raw text.
  • Keep provenance: source URL, retrieval time, parser version, and any transformations.

10. Separate extraction from reuse

Copyright, privacy law, contract terms, database rights, access restrictions, and jurisdiction can all affect what you may do with collected material. Facts and short metadata may present different issues from storing expressive article prose. Republishing an entire article, training a product on it, or selling a dataset deserves a separate permission analysis.

The University of Michigan’s copyright guide discusses these distinctions. For substantial academic or commercial work, consult qualified legal or institutional counsel rather than assuming that “public” means unrestricted. The Carpentries summarizes the practical precaution this way: “To avoid legal or ethical issues, it’s essential to check both the TOS and the site’s robots.txt file before scraping.”

11. Troubleshooting

403 or 429 responses

Cause: the server rejects the request, detects automation, or rate-limits you. Fix: stop, read the site’s rules, reduce request frequency, identify yourself, and seek an API or permissioned route. Do not rotate identities to evade a block.

200 response but no article text

Cause: JavaScript rendering, a consent wall, login requirement, or a bot-check document. Fix: inspect the saved HTML, verify an official route, and use authorized rendering only if the publisher permits it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Selector returns menus and footers

Cause: the selector is too broad or the template changed. Fix: inspect the DOM, target a stable article container, exclude known classes, and validate against multiple pages.

Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Dates or authors are missing

Cause: values may be in JSON-LD, visible labels, or a different meta property. Fix: inspect the page source and add a documented fallback; preserve null when the field is not present rather than guessing.

Timeouts and incomplete records

Cause: slow servers, oversized media, transient failures, or network idle that never occurs. Fix: use finite connect/read timeouts, bounded retries for transient errors, smaller scopes, and a queue of failed URLs for manual review. Never retry a forbidden request indefinitely.

12. A practical decision checklist

  1. Is there an API, feed, sitemap, export, or permission process?
  2. Have you read the current terms and privacy policy?
  3. Did you check the correct host’s robots.txt for your user-agent and paths?
  4. Is the URL set, field list, rate, and retention period bounded?
  5. Did a small sample produce correct titles, dates, authors, and body text?
  6. Are JavaScript rendering, cookies, or screenshots actually authorized and necessary?
  7. Can you explain how the collected data will be stored, protected, and reused?
  8. Do your logs and error handling let you stop quickly when the site objects?

Frequently Asked Questions

Does robots.txt make scraping legal?

No. It is a host- and protocol-specific crawler instruction. Terms, authorization, copyright, privacy, access controls, and local law must be assessed separately.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I scrape the rendered page or the raw HTML?

Use raw HTML when it contains the needed fields. If content is absent, check an official structured source first; use authorized browser rendering only when necessary and permitted.

How can I avoid collecting more personal data than needed?

Limit fields and URLs, redact or omit identifiers, restrict access, set a retention period, and avoid redistributing raw article text unless your permission and legal analysis cover it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.