What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Start with an official API or licensed feed. If you have permission to collect from HTML pages, build a conservative pipeline that checks the source’s rules, extracts structured data where available, preserves the original values and timestamps, and stops if the publisher blocks automated access. Travel prices, ticket availability, and property listings change quickly, so a scrape is only useful if its provenance and freshness are clear.

Can you scrape listings, or should you use an API?

Prefer a documented API or licensed feed when one is available for your intended use. An API can define permitted fields, authentication, quotas, and update behavior; it also gives the data owner more control over third-party collection. The Office of the Privacy Commissioner of Canada notes that APIs can help data owners manage lawful collection and detect unauthorized scraping.

HTML collection is a fallback, not a way around an unavailable API or an access restriction. Before collecting, review the site’s terms, API documentation, authentication requirements, and robots.txt. Digital.gov describes robots.txt as instructions to crawlers about which site areas they should or should not access. CNIL says web scraping is not inherently prohibited under GDPR, but advises against scraping sites that object through terms, robots instructions, CAPTCHAs, or comparable technical measures. Treat a CAPTCHA, access denial, or explicit block as a stop signal; do not build a workflow to defeat it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Permission to access a page does not by itself settle whether you may reuse, republish, or commercialize its contents. Confirm the scope of any license and the rules that apply to your use and location. If collection involves personal data, GDPR obligations may apply to operations such as collection, storage, organization, and retrieval, as the European Data Protection Board explains. Minimize personal data and avoid collecting details that are not necessary for the stated purpose.

What should a listing scraper collect?

Design separate records for each vertical rather than forcing unlike listings into one loose set of fields. Keep source-specific identifiers and raw source values alongside normalized fields; this makes corrections and parser changes auditable.

Listing type Useful fields Important modeling detail
Travel and lodging Property or lodging identity, address, amenities, coordinates; accommodation or room; offer price, currency, occupancy, stay dates, and terms Schema.org distinguishes the lodging business, accommodation or room, and offer. Do not attach a room-specific offer price as if it were a permanent property-wide price.
Events Unique event URL, event name, start and end dates, venue and location, organizer, ticket URL, price, currency, availability, and sale timing Google’s event guidance calls for a unique URL and accurate name, start date, and location. Ticket price information should include service charges and fees and be updated when price or availability changes.
Real estate Listing URL and ID, property type, sale or lease status, price, currency, bedrooms, bathrooms, floor area, lot size when provided, year built, address, coordinates, broker or agent, and listing/update timestamps Record whether a field was actually supplied. Schema.org accommodation examples include bedrooms, bathrooms, floor size, year built, address, latitude, and longitude; those fields can inform a property-listing model.

Schema.org supports structured representations such as JSON-LD, Microdata, and RDFa. Prefer a page’s documented API response or structured data over brittle selectors, but validate it against what a visitor can actually see and the source’s usage terms. A machine-readable value can still be stale, incomplete, or attached to the wrong offer.

How to build a cautious collection pipeline

  1. Discover permitted sources. Check for a licensed data feed or official API first. If HTML pages are allowed, use permitted index pages or sitemaps to identify targets. Record the source, the basis for access, and any documented quota.
  2. Fetch conservatively. Use a descriptive user agent, connection and read timeouts, low concurrency, caching, and conditional requests where the source supports them. Follow documented quotas. If there is no published rate, begin slowly and back off on errors. Repeated retries against a failing or blocking site are not a recovery strategy.
  3. Extract and retain provenance. Prefer API fields and JSON-LD before page-specific selectors. Store the canonical URL, source ID, retrieval time, source update time if published, and a parser or schema version. Where lawful, keep a raw response snapshot or hash so an extraction can be reviewed later.
  4. Normalize without erasing originals. Convert dates into UTC for comparison, but retain the source timezone and source date string. Normalize currency codes and numeric values while keeping the displayed value and currency. Never silently convert away the value the publisher presented.
  5. Validate before use. Require a stable source ID or canonical URL; check that end dates follow start dates, currency codes are valid, prices are nonnegative, locations are plausible, and mandatory-fee fields are present when known. Route failed checks to review rather than publishing questionable values.
  6. Deduplicate and refresh. Match on canonical URL and source ID first. A normalized title plus location and date can help identify duplicates across sources, but retain each source’s own ID. Set refresh intervals by source and volatility; recheck event availability and lodging prices frequently enough for the intended use.
  7. Monitor and stop on signals. Track HTTP errors, empty-result rates, blocks, schema changes, and field-level drift. Stop collection when the publisher prohibits it or imposes a technical block. Alert on stale prices or sold-out events before downstream pages or decisions rely on them.

Example: extract JSON-LD from a permitted listing page

The following Python example fetches one page and prints JSON-LD objects found in its HTML. It is a starting point for a source that permits automated requests, not a way to bypass a login, CAPTCHA, or block. Install dependencies with python -m pip install requests beautifulsoup4, replace the example URL and user agent with values appropriate to your permitted source, then run the script.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import requests
from bs4 import BeautifulSoup

URL = "https://example.com/permitted-listing"
HEADERS = {"User-Agent": "ExampleListingResearch/1.0 (contact: data@example.com)"}

response = requests.get(URL, headers=HEADERS, timeout=(5, 20))
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
found = 0
for script in soup.select('script[type="application/ld+json"]'):
    raw = script.string or script.get_text()
    try:
        data = json.loads(raw)
    except json.JSONDecodeError:
        continue
    print(json.dumps(data, ensure_ascii=False, indent=2))
    found += 1

if found == 0:
    print("No parseable JSON-LD found; do not assume the page has no listing data.")

This deliberately prints source data instead of pretending that every publisher uses the same schema. A production parser should inspect each source’s documented structure, handle JSON-LD arrays and @graph objects, map only known fields, and preserve the source URL and timestamps. If the page returns an error, requires a challenge, or has no structured block, do not respond by bypassing controls; verify permission and use an approved data source or a permitted extraction method.

How to keep prices and availability accurate

Prices are not interchangeable just because they share a currency. For lodging, a nightly rate may depend on dates, occupancy, room type, taxes, and terms. For tickets, a displayed base price may differ from the amount a buyer must pay. Store the date range, occupancy or ticket tier, currency, fee treatment, and availability state that accompanied each value.

In the United States, the Federal Trade Commission’s Unfair or Deceptive Fees Rule took effect on May 12, 2025, for covered businesses offering, displaying, or advertising live-event tickets and short-term lodging. Covered advertised prices must show the total mandatory price upfront; optional charges, taxes, government charges, and shipping have separate treatment under the rule. Do not label a base ticket price or nightly rate as the final total when mandatory fees are known. Verify current applicability for the business and transaction rather than assuming every listing or jurisdiction is covered identically.

Make freshness visible in your own system. Store retrieved_at for every record and source_updated_at when the publisher supplies it. A record with no source update time is not evidence that the offer has not changed; it only means that the source did not provide that timestamp.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to handle errors, blocks, and changing pages

  • HTTP 403, CAPTCHA, or access-denied page: Stop automated requests to that source. Check whether access is prohibited or requires an approved API, and contact the publisher if clarification is needed. Do not rotate identities or attempt to defeat the challenge.
  • HTTP 429 or quota response: Reduce request volume and honor any retry or quota instructions in the API documentation. Apply exponential backoff, cache prior results, and resume only within permitted limits.
  • Timeout or server error: Use a bounded retry policy with backoff for transient failures. If errors persist, pause the source and raise an alert; do not let retries create a request surge.
  • Empty or malformed extraction: Check whether the source changed its markup or structured-data format. Preserve the failed response or hash where lawful, mark the record as unverified, and update the parser only after confirming the new representation.
  • Price or date fails validation: Keep the original text and source context, quarantine the record, and check timezone, decimal, currency, and fee assumptions. Do not silently coerce an implausible value into a valid-looking one.
  • Duplicate listing: Compare canonical URLs and stable IDs before fuzzy title matching. Preserve provenance for each source even when your application merges equivalent records.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What changes performance, reliability, and operating cost?

The largest practical levers are request volume, page weight, source latency, permitted quotas, and how often each field needs refreshing. Fetch only the pages needed, use caching and conditional requests where available, and avoid parallelism until a source’s rules and behavior are understood. An API may have explicit quota or licensing costs; HTML access can still carry substantial engineering and maintenance costs because markup and anti-automation controls change.

Freshness should follow the data, not a single global timer. Ticket availability and accommodation offers are volatile; a property’s descriptive fields may change less often than its price or listing status. Separate refresh schedules by field and source if needed, and prevent stale values from appearing current just because a page was fetched recently. A recent retrieval timestamp says when you checked, not when the publisher last verified the underlying offer.

Or skip the browser setup

If you only need a visual capture of a permitted listing page, ScreenshotNeo is a screenshot API and MCP server, not a structured listing feed or replacement for checking source permission. It can return a PNG, JPEG, WebP, or PDF; its clean-shot steps can accept cookie/consent banners and remove known consent platforms, newsletter popups, and chat widgets before capture. Those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.

One GET request example (see the ScreenshotNeo API documentation):

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/permitted-listing -o shot.webp

The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan.

Questions developers still ask

Can a screenshot tell me the ticket price or property details reliably?

A screenshot preserves a visual state, not a normalized data record. For machine processing, prefer a permitted API or structured page data and validate extracted values; use a capture as visual context or for review, not as proof that a price or availability field is current.

Should I collect personal details shown beside a listing?

Only if they are necessary, permitted, and covered by an appropriate lawful basis and privacy process. For a listing workflow, avoid collecting contact or identity details that are not required; GDPR applicability depends on whether personal-data processing is involved and the circumstances of that processing.

Can I treat a missing fee as zero?

No. A missing fee field means the fee status is unknown unless the source explicitly says otherwise. Preserve that distinction so a downstream display does not turn incomplete source data into a claim that the total price is fee-free.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does scraping a publicly visible listing automatically allow me to republish it?

No. Public visibility alone does not establish permission to reuse or republish listing content. Check the source’s terms and license and the rules that apply to your intended use.

How should I represent an offer that has sold out or expired?

Keep the source’s status and the time it was observed, and distinguish unavailable, expired, and unknown rather than deleting or relabeling the record without provenance.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.