Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To collect job postings with Python, first confirm that the source permits your intended access. Prefer an official API or partner integration where one is available. For an allowed, server-rendered page, fetch its HTML with Requests and parse it with Beautiful Soup or lxml; use Scrapy for a larger crawl, or Playwright or Selenium only when the site permits browser automation. The example below shows the pattern for a page you are authorized to collect from—it is not a universal Indeed or LinkedIn scraper.

Choose a permitted source and access method

Before writing a scraper, check the source’s terms, applicable API rules, and robot-exclusion directives. Permission to view a page in a browser is not, by itself, permission to collect, store, or reuse its contents. Keep the collection within the permitted scope, and stop if the site blocks requests or changes its access rules.

If a source offers an API or partner program that covers your use case, investigate that first. Indeed’s developer documentation covers APIs for jobs, candidates, employers, and search integrations; its Job Sync API is a GraphQL API for ATS partners to create, update, expire, and check the status of job postings. Those descriptions do not mean every developer can freely query all Indeed listings. Indeed’s Developer Agreement restricts uses including copying, redistribution, permanent database creation, algorithmic query generation, and attempts to bypass access limits.

LinkedIn documents an approval and vetting process for Job Posting API integrations in its Job Posting API terms. Its crawling terms say automated crawling and indexing without express permission is prohibited; permitted crawling must follow authorized paths and robot-exclusion restrictions. LinkedIn’s prohibited software guidance also says third-party software, crawlers, bots, browser plug-ins, and scripts that scrape or automate activity are not permitted on its services. Do not treat browser automation as a workaround.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Match the tool to the page and scale

Approach Use it when Trade-off
Official API or partner integration The provider documents an API that covers your purpose and grants access. Access, fields, and usage depend on the provider’s terms and approval process.
Requests plus Beautiful Soup or lxml The page is server-rendered and HTML collection is permitted. Simple and light, but selectors can break when markup changes.
Scrapy You have permission to visit many pages and need a crawl queue, retries, and item pipelines. More setup than a one-page script; it does not grant access that the source disallows.
Playwright or Selenium The permitted page requires JavaScript to render the listing data. Uses a browser and more resources; automation must still be allowed by the site.

Beautiful Soup, Scrapy, Selenium, and Requests are among the Python scraping tools discussed in Web Scraping with Python. Tool choice does not change the permission question: use only the access method and pages the source allows.

Build a small, polite Requests scraper

For a permitted server-rendered listing page, start with one request and inspect the HTML before expanding to pagination. The example below expects each posting to use an article[data-job] element and the listed data-* attributes. Those selectors are an example contract, not selectors for Indeed, LinkedIn, or any other named job board; adapt them to the permitted source’s actual markup. Install the dependencies with python -m pip install requests beautifulsoup4.

import csv
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urldefrag

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

START_URL = "https://example.com/jobs"
OUTPUT_CSV = "jobs.csv"
MAX_PAGES = 5  # Keep within the source's permitted scope.

session = requests.Session()
session.headers.update({
    "User-Agent": "JobListingResearch/1.0 (contact: you@example.com)"
})
retry = Retry(
    total=3,
    backoff_factor=1,
    status_forcelist=(429, 500, 502, 503, 504),
    allowed_methods=frozenset(["GET"]),
    respect_retry_after_header=True,
)
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))


def text_or_empty(node):
    return node.get_text(" ", strip=True) if node else ""


def absolute_url(href, page_url):
    if not href:
        return ""
    return urldefrag(urljoin(page_url, href))[0]


def parse_page(html, page_url, retrieved_at):
    soup = BeautifulSoup(html, "html.parser")
    jobs = []
    for card in soup.select("article[data-job]"):
        link = card.select_one("a[data-job-url]")
        jobs.append({
            "title": text_or_empty(card.select_one("[data-job-title]")),
            "employer": text_or_empty(card.select_one("[data-employer]")),
            "location": text_or_empty(card.select_one("[data-location]")),
            "description": text_or_empty(card.select_one("[data-description]")),
            "employment_type": text_or_empty(card.select_one("[data-employment-type]")),
            "salary": text_or_empty(card.select_one("[data-salary]")),
            "posting_url": absolute_url(link.get("href") if link else "", page_url),
            "source_url": page_url,
            "retrieved_at_utc": retrieved_at,
        })
    next_link = soup.select_one("a[rel='next']")
    next_url = absolute_url(next_link.get("href") if next_link else "", page_url)
    return jobs, next_url


records = []
seen_urls = set()
page_url = START_URL

for page_number in range(MAX_PAGES):
    if not page_url:
        break
    response = session.get(page_url, timeout=(10, 30))
    response.raise_for_status()
    retrieved_at = datetime.now(timezone.utc).isoformat()
    page_jobs, next_url = parse_page(response.text, response.url, retrieved_at)

    for job in page_jobs:
        key = job["posting_url"] or (job["title"], job["employer"], job["location"])
        if key not in seen_urls:
            seen_urls.add(key)
            records.append(job)

    page_url = next_url
    if page_url:
        time.sleep(2)  # Set a conservative rate consistent with the source's rules.

fields = [
    "title", "employer", "location", "description", "employment_type",
    "salary", "posting_url", "source_url", "retrieved_at_utc",
]
with open(OUTPUT_CSV, "w", newline="", encoding="utf-8-sig") as csvfile:
    writer = csv.DictWriter(csvfile, fieldnames=fields)
    writer.writeheader()
    writer.writerows(records)

print(f"Saved {len(records)} postings to {OUTPUT_CSV}")

Change START_URL to a page you may collect and update the selectors after inspecting its HTML. A successful HTTP response with zero records usually means the page uses different markup, the content is JavaScript-rendered, or the page no longer contains listings; do not respond by guessing selectors or increasing crawl volume blindly. The script follows only a visible rel="next" link and caps the number of pages. Adjust both behavior and request interval to the source’s permitted scope.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

What the records preserve

The CSV separates posting data from collection metadata. Keep a canonical posting URL or stable source ID when available so later runs can deduplicate records. The script uses a URL as its preferred key and falls back to title, employer, and location only when no URL is present; that fallback is less reliable because two distinct jobs can share those fields.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Salary is collected only when the page shows it. Store the displayed value faithfully, including currency and pay period if present; do not infer compensation from a title, location, or similar listing. Likewise, an empty field means the source did not provide a value in the parsed markup, not that the value is zero. The UTC retrieval timestamp records when your script saw the page; it is not the employer’s publication or update time.

When the page needs JavaScript

If the permitted page’s listing cards do not appear in the HTML returned by Requests, first check for a documented API or a server-rendered version that the site authorizes. If browser automation is expressly allowed, Playwright can render the page before parsing its HTML. Install it with python -m pip install playwright and python -m playwright install chromium.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

url = "https://example.com/jobs"  # Use only a permitted page.
with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto(url, wait_until="domcontentloaded", timeout=30000)
    page.locator("article[data-job]").first.wait_for(timeout=10000)
    html = page.content()
    browser.close()

jobs, next_url = parse_page(html, url, "")
print(f"Rendered {len(jobs)} job cards; next page: {next_url or 'none'}")

This uses the same example selectors as the Requests script and calls its parse_page function, so keep that function in the same Python file or import it from your scraper module. A selector wait is preferable to an arbitrary long sleep, but it is not a way around access controls. Do not automate a source whose rules prohibit it, defeat bot checks or CAPTCHAs, or continue after a block.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not a job-data extraction API: a screenshot does not give you structured job fields or permission to collect postings. It can help capture an authorized listing page for visual review. Its API accepts a URL and returns an image or PDF; its clean-shot steps accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets, with each step able to be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a single visual capture, the Python request pattern is:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/jobs"},
    timeout=90,
)
r.raise_for_status()
open("job-listing.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for request options and response details. Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make collection useful without overstating the data

Normalize carefully

Preserve the original title, employer, location, and salary text alongside any normalized fields you add. Normalize whitespace and standardize formats only when the source supports the transformation. Locations may be ambiguous, and salary displays can use different currencies, ranges, and time units; keep the raw value so a later analyst can audit normalization. Do not turn missing values into estimates.

Choose storage for the job

CSV is convenient for a small export or spreadsheet workflow. For repeated collection, SQLite or a data warehouse makes it easier to track source IDs, retrieval history, and changes without overwriting earlier observations. Keep raw response metadata or an authorized archival representation when appropriate, but follow the source’s retention and redistribution terms. A posting’s presence in your dataset should not be mistaken for evidence that it remains open.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control freshness and duplicates

Deduplicate with the source’s stable posting ID or canonical URL where available. Listings can be reposted, updated, or removed, so record retrieval time separately from any publication or update time displayed on the page. If you need current openings, revisit only at a frequency allowed by the source and mark records that are no longer present instead of silently treating old rows as active.

Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Troubleshoot common failures

  • HTTP 403 or a bot check: The site is denying the request. Recheck permission and approved access routes, then stop if access is not authorized. Do not rotate identities, evade a CAPTCHA, or use browser automation to defeat the restriction.
  • HTTP 429 or repeated rate limits: Reduce request frequency, honor any Retry-After response, and narrow the collection. If the source offers an API or partner route, use its documented process instead.
  • HTTP 404: The URL may be stale or the listing removed. Check the source’s current navigation or API documentation; do not generate arbitrary URL patterns to probe for records.
  • 200 response but no cards: Inspect the returned HTML and confirm the page is the expected listing, then update selectors for actual markup. If data is rendered only in a browser, use an authorized API or an explicitly permitted browser workflow.
  • Timeout or connection error: Keep finite connect/read timeouts, retry only transient failures with backoff, and reduce concurrency. Repeated retries against a struggling or blocking site can make the problem worse.
  • Duplicate rows: Prefer stable IDs or canonical posting URLs. If only a composite of title, employer, and location is available, treat it as an imperfect deduplication key.
  • Broken CSV columns or garbled text: Use a CSV writer rather than concatenating strings, open the file as UTF-8, and preserve multiline descriptions as one quoted field. The sample uses utf-8-sig for spreadsheet compatibility.

Keep the scraper maintainable

Run a small permitted sample before expanding the crawl. Log the requested URL, response status, retrieval time, record count, and parsing errors; monitor changes in empty fields and duplicates as well as exceptions. If the page structure changes or the source signals that access is no longer allowed, pause collection and reassess instead of silently emitting incomplete records. Scrapy can help organize queues, retries, and item pipelines for an authorized multi-page project, but it cannot make prohibited collection permissible.

For a practical introduction to common Python scraping techniques and libraries, the linked Web Scraping with Python resource covers Requests, Beautiful Soup, Scrapy, and Selenium.

Frequently Asked Questions

Can I infer a missing salary from the job title or location?

No. Keep the salary field empty or mark it as not provided; an estimate based on title or location is not data reported by the employer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does a successful request mean a posting is still open?

No. It only shows that the page was retrievable when you collected it. Check the source again and keep retrieval time distinct from any posting status or date.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.