Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract URLs in two stages: first find URL-like text with a regular expression or a document parser, then trim surrounding punctuation, parse each candidate, and enforce your application’s scheme and host rules. A regex can locate likely links, but it cannot establish that a match is safe or usable.

Choose the right method for your input

For plain text, use a regular expression as a candidate finder and a URL parser for validation. If the source is HTML or Markdown you control, use its parser to read actual link nodes instead of searching all text: this avoids treating unrelated prose, code, or markup as links.

  • Plain text, logs, messages: use the examples below, then review trimming and validation rules.
  • HTML: parse the document and extract anchor destinations (and, if needed, other attributes such as image sources). A regex over markup can match text that is not a link.
  • Markdown: use a Markdown parser when you need actual link destinations, including reference-style links. A text regex may also find bare URLs, but it does not interpret Markdown structure.

The examples find absolute URLs beginning with http://, https://, or ftp://. Extend that policy only if your input and use case need other forms.

Extract and validate URLs in Python

This Python example finds candidates, removes common sentence punctuation, parses the result with urllib.parse.urlsplit(), accepts only the chosen schemes with a network location, removes fragments, and returns the resulting URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import re
from urllib.parse import urlsplit, urldefrag

candidate_re = re.compile(r'(?i)b(?:https?|ftp)://[^s<>"']+')

def extract_urls(text):
    found = []
    for raw in candidate_re.findall(text):
        cleaned = raw.rstrip('.,;:!?)]}')
        parts = urlsplit(cleaned)
        if parts.scheme in {'http', 'https', 'ftp'} and parts.netloc:
            url, _fragment = urldefrag(cleaned)
            found.append(url)
    return found

sample = 'Read https://example.com/docs, then visit (https://example.org/page#intro).'
print(extract_urls(sample))

urlsplit() separates scheme, authority, path, query, and fragment; urldefrag() removes the fragment. Python documents these APIs in urllib.parse. The returned values omit fragments by design. If fragments matter to your application, return cleaned instead of the URL from urldefrag().

Important limitation of the trimming line

rstrip('.,;:!?)]}') is a practical shortcut for common prose, not a universal punctuation algorithm. It strips closing brackets even when they may be part of a URL. A path can legitimately contain balanced parentheses, so blindly deleting every trailing ) can damage a valid link. If your input includes such URLs, track opening and closing delimiters or use source-specific parsing rules before trimming.

Extract and validate URLs in JavaScript

In JavaScript, use a regex to find candidates and the built-in URL constructor to parse them. This function returns normalized absolute URL strings for the allowed protocols.

function extractUrls(text, baseUrl) {
  const rough = text.match(/b(?:https?|ftp)://[^s<>"']+/gi) ?? [];
  return rough.flatMap(raw => {
    const cleaned = raw.replace(/[.,;:!?)]}+$/, "");
    try {
      const parsed = new URL(cleaned, baseUrl);
      if (!["http:", "https:", "ftp:"].includes(parsed.protocol)) return [];
      return [parsed.href];
    } catch {
      return [];
    }
  });
}

console.log(extractUrls("See https://example.com/docs, then https://example.org/a."));

The constructor returns a parsed URL or throws when it cannot parse the input. For a candidate-only check, MDN also documents URL.canParse(). See MDN: URL. As in the Python example, the closing-parenthesis trim is deliberately simple and should be adjusted if balanced parentheses are meaningful in your URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle punctuation, wrappers, and relative references

Trailing punctuation and line breaks

In ordinary writing, commas and periods often follow a URL without whitespace. Strip only delimiters that are outside the link. RFC 3986 notes that punctuation marks can be mistaken for URI content, and that line wrapping can make boundaries unclear. A URL split across lines may therefore need preprocessing based on the format that produced the text; do not remove whitespace indiscriminately if it could alter the intended value.

Quoted and bracketed URLs

Inputs may wrap links in quotation marks, angle brackets, or parentheses, or use a legacy URL: prefix. Handle these wrappers only when they are clearly outside the candidate. Python’s URL parsing documentation describes these wrapped forms and parsing behavior: urllib.parse. Do not strip a character merely because it is a delimiter in prose; it could be part of a legitimate path or query.

Relative references and protocol-relative forms

A value such as /docs/page is a relative reference, not a complete network URL. Keep it relative unless you have a trusted, relevant base URL. In Python, use urllib.parse.urljoin(base_url, reference); in JavaScript, use new URL(reference, baseUrl). RFC 3986 defines reference resolution; see Section 5. A reference beginning // inherits its scheme from a base URL. The regexes above do not locate relative or protocol-relative references; add them only if your input requires them and your code has a trusted base.

Fragments, internationalized domains, and percent encoding

Decide whether fragments identify distinct results in your application. The Python function drops them, while the JavaScript example preserves them. URL parsers also apply their own normalization rules, which can affect how a parsed value is represented. Preserve the original text separately if display fidelity matters. Do not lowercase paths or decode percent escapes as a blanket cleanup step: reserved characters and encoding can have meaning specific to a URL’s scheme or application. RFC 3986 covers URI syntax and percent encoding at RFC 3986.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apply a security policy before using extracted values

A syntactically parseable URL is not automatically safe to fetch, open, or trust. Before acting on extracted values, define an explicit policy:

  • Allow only the schemes the task requires, commonly https and sometimes http. Reject unexpected schemes such as javascript: when results may be navigated to or fetched.
  • For network URLs, require a nonempty hostname and validate acceptable ports and hostnames for your application.
  • Treat user information such as embedded usernames or passwords as sensitive; reject it if your use case does not allow credentials in URLs.
  • If your application fetches user-supplied URLs, consider where a hostname resolves and whether the destination is permitted by your network policy. Parsing alone is not a fetch-safety guarantee.
  • Keep the original string for display or auditing, and use a parsed representation for comparison and policy checks.

RFC 3986 discusses security considerations for URI construction and use. The rfc3986 library documentation describes validation options, including requiring schemes or hosts and forbidding passwords in user information. Use validation that matches your application; no single regex or parser policy fits every security boundary.

Regex or parser: what each part does

Approach Best use Limitation
Candidate-matching regex Finding likely absolute URLs in plain text Cannot reliably decide whether punctuation belongs to a URL or whether a match is valid and safe
URL parser Separating components and applying scheme, host, and other policy checks Does not identify every URL buried in arbitrary prose by itself
HTML or Markdown parser Extracting links from structured documents Must match the source format and the kinds of link-bearing elements you need

RFC 3986’s Appendix B provides a reference regular expression for decomposing a URI reference into components; that is different from finding every URL in arbitrary text. The standard describes parsing a URI reference into its major components. See Appendix B.

Troubleshoot common extraction problems

A period or comma appears at the end

The match includes sentence punctuation because there is no whitespace boundary. Trim likely delimiters after matching, but make the trimming logic sensitive to balanced brackets and the formats in your data.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A valid URL is missing

Check whether it uses a scheme your pattern does not include, is split across lines, or is a relative reference. Expand the candidate rules only for formats you need, then parse and validate the added forms separately.

A match is not a working link

Regex matching only found URL-like text. Confirm that parsing succeeds, that the scheme and host meet your policy, and that the destination itself is reachable if your application needs to fetch it.

A relative path is rejected

That is expected if the code is looking for absolute network URLs. Retain relative references as references, or resolve them against a trusted base URL; do not invent a base.

A URL changes after parsing

The parser may serialize or normalize the value. Keep the original substring if exact text is important, and avoid ad hoc case changes or percent-decoding. Compare normalized values only under rules appropriate to your application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If what you need is a screenshot of a page—not extracting links from text—ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. For example, save a WebP screenshot of a page with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses report the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

Further reading

The primary references for URL syntax and the APIs used here are RFC 3986, Python’s urllib.parse documentation, and MDN’s URL documentation.

Frequently Asked Questions

Should fragments be included in extracted URLs?

It depends on what the result represents: keep fragments when the destination within a page matters, or remove them when you want page-level URLs. The Python example removes fragments; the JavaScript example preserves them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can a regular expression validate every URL?

No. A regex can locate likely candidates, but parsing and explicit application rules are needed to check components and decide whether a URL is acceptable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.