Extract URLs in two stages: first find URL-like text with a regular expression or a document parser, then trim surrounding punctuation, parse each candidate, and enforce your application’s scheme and host rules. A regex can locate likely links, but it cannot establish that a match is safe or usable.
Choose the right method for your input
For plain text, use a regular expression as a candidate finder and a URL parser for validation. If the source is HTML or Markdown you control, use its parser to read actual link nodes instead of searching all text: this avoids treating unrelated prose, code, or markup as links.
- Plain text, logs, messages: use the examples below, then review trimming and validation rules.
- HTML: parse the document and extract anchor destinations (and, if needed, other attributes such as image sources). A regex over markup can match text that is not a link.
- Markdown: use a Markdown parser when you need actual link destinations, including reference-style links. A text regex may also find bare URLs, but it does not interpret Markdown structure.
The examples find absolute URLs beginning with http://, https://, or ftp://. Extend that policy only if your input and use case need other forms.
Extract and validate URLs in Python
This Python example finds candidates, removes common sentence punctuation, parses the result with urllib.parse.urlsplit(), accepts only the chosen schemes with a network location, removes fragments, and returns the resulting URLs.
Recommended Free Tools
#1 Best Overall
import re
from urllib.parse import urlsplit, urldefrag
candidate_re = re.compile(r'(?i)b(?:https?|ftp)://[^s<>"']+')
def extract_urls(text):
found = []
for raw in candidate_re.findall(text):
cleaned = raw.rstrip('.,;:!?)]}')
parts = urlsplit(cleaned)
if parts.scheme in {'http', 'https', 'ftp'} and parts.netloc:
url, _fragment = urldefrag(cleaned)
found.append(url)
return found
sample = 'Read https://example.com/docs, then visit (https://example.org/page#intro).'
print(extract_urls(sample))
urlsplit() separates scheme, authority, path, query, and fragment; urldefrag() removes the fragment. Python documents these APIs in urllib.parse. The returned values omit fragments by design. If fragments matter to your application, return cleaned instead of the URL from urldefrag().
Important limitation of the trimming line
rstrip('.,;:!?)]}') is a practical shortcut for common prose, not a universal punctuation algorithm. It strips closing brackets even when they may be part of a URL. A path can legitimately contain balanced parentheses, so blindly deleting every trailing ) can damage a valid link. If your input includes such URLs, track opening and closing delimiters or use source-specific parsing rules before trimming.
Extract and validate URLs in JavaScript
In JavaScript, use a regex to find candidates and the built-in URL constructor to parse them. This function returns normalized absolute URL strings for the allowed protocols.
function extractUrls(text, baseUrl) {
const rough = text.match(/b(?:https?|ftp)://[^s<>"']+/gi) ?? [];
return rough.flatMap(raw => {
const cleaned = raw.replace(/[.,;:!?)]}+$/, "");
try {
const parsed = new URL(cleaned, baseUrl);
if (!["http:", "https:", "ftp:"].includes(parsed.protocol)) return [];
return [parsed.href];
} catch {
return [];
}
});
}
console.log(extractUrls("See https://example.com/docs, then https://example.org/a."));
The constructor returns a parsed URL or throws when it cannot parse the input. For a candidate-only check, MDN also documents URL.canParse(). See MDN: URL. As in the Python example, the closing-parenthesis trim is deliberately simple and should be adjusted if balanced parentheses are meaningful in your URLs.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
Handle punctuation, wrappers, and relative references
Trailing punctuation and line breaks
In ordinary writing, commas and periods often follow a URL without whitespace. Strip only delimiters that are outside the link. RFC 3986 notes that punctuation marks can be mistaken for URI content, and that line wrapping can make boundaries unclear. A URL split across lines may therefore need preprocessing based on the format that produced the text; do not remove whitespace indiscriminately if it could alter the intended value.
Quoted and bracketed URLs
Inputs may wrap links in quotation marks, angle brackets, or parentheses, or use a legacy URL: prefix. Handle these wrappers only when they are clearly outside the candidate. Python’s URL parsing documentation describes these wrapped forms and parsing behavior: urllib.parse. Do not strip a character merely because it is a delimiter in prose; it could be part of a legitimate path or query.
Relative references and protocol-relative forms
A value such as /docs/page is a relative reference, not a complete network URL. Keep it relative unless you have a trusted, relevant base URL. In Python, use urllib.parse.urljoin(base_url, reference); in JavaScript, use new URL(reference, baseUrl). RFC 3986 defines reference resolution; see Section 5. A reference beginning // inherits its scheme from a base URL. The regexes above do not locate relative or protocol-relative references; add them only if your input requires them and your code has a trusted base.
Fragments, internationalized domains, and percent encoding
Decide whether fragments identify distinct results in your application. The Python function drops them, while the JavaScript example preserves them. URL parsers also apply their own normalization rules, which can affect how a parsed value is represented. Preserve the original text separately if display fidelity matters. Do not lowercase paths or decode percent escapes as a blanket cleanup step: reserved characters and encoding can have meaning specific to a URL’s scheme or application. RFC 3986 covers URI syntax and percent encoding at RFC 3986.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteApply a security policy before using extracted values
A syntactically parseable URL is not automatically safe to fetch, open, or trust. Before acting on extracted values, define an explicit policy:
- Allow only the schemes the task requires, commonly
httpsand sometimeshttp. Reject unexpected schemes such asjavascript:when results may be navigated to or fetched. - For network URLs, require a nonempty hostname and validate acceptable ports and hostnames for your application.
- Treat user information such as embedded usernames or passwords as sensitive; reject it if your use case does not allow credentials in URLs.
- If your application fetches user-supplied URLs, consider where a hostname resolves and whether the destination is permitted by your network policy. Parsing alone is not a fetch-safety guarantee.
- Keep the original string for display or auditing, and use a parsed representation for comparison and policy checks.
RFC 3986 discusses security considerations for URI construction and use. The rfc3986 library documentation describes validation options, including requiring schemes or hosts and forbidding passwords in user information. Use validation that matches your application; no single regex or parser policy fits every security boundary.
Regex or parser: what each part does
| Approach | Best use | Limitation |
|---|---|---|
| Candidate-matching regex | Finding likely absolute URLs in plain text | Cannot reliably decide whether punctuation belongs to a URL or whether a match is valid and safe |
| URL parser | Separating components and applying scheme, host, and other policy checks | Does not identify every URL buried in arbitrary prose by itself |
| HTML or Markdown parser | Extracting links from structured documents | Must match the source format and the kinds of link-bearing elements you need |
RFC 3986’s Appendix B provides a reference regular expression for decomposing a URI reference into components; that is different from finding every URL in arbitrary text. The standard describes parsing a URI reference into its major components. See Appendix B.
Troubleshoot common extraction problems
A period or comma appears at the end
The match includes sentence punctuation because there is no whitespace boundary. Trim likely delimiters after matching, but make the trimming logic sensitive to balanced brackets and the formats in your data.
Free tools Windows power users keep installed
One-click scans. No signup required.
A valid URL is missing
Check whether it uses a scheme your pattern does not include, is split across lines, or is a relative reference. Expand the candidate rules only for formats you need, then parse and validate the added forms separately.
A match is not a working link
Regex matching only found URL-like text. Confirm that parsing succeeds, that the scheme and host meet your policy, and that the destination itself is reachable if your application needs to fetch it.
A relative path is rejected
That is expected if the code is looking for absolute network URLs. Retain relative references as references, or resolve them against a trusted base URL; do not invent a base.
A URL changes after parsing
The parser may serialize or normalize the value. Keep the original substring if exact text is important, and avoid ad hoc case changes or percent-decoding. Compare normalized values only under rules appropriate to your application.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
Or skip the browser setup
If what you need is a screenshot of a page—not extracting links from text—ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. For example, save a WebP screenshot of a page with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses report the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Further reading
The primary references for URL syntax and the APIs used here are RFC 3986, Python’s urllib.parse documentation, and MDN’s URL documentation.
Frequently Asked Questions
Should fragments be included in extracted URLs?
It depends on what the result represents: keep fragments when the destination within a page matters, or remove them when you want page-level URLs. The Python example removes fragments; the JavaScript example preserves them.
Can a regular expression validate every URL?
No. A regex can locate likely candidates, but parsing and explicit application rules are needed to check components and decide whether a URL is acceptable.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

