Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python’s built-in re module lets you check whether text matches a pattern, find and extract matches, replace text, or split text. Use raw string literals such as r"d+" for patterns, then choose the function that fits the job: fullmatch for validating an entire string, search for locating a match anywhere, and findall or finditer for extraction.

Start with a raw string pattern

Import Python’s standard-library regular-expression module, re, and write the pattern as a raw string. The Python Regular Expression HOWTO describes regexes as a small, specialized pattern language embedded in Python.

import re

pattern = r"d+"

The r prefix tells Python to treat backslashes literally instead of interpreting them as string escapes first. That matters for patterns such as d, which means a digit to the regex engine. Without a raw string, you typically need to double the backslash, as in "\d+". The re module reference warns that invalid Python string escape sequences can produce a SyntaxWarning and may become a SyntaxError.

Choose the operation that matches your goal

Goal Function What it does
Check a prefix re.match(pattern, text) Attempts a match only at the beginning of the string.
Find the first occurrence re.search(pattern, text) Scans the string and returns the first match anywhere.
Validate an entire string re.fullmatch(pattern, text) Matches only if the whole string fits the pattern.
Collect all matches re.findall(pattern, text) Returns matching text, or captured group values when the pattern has capturing groups.
Iterate over matches and inspect details re.finditer(pattern, text) Returns an iterator of Match objects, including captured groups and match positions.
Replace matches re.sub(pattern, replacement, text) Returns text with matching sections replaced.
Split at matches re.split(pattern, text) Splits text wherever the pattern matches.

For example, match is not a substitute for checking an entire value: it only tests the beginning. Use fullmatch when the whole string must satisfy your pattern.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build patterns from literals, classes, and quantifiers

Literal characters match themselves. Character classes describe a set of allowed characters, quantifiers control repetition, and anchors constrain where a pattern can match.

  • [A-Z] matches one uppercase ASCII letter; d matches a digit.
  • * means zero or more repetitions, + one or more, and ? zero or one. Use {m,n} to specify a repetition range.
  • ^ and $ anchor a pattern to positions in the text. Their line behavior can be changed with re.MULTILINE.
  • Parentheses capture text. Use (?:...) when grouping is needed without capturing, or (?P<name>...) to give a captured field a name.

The full syntax, including backreferences, is documented in the Python re reference. Prefer explicit boundaries and targeted character classes to an unrestricted .*, which can make a pattern harder to reason about.

Extract values with groups

Capturing groups let you pull structured fields from a match. Named groups are especially useful when extracted values have stable meanings, because code can refer to a descriptive name instead of a group number.

import re

text = "Order IDs: AB-123, CD-456"

m = re.search(r"(?P<code>[A-Z]{2})-(?P<number>d{3})", text)
if m:
    print(m.group("code"), m.group("number"))

A Match object’s .group() or .group(0) returns the complete match. Use .group(1) or a group name for captured text. .start(), .end(), and .span() provide the match’s positions in the string.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand findall’s result shape

With no capturing groups, findall returns the complete text for every match. If the pattern has one capturing group, it returns that group’s text for each match; with multiple groups, it returns tuples of captured values. When you need Match objects, spans, or several named fields, use finditer.

ids = re.findall(r"[A-Z]{2}-d{3}", text)
print(ids)  # ['AB-123', 'CD-456']

Replace text or split it at a pattern

Use re.sub for replacement and re.split for dividing text wherever a pattern matches. This example uses a whitespace pattern to reduce runs of spaces to one space:

clean = re.sub(r"s+", " ", "too   many spaces").strip()
print(clean)  # too many spaces

For replacements driven by captured fields, include groups in the pattern and use the replacement features documented in the re module reference.

Use flags to change matching behavior

Flags adjust how a pattern is interpreted. Pass one flag or combine several with bitwise OR:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
pattern = re.compile(r"^error", re.IGNORECASE | re.MULTILINE)
  • re.IGNORECASE (or re.I) ignores case distinctions.
  • re.MULTILINE (or re.M) makes ^ and $ work at line boundaries as well as string boundaries.
  • re.DOTALL (or re.S) lets . match newline characters.
  • re.ASCII (or re.A) limits shorthand character classes such as d to ASCII behavior.
  • re.VERBOSE (or re.X) lets you lay out a pattern with whitespace and comments for readability.

See the Python re documentation for the exact behavior of each flag and how it interacts with a pattern.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compile patterns reused in loops

re.compile(pattern, flags=0) creates a reusable Pattern object whose methods include matching, searching, finding, replacing, and splitting. It can make repeated use clearer and is useful when a pattern is accessed repeatedly in a loop.

order_id = re.compile(r"[A-Z]{2}-d{3}")

for line in lines:
    match = order_id.search(line)
    if match:
        print(match.group())

The HOWTO notes that module-level functions are convenient shortcuts and that the module caches recent patterns, so compiling is not necessarily faster for occasional calls outside a loop. Choose compilation mainly for reuse and clarity rather than assuming it improves every use.

Keep text types consistent and patterns bounded

Python’s regex engine works with Unicode str or 8-bit bytes, but a pattern and the value being searched must be the same type. Mixing a string pattern with bytes data, or the reverse, raises a type error.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use explicit character classes and boundaries instead of broad, unbounded matching when the accepted format is known.
  • If literal user input must become part of a pattern, pass it through re.escape() so regex metacharacters in that input are treated literally.
  • Test representative valid and invalid examples for the format you actually accept. A pattern does not automatically validate every possible email address, URL, or international format unless its accepted grammar is defined.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.