๐Ÿ” Regular Expressions

โ† Back to Cheatsheets

A quick-reference for regex syntax. Regular expressions are patterns for matching and manipulating text. The core syntax is consistent across most tools โ€” Python's re, JavaScript's RegExp, grep -E, sed, and PCRE all share the same fundamentals, with minor differences in supported features.

Category: Pattern Language / DSL Syntax: PCRE / ERE compatible core Common engines: Python re ยท JavaScript RegExp ยท PCRE ยท grep -E
Resources: regex101 โ€” Live tester with explanation regexr.com โ€” Visual reference Python re module MDN โ€” JS Regex regular-expressions.info

Basics & Escaping

โ–ผ

Most characters match themselves literally. A handful of characters โ€” . ^ $ * + ? { } [ ] \ | ( ) โ€” are special and must be escaped with \ to be treated as literals.

Dot & Alternation
a.c        # "a" + any char (except \n) + "c"
a|b        # "a" OR "b"
use the s flag to make . match newlines too
Literal Characters
abc        # literal "abc"
(ab)+      # one or more repetitions of "ab"
any non-special character matches itself exactly
Escaping Special Characters
\.         # literal dot
\(         # literal parenthesis
\$         # literal dollar sign
https?://  # matches "http://" or "https://"
\d+\.\d+   # decimal number like "3.14"
any of . ^ $ * + ? { } [ ] \ | ( ) must be escaped to match literally
Raw Strings in Python
r"\d+"   # correct: raw string, \ passed to regex engine
"\\d+"   # also works but harder to read
always use raw strings (r"...") โ€” Python would otherwise consume the \ before the regex engine sees it

Anchors & Boundaries

โ–ผ

Anchors are zero-width โ€” they match a position, not a character. They don't consume input.

Position Anchors
^      # start of string (or line in multiline mode)
$      # end of string (or line in multiline mode)
\b     # word boundary: between \w and \W
\B     # non-word boundary
\A     # start of string (Python/Perl; ignores multiline flag)
\Z     # end of string (Python/Perl)
\A and \Z always refer to the full string regardless of flags
Examples
^\d+$        # entire string must be digits
\bword\b     # "word" as a whole word
^Error       # line starts with "Error" (multiline flag)
\bword\b won't match inside "sword" or "words"

Character Classes

โ–ผ

A character class matches one character from a defined set. Shorthands like \d are equivalent to their bracket form but more concise.

Bracket Expressions
[abc]        # matches a, b, or c
[^abc]       # matches anything except a, b, c
[a-z]        # lowercase letter
[A-Za-z0-9]  # alphanumeric
[a-z&&[^aeiou]]  # consonants (Java-style intersection)
^ inside [...] negates the class; a literal - must be first or last
Shorthand Classes
\d   # digit           [0-9]
\D   # non-digit        [^0-9]
\w   # word character   [a-zA-Z0-9_]
\W   # non-word         [^a-zA-Z0-9_]
\s   # whitespace       [ \t\n\r\f]
\S   # non-whitespace
\b   # word boundary    (zero-width, see Anchors)
.    # any char except newline  (see Basics)

Quantifiers

โ–ผ

Quantifiers apply to the preceding element (character, group, or class). By default they are greedy โ€” they match as much as possible. Append ? to make them lazy โ€” match as little as possible.

Main Quantifiers
*      # zero or more      colou*r  โ†’ "colour", "colouur", "colr"
+      # one or more       colou+r  โ†’ "colour", "colouur"  (not "colr")
?      # zero or one       colou?r  โ†’ "colour", "colr"
{n}    # exactly n         \d{4}    โ†’ exactly 4 digits
{n,}   # n or more         \d{2,}   โ†’ 2 or more digits
{n,m}  # between n and m   \d{2,4}  โ†’ 2, 3, or 4 digits
Greedy vs Lazy
<.+>   # greedy โ†’ matches "<b>bold</b>" (entire string)
<.+?>  # lazy   โ†’ matches "<b>" then "</b>" separately
lazy quantifiers stop at the earliest possible match; add ? after any quantifier to make it lazy

Groups & Capturing

โ–ผ

Parentheses group sub-expressions and capture their match for later use. Captured groups are numbered left-to-right by their opening parenthesis.

Capturing vs Non-Capturing
(abc)           # capturing group โ€” match available as \1 / $1
(?:abc)         # non-capturing โ€” grouping without saving
(?<name>abc)    # named group โ€” available as \k<name>
prefer (?:...) when you only need grouping, not a reference
Back-references
(\w+)\s+\1          # detect repeated word ("the the")
<(\w+)>.*?</\1>     # match opening + matching closing tag
(?<q>['"]).*?\k<q>  # string in matching quotes
back-reference must match the same text, not just the same pattern
Alternation
cat|dog           # "cat" or "dog"
(jpg|jpeg|png)$   # image file extension at end of string
| has the lowest precedence โ€” use (?:...) to scope alternation

Lookahead & Lookbehind

โ–ผ

Lookaround assertions are zero-width โ€” they check what's around the current position without consuming characters. Useful for adding conditions without including context in the match.

Syntax
(?=...)    # positive lookahead  โ€” must be followed by ...
(?!...)    # negative lookahead  โ€” must NOT be followed by ...
(?<=...)   # positive lookbehind โ€” must be preceded by ...
(?<!...)   # negative lookbehind โ€” must NOT be preceded by ...
Examples
\d+(?= dollars)      # digits followed by " dollars"
\bfoo(?!bar)\b       # "foo" not followed by "bar"
(?<=@)\w+           # word after "@" (e.g. domain in email)
(?<!\d)\d{4}(?!\d)  # exactly 4-digit number not adjacent to others
the match result contains only the part outside the lookaround assertion

Flags / Modifiers

โ–ผ

Flags change how the pattern is applied. In most languages they go after the closing delimiter: /pattern/flags. Python passes them as a third argument or via re.compile().

Common Flags
g   global       find ALL matches, not just the first
i   ignoreCase   [a-z] also matches [A-Z]
m   multiline    ^ and $ match start/end of each line
s   dotAll       . matches \n too (Python: re.DOTALL)
x   verbose      allow whitespace + comments in pattern
Python uses re.VERBOSE for x; not all engines support all flags
JavaScript Usage
const re = /pattern/gi;
"Hello World".replace(/o/gi, "0");  // "Hell0 W0rld"
grep Usage
# use -E for extended regex, -i for case-insensitive
grep -Ei "error|warning" app.log

Common Patterns

โ–ผ

Ready-to-use patterns for common tasks. Test and adjust for your specific requirements โ€” general-purpose patterns always have edge cases.

Validation Patterns
# Email (simplified โ€” RFC 5322 is far more complex)
[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}

# IPv4 address
\b(?:\d{1,3}\.){3}\d{1,3}\b

# Date  YYYY-MM-DD
\d{4}-(?:0[1-9]|1[0-2])-(?:0[1-9]|[12]\d|3[01])

# Hex colour  #fff or #ffffff
#(?:[0-9a-fA-F]{3}){1,2}

# URL slug
[a-z0-9]+(?:-[a-z0-9]+)*
Extraction Patterns
# Version number  e.g. "v1.2.3"
v?(\d+)\.(\d+)(?:\.(\d+))?

# Key=value pair
(\w+)\s*=\s*(.+)

# Content inside brackets
\[([^\]]+)\]     # [like this]
\(([^)]+)\)      # (like this)
Text Processing Patterns
# Trim leading/trailing whitespace
^\s+|\s+$

# Collapse multiple spaces to one
\s+  โ†’  replace with single space

# Detect duplicate adjacent words
\b(\w+)\s+\1\b

# Remove HTML tags
<[^>]+>

# Match a whole line containing "word"
^.*word.*$

Python re Module

โ–ผ

Python's built-in re module provides full PCRE-compatible regex support. Always use raw strings (r"...") for patterns to avoid double-escaping backslashes.

Core Functions
import re

re.match(r'\d+', text)        # match at START of string
re.search(r'\d+', text)       # search anywhere in string
re.findall(r'\d+', text)      # return list of all matches
re.finditer(r'\d+', text)     # return iterator of match objects
re.sub(r'\s+', ' ', text)     # replace matches with string
re.compile(r'\d+', re.I)      # compile pattern for reuse
re.match() only checks the beginning; use re.search() to find anywhere
Match Object
m = re.search(r'(\d+)-(\w+)', text)
if m:
    m.group(0)    # entire match
    m.group(1)    # first captured group
    m.groups()    # all groups as tuple
    m.groupdict() # named groups as dict
    m.start(), m.end(), m.span()
Flags
re.findall(r'\d+', text, re.IGNORECASE)
re.sub(r'\s+', ' ', text, flags=re.MULTILINE)

re.IGNORECASE  # re.I  โ€” case-insensitive
re.MULTILINE   # re.M  โ€” ^ and $ match each line
re.DOTALL      # re.S  โ€” dot matches \n too
re.VERBOSE     # re.X  โ€” allow whitespace + comments
combine flags with |: re.I | re.M
Verbose Pattern (re.VERBOSE)
pattern = re.compile(r"""
    (?P<year>  \d{4} )   # 4-digit year
    -
    (?P<month> 0[1-9] | 1[0-2] )  # month 01-12
    -
    (?P<day>   \d{2} )   # day
""", re.VERBOSE)
re.VERBOSE strips unescaped whitespace and # comments from the pattern