Regex pattern library
Ten patterns for tasks that come up in real code. Each page has a table of at least eight test cases that the automated tests run, a list of what the pattern wrongly accepts or rejects, flavor notes, and a note on when a regex is the wrong tool.
-
ISO 8601 date (YYYY-MM-DD)
YYYY-MM-DD with month and day ranges, plus a leap-year-aware version checked against an independent calendar for every day number in the years 0000 to 2500.
-
IPv4 address
Four octets from 0 to 255, strict about leading zeros, with a lenient variant and an exhaustive octet test.
-
UUID
The 8-4-4-4-12 text form, a strict version/variant check, and the Nil and Max special cases.
-
Hex color
The four CSS hex notations (3, 4, 6 or 8 digits) and a six-digit variant, with a test of every digit count.
-
URL slug
Lowercase words joined by single hyphens: an ASCII version and a Unicode-aware one.
-
Semantic version
The official semver.org pattern, tested on every valid example in the specification and a list of near misses.
-
Double-quoted string with escapes
The JSON string grammar as a regex, and the naive "(...|\.)*" shortcut that accepts bad input and backtracks.
-
Trim whitespace
Strip spaces from both ends: what \s matches in each flavor, a portable ASCII version, and why trim() wins.
-
Email-ish address
The pattern browsers use for input type=email, a stricter variant, and an honest list of what neither can know.
-
URL path segment
One path segment as RFC 3986 defines it: unreserved, sub-delims, ":" "@" and %XX escapes.
How to use
- Pick the task from the shelf above and open its page.
- Read the rule the pattern is judged against, with its source: an RFC, the WHATWG HTML Standard, semver.org or MDN. The pattern is only as right as that rule.
- Check the pass/fail table. Rows marked as a false positive or false negative are the pattern's known limits, each with a note.
- Press Open in tester to load the pattern and sample text into the tester, change flags or flavor, and see the token-by-token explanation.
- Read flavor notes before pasting the pattern into PCRE2, Python, Java or Go code, and Why not regex here before deciding a regex is the right tool at all.
Worked examples
What a table row tells you
On the IPv4 page the row 01.2.3.4 has the rule saying "invalid" (RFC 3986 allows no leading zero) and the strict pattern saying "no match": a pass. The lenient variant says "match" for the same input, so there the row is a known false positive, labelled as such. Reading the same input in two variants shows exactly what each pattern gives up.
Why the end anchor changes between flavors
Every table is produced with JavaScript's $, which matches only at the end of the input. PCRE2 and Go are run with \z in its place, because in PCRE2 $ also matches before a final newline. A pattern that is correct in JavaScript can therefore accept "1.2.3.4\n" in PCRE2 if you paste it unchanged.
Limits & gotchas
- A pattern checks shape. It cannot check that a date exists on the calendar, that an address is reachable, or that a mailbox is real.
- The tables are evidence for the listed inputs, not proof for all inputs. Where an exhaustive check is possible (octet values, calendar dates) the tests do it and the page says so.
- Only JavaScript, PCRE2 and Go are executed for the cross-checks. Python and Java statements come from their documentation and are labelled as notes.
- Standards change. Each page cites the document version it was read from, with the date it was read.
FAQ
How were the pass/fail tables checked?
Each row's expected answer was worked out by hand from the cited standard before it was compared with the browser's RegExp. The automated tests then run every row, run the same rows on PCRE2 and Go where the pattern is portable, and compare the patterns against independent reference code (for example every calendar date over 2501 years for the ISO date pattern).
Why does a page say "known false positive"?
A false positive is an input the pattern accepts but the rule rejects. Every pattern here has some, because regular expressions cannot check things like whether 31 February exists. The tables show those rows instead of hiding them.
Why are there only ten patterns?
A page is published only if it has at least eight tested cases per pattern, listed false positives and negatives, flavor notes and an honest note on when not to use a regex. More patterns are planned only when they meet the same bar.
Can I use these patterns in other languages?
Usually yes with small changes, and each page lists them. The most common change is replacing the end anchor $ with \z (PCRE2, Go) or using a full-match function, because $ accepts a trailing newline in several flavors.
Sources
- MDN: Regular expressions (reference) Used for: List of flags (d g i m s u v y), the groups of syntax, and which syntaxes are assertions.
- MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
- PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
- Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
- Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
- Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.