Regex pattern library

Ten patterns for tasks that come up in real code. Each page has a table of at least eight test cases that the automated tests run, a list of what the pattern wrongly accepts or rejects, flavor notes, and a note on when a regex is the wrong tool.

  • ISO 8601 date (YYYY-MM-DD)

    YYYY-MM-DD with month and day ranges, plus a leap-year-aware version checked against an independent calendar for every day number in the years 0000 to 2500.

    2 variants, 28 tested rows

  • IPv4 address

    Four octets from 0 to 255, strict about leading zeros, with a lenient variant and an exhaustive octet test.

    2 variants, 26 tested rows

  • UUID

    The 8-4-4-4-12 text form, a strict version/variant check, and the Nil and Max special cases.

    2 variants, 29 tested rows

  • Hex color

    The four CSS hex notations (3, 4, 6 or 8 digits) and a six-digit variant, with a test of every digit count.

    2 variants, 25 tested rows

  • URL slug

    Lowercase words joined by single hyphens: an ASCII version and a Unicode-aware one.

    2 variants, 24 tested rows

  • Semantic version

    The official semver.org pattern, tested on every valid example in the specification and a list of near misses.

    2 variants, 38 tested rows

  • Double-quoted string with escapes

    The JSON string grammar as a regex, and the naive "(...|\.)*" shortcut that accepts bad input and backtracks.

    2 variants, 28 tested rows

  • Trim whitespace

    Strip spaces from both ends: what \s matches in each flavor, a portable ASCII version, and why trim() wins.

    2 variants, 23 tested rows

  • Email-ish address

    The pattern browsers use for input type=email, a stricter variant, and an honest list of what neither can know.

    2 variants, 34 tested rows

  • URL path segment

    One path segment as RFC 3986 defines it: unreserved, sub-delims, ":" "@" and %XX escapes.

    2 variants, 28 tested rows

How to use

  1. Pick the task from the shelf above and open its page.
  2. Read the rule the pattern is judged against, with its source: an RFC, the WHATWG HTML Standard, semver.org or MDN. The pattern is only as right as that rule.
  3. Check the pass/fail table. Rows marked as a false positive or false negative are the pattern's known limits, each with a note.
  4. Press Open in tester to load the pattern and sample text into the tester, change flags or flavor, and see the token-by-token explanation.
  5. Read flavor notes before pasting the pattern into PCRE2, Python, Java or Go code, and Why not regex here before deciding a regex is the right tool at all.

Worked examples

What a table row tells you

On the IPv4 page the row 01.2.3.4 has the rule saying "invalid" (RFC 3986 allows no leading zero) and the strict pattern saying "no match": a pass. The lenient variant says "match" for the same input, so there the row is a known false positive, labelled as such. Reading the same input in two variants shows exactly what each pattern gives up.

Why the end anchor changes between flavors

Every table is produced with JavaScript's $, which matches only at the end of the input. PCRE2 and Go are run with \z in its place, because in PCRE2 $ also matches before a final newline. A pattern that is correct in JavaScript can therefore accept "1.2.3.4\n" in PCRE2 if you paste it unchanged.

Limits & gotchas

  • A pattern checks shape. It cannot check that a date exists on the calendar, that an address is reachable, or that a mailbox is real.
  • The tables are evidence for the listed inputs, not proof for all inputs. Where an exhaustive check is possible (octet values, calendar dates) the tests do it and the page says so.
  • Only JavaScript, PCRE2 and Go are executed for the cross-checks. Python and Java statements come from their documentation and are labelled as notes.
  • Standards change. Each page cites the document version it was read from, with the date it was read.

FAQ

How were the pass/fail tables checked?

Each row's expected answer was worked out by hand from the cited standard before it was compared with the browser's RegExp. The automated tests then run every row, run the same rows on PCRE2 and Go where the pattern is portable, and compare the patterns against independent reference code (for example every calendar date over 2501 years for the ISO date pattern).

Why does a page say "known false positive"?

A false positive is an input the pattern accepts but the rule rejects. Every pattern here has some, because regular expressions cannot check things like whether 31 February exists. The tables show those rows instead of hiding them.

Why are there only ten patterns?

A page is published only if it has at least eight tested cases per pattern, listed false positives and negatives, flavor notes and an honest note on when not to use a regex. More patterns are planned only when they meet the same bar.

Can I use these patterns in other languages?

Usually yes with small changes, and each page lists them. The most common change is replacing the end anchor $ with \z (PCRE2, Go) or using a full-match function, because $ accepts a trailing newline in several flavors.

Sources

  1. MDN: Regular expressions (reference) Used for: List of flags (d g i m s u v y), the groups of syntax, and which syntaxes are assertions.
  2. MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
  3. PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
  4. Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
  5. Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
  6. Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.

Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.