Regex flavor cheat sheet
Where JavaScript, PCRE2, Python, Java and Go/RE2 disagree, one row per difference, with the official source for every cell and an honest label when the source does not settle the question.
Column headers say whether the flavor is executed by this site's tester or notes only. Cells are marked documented, observed or unclear. Bracketed names link to the source list at the bottom of the page.
| Feature | JavaScript executed | PCRE2 executed | Python re notes only | Java Pattern notes only | Go regexp / RE2 executed |
|---|---|---|---|---|---|
| Named group syntaxThe most common reason a pattern copied from another language fails with "unknown extension" or "invalid group". | (?<name>...). Names must be unique in a pattern, except in different alternatives. documented [MDN] | All three spellings: (?<name>...), (?'name'...) and (?P<name>...). documented [PCRE2] | (?P<name>...) is the only form the docs list. Observed on Python 3.13.5: (?<n>a) raises "unknown extension ?<n". documented [Python docs] | (?<name>X). Observed on JDK 17: (?P<n>a) raises "Unknown inline modifier". documented [Oracle (Java SE 21)] | (?P<name>re) and (?<name>re). (?'name're) is listed as NOT SUPPORTED by RE2. documented [Go, Google RE2] |
| BackreferencesRE2-family engines drop backreferences so that matching stays linear-time. | \1 and \k<name>. In Unicode-unaware mode an invalid \1 is a legacy octal escape. documented [MDN, MDN] | \1, \g{1}, \g{name}, \k<name>, \k'name' and (?P=name). documented [PCRE2, PCRE2] | \1 and (?P=name). documented [Python docs] | \n and \k<name>. documented [Oracle (Java SE 21)] | Not supported (listed as NOT SUPPORTED by RE2). Go reports a syntax error. documented [Go, Google RE2] |
| Lookahead (?=...) (?!...)Lookahead is the usual way to say "followed by" without consuming text. | Supported. There is no backtracking into a lookahead. documented [MDN] | Supported: (?=...) and (?!...) assertions. documented [PCRE2] | Supported: (?=...) and (?!...). documented [Python docs] | Supported: (?=X) and (?!X). documented [Oracle (Java SE 21)] | Not supported (listed as NOT SUPPORTED by RE2). documented [Go, Google RE2] |
| Lookbehind (?<=...) (?<!...)Lookbehind length rules are the biggest difference between flavors that all "support" it. | Supported, and the contents may be variable length: the matcher runs the lookbehind backwards. documented [MDN] | Each top-level alternative must have a fixed length (up to 65535). From PCRE2 10.43, alternatives may differ in length, up to a limit set by the program (default 255); unlimited repetition such as \d* is not supported. documented [PCRE2] | The contained pattern must match strings of some fixed length. Observed on Python 3.13.5: (?<=a|bc)d is rejected with "look-behind requires fixed-width pattern". documented [Python docs] | Supported. The Javadoc read does not state a length rule. Observed on JDK 17 only: (?<=a+)b and (?<=a*)b compile and match. unclear [Oracle (Java SE 21)] | Not supported (listed as NOT SUPPORTED by RE2). documented [Go, Google RE2] |
| Atomic groups (?>...)Atomic groups discard backtracking points; they are one of the standard fixes for catastrophic backtracking. | Not part of JavaScript regular expression syntax (not in the MDN syntax reference). Observed in Node 20: (?>a) raises "Invalid group". observed [MDN] | Supported: (?>...). documented [PCRE2] | Supported from Python 3.11: (?>...). documented [Python docs] | Supported: (?>X), "an independent, non-capturing group". documented [Oracle (Java SE 21)] | Not supported (listed as NOT SUPPORTED by RE2). documented [Go, Google RE2] |
| Possessive quantifiers a++ a*+ a?+Possessive quantifiers are shorthand for an atomic group around a repeat. | Not part of JavaScript syntax. Observed in Node 20: a++ raises "Nothing to repeat". observed [MDN] | Supported: a++, a*+, a?+ and the {n,m}+ forms. documented [PCRE2] | Supported from Python 3.11: x*+, x++, x?+ and x{m,n}+ are equivalent to (?>x*) and so on. documented [Python docs] | Supported: X?+, X*+, X++ and the {n,m}+ forms. documented [Oracle (Java SE 21)] | Not supported (listed as NOT SUPPORTED by RE2). documented [Go, Google RE2] |
| Inline flags (?i) (?i:...)Where a flag may be switched on inside the pattern differs, and JavaScript only recently gained any form of it. | Scoped modifiers (?i:...), (?-i:...) for the flags i, m and s only. MDN marks them Baseline 2025 ("newly available"), so older browsers reject them. A bare (?i) is not allowed (observed in Node 20). documented [MDN] | (?i) and similar can appear inside the pattern, and scoped forms are supported. The PCRE2_CASELESS option or (?i) turns case-insensitive matching on. documented [PCRE2] | (?aiLmsux) is only allowed at the start of the expression (3.11+); scoped (?aiLmsux-imsx:...) is allowed anywhere. documented [Python docs] | (?idmsuxU-idmsuxU) and the scoped (?idmsuxU-idmsuxU:X). documented [Oracle (Java SE 21)] | (?flags) and (?flags:re) with the flags i, m, s and U. Flag syntax is xyz (set) or -xyz (clear). documented [Go, Google RE2] |
| Comments and free-spacingLong patterns are only maintainable if you can annotate them, and not every flavor lets you. | No comment syntax in a pattern; comment in the surrounding code. Observed in Node 20: (?#x)a raises "Invalid group". observed [MDN] | (?#...) comments, and the extended option (PCRE2_EXTENDED, flag x on this site) ignores most white space and allows # comments. documented [PCRE2] | (?#...) comments and the re.VERBOSE (re.X) flag. documented [Python docs] | The COMMENTS flag, (?x), permits white space and comments. Observed on JDK 17: (?#c)a raises "Unknown inline modifier". documented [Oracle (Java SE 21)] | (?#text) is listed as NOT SUPPORTED by RE2. documented [Go, Google RE2] |
| \d \w \s: ASCII or Unicode?The same \d matches Arabic-Indic digits in one flavor and only 0-9 in another. | \d is [0-9]. \w is A-Z, a-z, 0-9 and underscore. \s is white space plus line terminators (MDN lists NBSP and U+FEFF among white space; observed in Node 20: \s does not match U+0085). documented [MDN, MDN] | Characters above 127 never match \d, \s or \w unless PCRE2_UCP is set (flag U on this site); with UCP they use Unicode properties (\d is \p{Nd}). documented [PCRE2] | For str patterns \d matches any Unicode decimal digit (category Nd), \s Unicode white space and \w Unicode alphanumerics plus underscore; the re.ASCII flag restricts them to ASCII. documented [Python docs] | \d is [0-9], \s is [ \t\n\x0B\f\r] and \w is [a-zA-Z_0-9], unless UNICODE_CHARACTER_CLASS is set. documented [Oracle (Java SE 21)] | ASCII only: \d is [0-9], \s is [\t\n\f\r ], \w is [0-9A-Za-z_]. documented [Go, Google RE2] |
| Meaning of $ (without the multiline flag)In several flavors $ also matches before a final newline, which silently lets "abc\n" through a validation pattern. | End of input only. documented [MDN] | End of the subject, or immediately before a newline at the end, unless PCRE2_DOLLAR_ENDONLY is set. documented [PCRE2] | End of the string or just before the newline at the end of the string. documented [Python docs] | End of input, and also just before the last line terminator if nothing follows it. documented [Oracle (Java SE 21)] | End of text, "like \z not \Z". documented [Go, Google RE2] |
| Absolute anchors \A \z \ZThese are the safe way to say "whole string" in flavors where $ is lenient, but JavaScript barely has them. | \A, \z and \Z are marked experimental by MDN and apply in Unicode-aware mode only. Observed in Node 20: \A without the u flag is an identity escape for the letter A, and with u it is a syntax error. documented [MDN] | \A start, \z end, \Z end or before a final newline. documented [PCRE2] | \A and \Z (end of string only). \z is documented as added in Python 3.14; observed on Python 3.13.5: \z is a "bad escape". documented [Python docs] | \A, \Z (end but for the final terminator) and \z. documented [Oracle (Java SE 21)] | \A and \z. \Z is listed as NOT SUPPORTED. documented [Go, Google RE2] |
| Unicode properties \p{L}Matching "a letter" in any script needs a Unicode property in most flavors. | \p{...} and \P{...} need the u or v flag; without it they are identity escapes for the letter p or P. documented [MDN] | \p{xx} is supported (with UTF mode; this site always turns UTF mode on). documented [PCRE2] | The re module docs do not list \p. Observed on Python 3.13.5: \p{L} is a "bad escape". observed [Python docs] | \p{L} and the category forms are supported. documented [Oracle (Java SE 21)] | \pN, \p{Greek} and \P{...} are supported. documented [Go, Google RE2] |
| What . matchesWhether . crosses newlines, and whether it means a UTF-16 unit or a whole character, changes real results. | Any character except line terminators; with the s flag it also matches them. Without the u flag it matches one UTF-16 code unit, so an emoji needs two dots. documented [MDN] | Any one character except a newline; PCRE2_DOTALL (flag s) removes the exception. Which characters count as newline depends on the newline convention. documented [PCRE2] | Any character except a newline; the DOTALL flag also matches newline. documented [Python docs] | Any character except a line terminator unless DOTALL is specified. documented [Oracle (Java SE 21)] | "any character, possibly including newline (s=true)": the default is false, so . does not match \n. documented [Go, Google RE2] |
| Empty matches in a global search or replaceReplacing x* in "abxd" with "-" gives a different result in Go than in the other four. | x* over "abxd" with replacement "-" gives "-a-b--d-": an empty match is found right after the match "x" (observed in Node 20, matches ECMA-262 behaviour). observed [MDN, ECMA-262] | Observed through this site's PCRE2 engine: x* over "abxd" finds 5 matches (0-0, 1-1, 2-3, 3-3, 4-4), the same list as JavaScript. This depends on how the wrapper restarts the search after a match. observed [PCRE2] | re.sub('x*', '-', 'abxd') returns '-a-b--d-' (the docs say an empty match can occur immediately after a non-empty match). documented [Python docs] | Observed on JDK 17: Pattern.compile("x*").matcher("abxd").replaceAll("-") returns "-a-b--d-". The Javadoc pages read do not state the rule. observed [Oracle (Java SE 21)] | "Empty matches abutting a preceding match are ignored", so x* over "abxd" finds 4 matches (0-0, 1-1, 2-3, 4-4) and a replacement gives "-a-b-d-". documented [Go] |
| Limits on repeat countsA pattern with a large {n,m} can compile in one flavor and fail in another. | The MDN quantifier page read states no upper limit. unclear [MDN] | The numbers in {n,m} must be less than 65536. The maximum number of capture groups is 65535. documented [PCRE2] | The re docs read state no limit. unclear [Python docs] | The Javadoc read states no limit. unclear [Oracle (Java SE 21)] | Counting forms {n}, {n,}, {n,m} reject a minimum or maximum above 1000; unlimited repetition is not restricted. documented [Go, Google RE2] |
| Matching algorithm and runaway protectionThis decides whether a bad pattern or bad input can hang your program. | A backtracking engine: observed in Node 20, /^(a+)+$/ on 24 "a" characters plus "b" takes about 0.66 s, and each extra "a" doubles it. No built-in step limit, which is why this site runs patterns in a Worker it can terminate. observed [OWASP, MDN] | A backtracking matcher with a configurable match limit: an internal counter incremented each time round the main matching loop; reaching it returns PCRE2_ERROR_MATCHLIMIT. documented [PCRE2] | Backtracking (the Python HOWTO describes backtracking). Observed on Python 3.13.5: re.match('^(a+)+$', 'a'*26+'b') took about 2.4 s, about 4 times the 24-character case. observed [Python docs, OWASP] | The Javadoc read does not state the algorithm's complexity. Observed on JDK 17: ^(a+)+$ on 28 and 32 "a" characters plus "b" returned in about 2 ms, so do not assume the pattern is safe or unsafe; test your own. unclear [Oracle (Java SE 21)] | "guaranteed to run in time linear in the size of the input"; RE2 states the same linear guarantee and omits constructs that need backtracking. documented [Go, Google RE2] |
How to use
- Scan the Feature column for the construct that is failing in your program, or open the tester: it lists the rows that apply to the pattern you typed.
- Read across the row to see what each flavor does. The highlighted column is JavaScript. Every cell ends with its status and source.
- Follow the bracketed source name to the list at the bottom of this page, which links to the official document and says what it was used for.
- Paste a candidate pattern into the tester with the flavor you target. For JavaScript, PCRE2 and Go the answer comes from the real engine; the tests for this table run its probes the same way.
Worked examples
Porting a named group
A pattern written for Python, (?P<year>[0-9]{4}), is a syntax error in JavaScript and Java and works unchanged in PCRE2 and Go. In reverse, (?<year>…) is a syntax error in Python 3.13.5. The simplest portable choice between JavaScript, PCRE2, Java and Go is (?<name>…); between Python, PCRE2 and Go it is (?P<name>…).
The invisible trailing newline
The pattern ^[0-9]+$ rejects "123\n" in JavaScript and Go, but accepts it in PCRE2 by default, and (by their documentation) in Python and Java. Use \z in PCRE2 and Go, \Z or fullmatch() in Python, or matches() in Java when you mean "the whole string". The library pages run their patterns with \z on PCRE2 and Go for this reason.
Empty matches in a replace
Replacing x* with - in abxd gives -a-b--d- in JavaScript, Python and (observed) Java, and -a-b-d- in Go, because Go ignores an empty match right after the previous match. The tester shows the five spans in JavaScript and four in Go.
Limits & gotchas
- Only the five flavors on the page are covered. Perl, Ruby, .NET, PHP (PCRE) and others are not.
- Versions matter: Python's atomic groups and possessive quantifiers arrived in 3.11, PCRE2's variable-length lookbehind in 10.43, and JavaScript's inline modifiers are newly available. The table names the version where the source does.
- "Observed" results come from a single run of one version (Node 20, PCRE2 10.48 through this site's WebAssembly build, Go 1.24.4, Python 3.13.5, JDK 17). Your version may differ.
- The Python 3.14 documentation was read (the page header says 3.14.8) but the runtime used for observations was 3.13.5, so where the docs mention 3.14-only features (for example
\z) the observation shows the older behaviour. - The table is a map, not a substitute for the manual of your engine. Options set outside the pattern (newline conventions, UCP, locale) change results.
FAQ
What does "documented", "observed" and "unclear" mean in the table?
Documented means the cited official page states it. Observed means the page does not say, so we ran the real engine (the version is in the cell) and recorded what happened. Unclear means we read the page and it does not settle the question, so the cell says no more than that.
Why are Python and Java marked "notes only"?
This site runs JavaScript (your browser), PCRE2 and Go regexp as real engines. No Python or Java engine runs in the page. Their cells come from the official documentation, and the few "observed" notes come from running Python 3.13.5 and JDK 17 once on a developer machine.
Is Go the same as RE2?
Go's regexp package uses RE2 syntax and has the same linear-time guarantee, and its syntax page is a copy of the RE2 syntax with Go's own differences, but it is a separate implementation written in Go. This site cites both the Go regexp/syntax page and the RE2 wiki, and runs Go's implementation.
Which row causes the most bugs when porting a pattern?
Named-group syntax, lookbehind rules, the meaning of $ at a trailing newline, and whether \d \s \w are ASCII or Unicode. Each has a row, and every cell cites its source.
Sources
- MDN: Named capturing group: (?<name>...) Used for: Names must be unique within a pattern, except in different alternatives; groups object; unmatched named group is undefined.
- PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
- Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
- Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
- Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
- Google RE2: RE2 Syntax (wiki) Used for: RE2 syntax with explicit NOT SUPPORTED markers: lookaround, backreferences, possessive quantifiers, \Z, atomic groups.
- MDN: Backreference: \1, \2 Used for: Numeric backreferences; invalid ones become legacy octal escapes in Unicode-unaware mode.
- MDN: Named backreference: \k<name> Used for: \k<name> syntax and the Unicode-unaware-mode fallback to a literal k.
- PCRE2: pcre2syntax Used for: Quick syntax reference.
- MDN: Lookahead assertion: (?=...), (?!...) Used for: Zero-width assertions; no backtracking into a lookahead.
- MDN: Lookbehind assertion: (?<=...), (?<!...) Used for: Lookbehind matches backwards; JavaScript allows variable length where some languages forbid it; baseline since March 2023.
- MDN: Regular expressions (reference) Used for: List of flags (d g i m s u v y), the groups of syntax, and which syntaxes are assertions.
- MDN: Quantifier Used for: Greedy and lazy quantifiers, {n}, {n,}, {n,m}, no spaces inside braces, braces as literals in Unicode-unaware mode.
- MDN: Modifier: (?ims-ims:...) Used for: Inline i, m, s modifiers (bounded form only); Baseline 2025.
- MDN: Character class escape: \d, \D, \w, \W, \s, \S Used for: \d is [0-9]; \w is letters, digits and underscore; \s is whitespace plus line terminators.
- MDN: Whitespace (glossary) Used for: JavaScript white space: tab, vertical tab, form feed, space, no-break space, U+FEFF and other Space_Separator code points.
- MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
- MDN: Buffer boundary assertion: \A, \z, \Z Used for: Experimental in JavaScript, Unicode-aware mode only.
- MDN: Unicode character class escape: \p{...}, \P{...} Used for: \p and \P need the u or v flag.
- MDN: Wildcard: . Used for: . excludes line terminators unless the s flag is set; code units vs code points with the u flag.
- MDN: String.prototype.replace() Used for: Replacement patterns $$ $& $` $' $n $<Name> and their meaning.
- ECMA-262: Text processing: RegExp objects Used for: Regular expression semantics, CharacterClassEscape \s (WhiteSpace and LineTerminator), assertions.
- PCRE2: pcre2api Used for: Match limit (PCRE2_ERROR_MATCHLIMIT), default newline, DOLLAR_ENDONLY, pcre2_substitute replacement syntax ($1, ${1}, $<name>, $0 or $&).
- Oracle (Java SE 21): java.util.regex.Matcher: appendReplacement Used for: Replacement syntax $g, ${name}, backslash escapes.
- Go: regexp package Used for: Linear-time guarantee, leftmost-first semantics, Expand template syntax ($1, ${name}, $name longest-name rule).
- OWASP: Regular expression Denial of Service - ReDoS Used for: Explanation of backtracking, "evil" patterns such as (a+)+$, and the doubling of paths per extra character.
- MDN: Worker: terminate() Used for: terminate() stops a Worker immediately without letting it finish.
- Python docs: Regular Expression HOWTO Used for: Greedy matching and backtracking behaviour.
- Google RE2: README Used for: RE2 is a safe alternative to backtracking engines; linear running time; no constructs that need backtracking.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.