Regex for a UUID (RFC 9562 text format)
Two tested patterns for the hexadecimal 8-4-4-4-12 UUID text form: one that checks only the shape, and one that also checks the version digit (1 to 8) and the variant bits defined by RFC 9562, with the Nil and Max UUIDs covered as the special cases they are.
Shape only (any 128-bit UUID text)
Rule: RFC 9562 text format: 32 hexadecimal digits in 8-4-4-4-12 groups separated by hyphens; upper, lower or mixed case.
/^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$/ Open in tester Loads the pattern with sample text, JavaScript flavor.
Parts of the pattern
[0-9a-fA-F]{8}- Eight hexadecimal digits (time_low in version 1). RFC 9562: the letters may be upper, lower or mixed case.
-[0-9a-fA-F]{4}- Hyphen and four digits, repeated for the second, third and fourth groups.
-[0-9a-fA-F]{12}- Hyphen and twelve digits for the last group.
Test cases
14 of 14 rows agree with the rule.
| Input | Rule says | Pattern says | Result | Note |
|---|---|---|---|---|
f81d4fae-7dec-11d0-a765-00a0c91e6bf6 | valid | match | pass | The example UUID in RFC 9562 section 4 (version 1, variant 10xx). |
F81D4FAE-7DEC-11D0-A765-00A0C91E6BF6 | valid | match | pass | Upper case is allowed by the ABNF. |
f81D4fae-7DEC-11d0-a765-00A0c91e6bf6 | valid | match | pass | Mixed case is allowed. |
00000000-0000-0000-0000-000000000000 | valid | match | pass | The Nil UUID (RFC 9562 section 5.9). |
ffffffff-ffff-ffff-ffff-ffffffffffff | valid | match | pass | The Max UUID (section 5.10). |
919108f7-52d1-4320-9bac-f847db4148a8 | valid | match | pass | The UUIDv4 test vector in RFC 9562 Appendix A.3. |
f81d4fae7dec11d0a76500a0c91e6bf6 | invalid | no match | pass | No hyphens. A common storage form, but not the RFC 9562 text format. |
f81d4fae-7dec-11d0-a765-00a0c91e6bf | invalid | no match | pass | Last group has 11 digits. |
f81d4fae-7dec-11d0-a765-00a0c91e6bf6a | invalid | no match | pass | Last group has 13 digits. |
g81d4fae-7dec-11d0-a765-00a0c91e6bf6 | invalid | no match | pass | "g" is not a hexadecimal digit. |
{f81d4fae-7dec-11d0-a765-00a0c91e6bf6} | invalid | no match | pass | Braces: a Microsoft-style rendering, not the RFC format. |
urn:uuid:f81d4fae-7dec-11d0-a765-00a0c91e6bf6 | invalid | no match | pass | The URN form (RFC 9562 Figure 4) has the prefix urn:uuid:. |
f81d4fae-7dec-11d0-a765-00a0c91e6bf6\n | invalid | no match | pass | Trailing newline: JavaScript's $ rejects it, PCRE2, Python and Java's $ would accept it. |
00000000-0000-0000-0000-000000000001 | valid | match | pass | Shape-valid. Variant bits 0 mean it is in the reserved NCS range, but it is still 128 bits in the text form. |
RFC 9562 variant, version 1-8 (plus Nil and Max)
Rule: A UUID in the variant defined by RFC 9562 (top bits 10, so the first digit of the fourth group is 8, 9, a or b) with a version digit 1 to 8, or the Nil UUID or the Max UUID.
/^(?:[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|[fF]{8}-[fF]{4}-[fF]{4}-[fF]{4}-[fF]{12})$/ Open in tester Loads the pattern with sample text, JavaScript flavor.
Parts of the pattern
[1-8]- The first digit of the third group is the version. RFC 9562 Table 2 defines versions 1 to 8 (version 0 is "Unused" and 9 to 15 are "Reserved for future definition").
[89abAB]- The first digit of the fourth group carries the variant bits. For the variant specified in RFC 9562 the two top bits are 10, which is the hex digits 8, 9, A and B (Table 1).
0{8}-0{4}-0{4}-0{4}-0{12}- The Nil UUID, written out in full: it has variant bits 0 and version 0, so the first alternative cannot match it.
[fF]{8}-…-[fF]{12}- The Max UUID, all ones. Its version nibble is F and its variant nibble is F, so it also needs its own alternative.
Test cases
15 of 15 rows agree with the rule.
| Input | Rule says | Pattern says | Result | Note |
|---|---|---|---|---|
f81d4fae-7dec-11d0-a765-00a0c91e6bf6 | valid | match | pass | RFC example: version 1, variant digit "a". |
919108f7-52d1-4320-9bac-f847db4148a8 | valid | match | pass | RFC 9562 Appendix A.3 UUIDv4: version 4, variant digit "9". |
5df41881-3aed-3515-88a7-2f4a814cf09e | valid | match | pass | RFC 9562 Appendix A.2 UUIDv3: version 3, variant digit "8". |
2ed6657d-e927-568b-95e1-2665a8aea6a2 | valid | match | pass | Appendix A.4 UUIDv5. |
1EC9414C-232A-6B00-B3C8-9F6BDECED846 | valid | match | pass | Appendix A.5 UUIDv6, upper case, variant digit "B". |
017F22E2-79B0-7CC3-98C4-DC0C0C07398F | valid | match | pass | Appendix A.6 UUIDv7. |
2489E9AD-2EE2-8E00-8EC9-32D5F69181C0 | valid | match | pass | Appendix B.1 UUIDv8. |
00000000-0000-0000-0000-000000000000 | valid | match | pass | Nil UUID: allowed through its own alternative. |
FFFFFFFF-FFFF-FFFF-FFFF-FFFFFFFFFFFF | valid | match | pass | Max UUID. |
f81d4fae-7dec-01d0-a765-00a0c91e6bf6 | invalid | no match | pass | Version 0 is "Unused" in RFC 9562 Table 2. |
f81d4fae-7dec-91d0-a765-00a0c91e6bf6 | invalid | no match | pass | Version 9 is reserved for future definition. |
f81d4fae-7dec-11d0-c765-00a0c91e6bf6 | invalid | no match | pass | Variant digit "c" (110x) is the Microsoft backward-compatibility range, not the RFC 9562 variant. |
f81d4fae-7dec-11d0-0765-00a0c91e6bf6 | invalid | no match | pass | Variant digit "0" is the reserved NCS range. |
f81d4fae-7dec-11d0-a765-00a0c91e6bf | invalid | no match | pass | Too short. |
ffffffff-ffff-ffff-ffff-fffffffffffe | invalid | no match | pass | Almost Max: not the Max UUID, and version F / variant F fail the main alternative. |
How it works
RFC 9562 defines the text format as 4hexOctet "-" 2hexOctet "-" 2hexOctet "-" 2hexOctet "-" 6hexOctet, that is, 32 hexadecimal digits in groups of 8, 4, 4, 4 and 12, and notes that the letters may be upper case, lower case or mixed. The shape pattern is that ABNF translated directly.
The strict pattern adds two checks that live in single digits. The version is the top nibble of octet 6, which is the first hex digit of the third group. The variant is the top bits of octet 8, which is the first hex digit of the fourth group: RFC 9562 Table 1 lists 8-9 and A-B for the variant it defines, C-D as reserved for Microsoft backward compatibility, E-F as reserved for future definition (and the Max UUID), and 0-7 as the reserved NCS range (which includes the Nil UUID).
Because the Nil UUID has version 0 and variant 0, and the Max UUID has version F and variant F, they cannot satisfy the version and variant check. RFC 9562 defines them as special forms, so the pattern lists them as separate alternatives instead of weakening the check for everyone.
The strict table includes the example UUIDs printed in the RFC 9562 appendices for versions 3, 4, 5, 6, 7 and 8, plus the section 4 example for version 1. For each one the version digit and variant digit were read off by hand and written down before the pattern was run.
Known false positives
- The shape pattern accepts any 128 bits written in the right layout: it does not know that a version-4 UUID should have the digit 4 in the third group.
- The strict pattern only checks the version and variant digits. It does not check that the contents make sense for the version: a version-1 UUID is supposed to carry a timestamp and node, a version-5 UUID a SHA-1 hash, and the pattern cannot verify either.
Known false negatives
- Neither pattern accepts the compact 32-digit form without hyphens, the URN form urn:uuid:…, or the braced form. Normalise those before matching if you need to accept them.
- The strict pattern rejects UUIDs in the reserved NCS, Microsoft and future variants (variant digits 0-7 and C-F), apart from Nil and Max. Some older systems generated UUIDs with variant digits in those ranges; if you must accept them use the shape pattern.
- Versions above 8 are rejected because RFC 9562 reserves them. If a future specification defines version 9, the strict pattern will need an update.
Flavor notes
- All of the character classes use explicit ranges, so they mean the same in every flavor. A case-insensitive flag can replace the A-F ranges in JavaScript (the i flag), PCRE2 (caseless option), Python (re.IGNORECASE), Java (CASE_INSENSITIVE) and Go ((?i)), but then every other letter in the pattern becomes case-insensitive too, so this page spells the cases out.
- Trailing newline: JavaScript and Go's $ match only at the very end of the input; PCRE2 (default), Python and Java's $ also match before a final newline. Use \z, \Z, fullmatch() or matches() as appropriate when validating.
- The strict pattern uses only alternation, classes and counted repeats, so JavaScript, PCRE2 and Go all accept it, and the test suite runs it on those three real engines. Python and Java are notes only on this site; nothing is executed on them.
Why not regex here?
If you only need a fresh random identifier, do not parse and pattern-check strings you generated yourself: crypto.randomUUID() returns a version 4 UUID from a cryptographically secure generator (MDN, Crypto.randomUUID), and your database or library already knows the type.
Use the regex at a boundary where text arrives from outside, to reject malformed input before it reaches code that assumes a UUID. Prefer the shape pattern there unless you really need to restrict versions.
How to use
- Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
- Check how it is anchored: the patterns use
^and$to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases. - If your language is not JavaScript, read Flavor notes and change
$to\zor use a full-match function. - Open it in the tester to see the explanation of each token and try your own inputs.
Worked examples
Accept any UUID from an API field
Use the shape pattern with the whole-string anchors. Lower-case the value afterwards if you store it, because RFC 9562 treats the letters as case-insensitive in the text form.
Only accept modern time-ordered UUIDs
Narrow the version class to [67] to accept only version 6 and version 7, which RFC 9562 defines as time-based (reordered Gregorian and Unix Epoch). Try it in the tester with the RFC 9562 test vectors.
Limits & gotchas
- A pattern match proves the text is well-formed, not that the identifier exists or is unique.
- The strict pattern is tied to RFC 9562 as read on 2026-10-02. A later revision that defines new versions or variants changes what "correct" means.
FAQ
Are UUIDs case-sensitive?
In the RFC 9562 text format the hexadecimal letters may be upper case, lower case or mixed. Both patterns here accept all three. If you compare UUIDs as strings, normalise the case first.
Why does the strict pattern need separate alternatives for the Nil and Max UUIDs?
The Nil UUID is all zeros and the Max UUID is all ones. They are defined in RFC 9562 sections 5.9 and 5.10 as special forms, but their version and variant digits (0 and F) fall outside the ranges that identify an ordinary UUID.
Does the pattern tell me which version a UUID is?
Not by itself, but the version is the first digit of the third group, so you can capture it with a group such as ^[0-9a-f]{8}-[0-9a-f]{4}-([1-8])[0-9a-f]{3}- and read group 1.
Should I validate UUIDs with a regex in production code?
It is reasonable as a boundary check for untrusted text. Parse the value into your platform's UUID type for everything after that, because the type guarantees well-formedness for the rest of the program.
Sources
- IETF: RFC 9562: UUIDs Used for: UUID text format (hex-and-dash, any case), variant bits 10xx (8-B), versions table, Nil and Max UUID.
- MDN: Crypto: randomUUID() method Used for: randomUUID() generates a v4 UUID using a cryptographically secure random number generator.
- MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
- PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
- Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
- Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
- Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
- Google RE2: RE2 Syntax (wiki) Used for: RE2 syntax with explicit NOT SUPPORTED markers: lookaround, backreferences, possessive quantifiers, \Z, atomic groups.
- MDN: Regular expressions (reference) Used for: List of flags (d g i m s u v y), the groups of syntax, and which syntaxes are assertions.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.