Regex for an email-ish address (and what it cannot tell you)

A regular expression can check that text looks like an email address. It cannot check that the address exists, that the mailbox belongs to the person typing, or that it follows the full RFC 5322 grammar. This page tests the pattern the HTML standard uses, a stricter variant, and lists exactly what each accepts and rejects.

WHATWG HTML pattern (input type=email)

Rule: The WHATWG HTML Standard's "valid email address": one or more atext characters or dots, an at sign, then dot-separated labels of letters, digits and hyphens (1 to 63 characters, not starting or ending with a hyphen).

/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

[a-zA-Z0-9.!#$%&'*+/=?^_`{|}~-]+
The local part: one or more of the RFC 5322 atext characters or a dot. Note that a dot is allowed anywhere, including first, last or doubled.
@
A literal at sign.
[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?
One domain label: starts and ends with a letter or digit, up to 61 letters, digits or hyphens in between, so at most 63 characters (the limit from RFC 1034 section 3.5, per the standard).
(?:\.label)*
Zero or more further labels, each after a dot. Zero is allowed: user@localhost passes.

Test cases

23 of 23 rows agree with the rule.

Test cases for WHATWG HTML pattern (input type=email)
Input Rule says Pattern says Result Note
user@example.com valid match pass Ordinary address.
first.last+tag@sub.example.co.uk valid match pass Dots and plus in the local part, several labels.
o'brien@example.org valid match pass An apostrophe is atext.
a@b valid match pass Valid under the WHATWG pattern: no dot is required in the domain, so user@localhost passes.
x@example.com valid match pass One-character local part.
a@ + 63-character label + .com valid match pass A 63-character label is the maximum.
a@ + 64-character label + .com invalid no match pass 64 characters in one label is too long.
user@@example.com invalid no match pass Two at signs.
@example.com invalid no match pass Empty local part.
user@ invalid no match pass Empty domain.
user@-example.com invalid no match pass Label starts with a hyphen.
user@example-.com invalid no match pass Label ends with a hyphen.
user@exam_ple.com invalid no match pass Underscore is not allowed in a domain label.
user@example..com invalid no match pass Empty label between dots.
user@example.com. invalid no match pass Trailing dot (the fully qualified root form) is not accepted.
user␣name@example.com invalid no match pass Space.
user@example.com\n invalid no match pass Trailing newline: rejected by JavaScript's $; PCRE2, Python and Java's $ would accept it.
"quoted␣name"@example.com invalid no match pass Valid under RFC 5322 (quoted-string local part) but a "willful violation" in the WHATWG definition, so rejected.
üser@example.com invalid no match pass Valid for internationalized mail (RFC 6531 extends atext with non-ASCII), rejected here: ASCII only.
user@[192.168.0.1] invalid no match pass Domain literal: valid RFC 5322, rejected here.
.user@example.com valid match pass Accepted by the WHATWG rule even though RFC 5322 dot-atom forbids a leading dot. See the stricter variant.
user.@example.com valid match pass Same: trailing dot in the local part accepted by WHATWG.
us..er@example.com valid match pass Same: doubled dot accepted by WHATWG.

Stricter local part (RFC 5322 dot-atom)

Rule: Local part is an RFC 5322 dot-atom-text (atext runs separated by single dots, so no leading, trailing or doubled dot); domain as in the WHATWG definition.

/^[a-zA-Z0-9!#$%&'*+\/=?^_`{|}~-]+(?:\.[a-zA-Z0-9!#$%&'*+\/=?^_`{|}~-]+)*@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

[atext]+(?:\.[atext]+)*
A run of atext characters, then zero or more groups of one dot and another run. This is RFC 5322 dot-atom-text = 1*atext *("." 1*atext).
rest
Same @ and domain labels as the WHATWG pattern.

Test cases

11 of 11 rows agree with the rule.

Test cases for Stricter local part (RFC 5322 dot-atom)
Input Rule says Pattern says Result Note
user@example.com valid match pass Ordinary address.
first.last@example.com valid match pass Single dot.
a.b.c.d@example.com valid match pass Several dots.
user+tag@example.com valid match pass Plus sign.
.user@example.com invalid no match pass Leading dot rejected.
user.@example.com invalid no match pass Trailing dot rejected.
us..er@example.com invalid no match pass Doubled dot rejected.
user@example..com invalid no match pass Empty domain label.
user@-example.com invalid no match pass Label starts with a hyphen.
"quoted"@example.com invalid no match pass Quoted local part not supported (valid in RFC 5322).
user@localhost valid match pass No dot required in the domain.

How it works

The WHATWG HTML Standard defines a valid email address with an ABNF: 1*( atext / "." ) "@" label *( "." label ), where atext is the RFC 5322 set of printable characters and the label comes from RFC 1034 (letters, digits, hyphens, at most 63 characters, starting and ending with a letter or digit). It then says this requirement is "a willful violation of RFC 5322", which it calls too strict before the at sign, too vague after it, and too lax in allowing comments, whitespace and quoted strings. The first pattern is the JavaScript- and Perl-compatible expression the standard prints for that definition.

The second pattern tightens one thing only: the local part must follow RFC 5322's dot-atom-text, which does not allow a leading, trailing or doubled dot. That is a change this page makes to the standard's pattern; the standard itself says a dot may appear anywhere. It does not change the domain part.

What neither pattern decides: whether the domain exists, whether it accepts mail, whether the mailbox exists, or whether the person typing it owns it. Only sending a message and having the owner act on it (for example, click a link) proves that. This is why so many sites send a confirmation email and treat the pattern as a typo filter.

Known false positives

  • a@b and user@localhost pass: neither has a dot in the domain, which the standard allows. Add (?:\.label)+ instead of * if you need at least one dot.
  • The WHATWG pattern accepts a leading, trailing or doubled dot in the local part; the stricter variant rejects those.
  • Both patterns accept addresses that look right but cannot receive mail: no-one@example.com, or a well-formed address on a domain that does not exist.

Known false negatives

  • Addresses with quoted local parts ("john doe"@example.com), domain literals (user@[192.168.0.1]) and comments are valid under RFC 5322 but rejected by both patterns.
  • Internationalized addresses with non-ASCII characters in the local part are rejected. RFC 6531 extends atext with non-ASCII characters (UTF8-non-ascii) for mail systems that support it. Non-ASCII domain names must be converted to their ASCII (punycode) form before the pattern sees them; browsers do this in input type=email, and the standard describes that conversion.
  • A trailing dot (user@example.com.) is rejected.

Flavor notes

  • The patterns use a character class with many punctuation characters. Inside a class, only ], \, ^ (first) and - (between characters) are special in the flavors on this site, which is why the hyphen is placed last. The backslash before the slash is only needed in a JavaScript regex literal; it is harmless in a string.
  • Neither pattern uses lookaround, so both compile in JavaScript, PCRE2 and Go (RE2 syntax) and are run on those real engines by the tests. Python and Java are notes only on this site.
  • Trailing newline: PCRE2 (default), Python and Java let $ match before a final newline, so "user@example.com" followed by a line feed would pass; JavaScript and Go reject it. Use \z, \Z, fullmatch() or matches() in those flavors.

Why not regex here?

For a form on a web page, let the browser do the first check: <input type="email"> applies the same WHATWG definition and handles internationalized domain names for you. Then send a confirmation email. The regex is useful where no browser is involved, for example when a server must reject obvious junk before queueing mail.

Do not write a regex that tries to implement all of RFC 5322. The standard's authors judged the full grammar unfit for validation and chose a simpler definition on purpose.

How to use

  1. Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
  2. Check how it is anchored: the patterns use ^ and $ to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases.
  3. If your language is not JavaScript, read Flavor notes and change $ to \z or use a full-match function.
  4. Open it in the tester to see the explanation of each token and try your own inputs.

Worked examples

Check a pasted list

Remove the anchors, add the g and m flags and run the WHATWG pattern over a text with one address per line: each valid address is highlighted.

Require a dot in the domain

Change the final (?:\.label)* to (?:\.label)+ so user@localhost is rejected but user@example.com still passes.

Limits & gotchas

  • A match means "has the shape of an address", nothing more.
  • The 63-character label limit is checked; the 254-character limit for a whole address (a separate rule that this page did not read an official source for) is not.

FAQ

Is there a regex that fully validates email addresses?

Not one that is useful. RFC 5322 allows comments, quoted strings and domain literals, and the WHATWG standard calls that grammar unsuitable for form validation. Use a simpler pattern as a typo filter and confirm the address by sending mail.

Why does user@localhost pass?

The WHATWG definition allows a domain made of a single label. If you need a dot, change the domain part to require at least one.

Why is "john doe"@example.com rejected?

A quoted local part is legal in RFC 5322, but the WHATWG definition deliberately excludes it, and so does this pattern. Real-world use of such addresses is very rare.

Should I use this pattern or type=email?

In a browser form use type=email, which applies the same definition. Use the pattern for server-side checks and tests, and treat it as a first filter only.

Sources

  1. WHATWG HTML Standard: Valid email address (input type=email) Used for: The email production, the JavaScript- and Perl-compatible pattern, and the "willful violation of RFC 5322" statement.
  2. IETF: RFC 5322: Internet Message Format Used for: dot-atom-text and the full address grammar that is too complex for a short pattern.
  3. IETF: RFC 6531: SMTP Extension for Internationalized Email Used for: atext extended with UTF8-non-ascii, so non-ASCII characters can appear in an internationalized local part.
  4. MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
  5. PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
  6. Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
  7. Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
  8. Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
  9. Google RE2: RE2 Syntax (wiki) Used for: RE2 syntax with explicit NOT SUPPORTED markers: lookaround, backreferences, possessive quantifiers, \Z, atomic groups.

Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.