Regex for one URL path segment (RFC 3986)

The text between two slashes in a URL path has a precise definition in RFC 3986. This page turns it into a pattern, tests it, and adds a variant that rejects the dot-segments "." and "..", which matter for routing and file access.

RFC 3986 segment-nz (one or more pchar)

Rule: RFC 3986 segment-nz = 1*pchar, where pchar = unreserved / pct-encoded / sub-delims / ":" / "@".

/^(?:[A-Za-z0-9\-._~!$&'()*+,;=:@]|%[0-9A-Fa-f]{2})+$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

[A-Za-z0-9\-._~]
unreserved: letters, digits and the four marks - . _ ~ (RFC 3986 unreserved).
[!$&'()*+,;=]
sub-delims: the eleven punctuation marks RFC 3986 lists as sub-delims.
[:@]
The two extra pchar characters.
%[0-9A-Fa-f]{2}
pct-encoded: a percent sign and exactly two hexadecimal digits.
(?:…)+
One or more of those. segment-nz forbids an empty segment; the RFC's plain segment = *pchar allows it.

Test cases

19 of 19 rows agree with the rule.

Test cases for RFC 3986 segment-nz (one or more pchar)
Input Rule says Pattern says Result Note
about valid match pass Plain word.
hello-world_2026.html valid match pass Hyphen, underscore and dot are unreserved.
a%20b valid match pass Percent-encoded space.
%E2%82%AC valid match pass Percent-encoded UTF-8 for the euro sign.
user@host valid match pass "@" is a pchar.
a:b valid match pass ":" is a pchar.
a=b&c valid match pass "=" and "&" are sub-delims.
~tilde valid match pass "~" is unreserved.
a␣b invalid no match pass A raw space must be written %20.
a%2 invalid no match pass Incomplete percent escape.
a%zz invalid no match pass Percent must be followed by two hex digits.
a/b invalid no match pass A slash separates segments; it is not part of one.
a?b invalid no match pass "?" starts the query.
a#b invalid no match pass "#" starts the fragment.
café invalid no match pass Non-ASCII must be percent-encoded in a URI.
[x] invalid no match pass Square brackets are gen-delims, not pchar.
(empty string) invalid no match pass Empty: rejected by segment-nz (allowed by plain segment).
.. valid match pass Valid pchar text, but ".." is a dot-segment with special meaning: see the second variant.
a\n invalid no match pass Trailing newline: rejected by JavaScript's $; PCRE2, Python and Java's $ would accept it.

pchar segment that is not "." or ".."

Rule: segment-nz that is not one of the dot-segments "." and ".." (RFC 3986 section 3.3: they are interpreted and removed during reference resolution).

/^(?!\.{1,2}$)(?:[A-Za-z0-9\-._~!$&'()*+,;=:@]|%[0-9A-Fa-f]{2})+$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

(?!\.{1,2}$)
A negative lookahead at the start: fail if the whole segment is one dot or two dots. It consumes nothing.
rest
Same pchar sequence as the first pattern.

Test cases

8 of 9 rows agree with the rule. The other 1 are known limits of this pattern, explained in their notes.

Test cases for pchar segment that is not "." or ".."
Input Rule says Pattern says Result Note
about valid match pass Plain.
. invalid no match pass Dot-segment.
.. invalid no match pass Dot-segment.
... valid match pass Three dots are an ordinary segment, not a dot-segment.
.hidden valid match pass Begins with a dot but is not a dot-segment.
a.. valid match pass Ends with dots.
%2e%2e invalid match false positive KNOWN FALSE POSITIVE for security checks: %2e%2e decodes to "..", a dot-segment, but the pattern only sees the encoded text. Decode first, then test.
a␣b invalid no match pass Space.
(empty string) invalid no match pass Empty.

How it works

RFC 3986 section 3.3 defines segment = *pchar, segment-nz = 1*pchar and pchar = unreserved / pct-encoded / sub-delims / ":" / "@". Section 2.3 gives unreserved = ALPHA / DIGIT / "-" / "." / "_" / "~", and sub-delims are the eleven characters ! $ & ' ( ) * + , ; =. The first pattern writes those out as one character class plus the percent-encoded alternative.

The characters not in pchar are exactly the ones that would end or change the meaning of a segment: "/" starts the next segment, "?" starts the query and "#" starts the fragment. A space, a double quote, a backslash, angle brackets and the square brackets are not allowed either, so they appear percent-encoded.

RFC 3986 also says the path segments "." and ".." are dot-segments, intended for relative references and removed as part of reference resolution (section 5.2.4 describes the algorithm). A segment of just those characters is syntactically fine but behaves differently, which is why the second variant rejects it.

Known false positives

  • Percent-encoded dot-segments such as %2e%2e pass both patterns. Decode the text first if you need to decide on the decoded value.
  • A matching segment may still be meaningless for your application, or may decode to something your code does not expect, such as a slash (%2F) or a NUL byte (%00).

Known false negatives

  • Browsers and servers accept some text that is not a strict RFC 3986 URI (for example, raw non-ASCII letters, which the URL standard percent-encodes for you). The patterns reject those; encode them first with encodeURIComponent, which leaves A-Z a-z 0-9 - _ . ! ~ * ' ( ) unescaped.
  • The empty segment (a double slash in a path) is rejected by segment-nz. If you are matching the RFC's segment production, use * instead of + in the pattern.

Flavor notes

  • Inside the character class the hyphen is escaped and everything else is a literal in all flavors on this site. The percent-encoded alternative uses an explicit [0-9A-Fa-f]{2}.
  • The second variant uses a negative lookahead, which JavaScript and PCRE2 support and Go/RE2 does not (RE2 lists lookaround as not supported). For Go, test for "." and ".." in code and use the first pattern for the rest. The tests run it on JavaScript and PCRE2 only.
  • Trailing newline: PCRE2 (default), Python and Java let $ match before a final newline; JavaScript and Go do not. Use \z, \Z, fullmatch() or matches() in those flavors.

Why not regex here?

If you already hold a URL, do not split it with a regex. The URL interface in browsers parses, normalizes and encodes URLs and gives you the pathname directly (MDN); split that on "/" and decode each piece. A regex is the right tool for checking a single segment you have already isolated, such as a route parameter.

Never use a segment check alone as a path-traversal defence. Decode first, then reject dot-segments, then resolve against a fixed base and verify the result is inside it.

How to use

  1. Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
  2. Check how it is anchored: the patterns use ^ and $ to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases.
  3. If your language is not JavaScript, read Flavor notes and change $ to \z or use a full-match function.
  4. Open it in the tester to see the explanation of each token and try your own inputs.

Worked examples

Validate a route parameter

Use the dot-segment-safe variant on the id part of /items/:id before looking it up.

Encode then check

encodeURIComponent("a b/c") returns a%20b%2Fc; that passes the pattern, whereas the raw string "a b/c" does not.

Limits & gotchas

  • The patterns check characters, not meaning. They do not decode the segment or resolve it against a base.
  • Length is unchecked. Servers and browsers have their own limits on URL length (not read for this page).

FAQ

Why is "@" allowed in a path segment?

RFC 3986 includes "@" and ":" in pchar. They are only restricted in the first segment of a relative-path reference (segment-nz-nc), which this pattern does not model.

Why is "café" rejected?

URIs are ASCII. A non-ASCII character is sent as its UTF-8 bytes percent-encoded, so café becomes caf%C3%A9, which the pattern accepts.

Does encodeURIComponent always produce something this pattern accepts?

Its output uses only the characters A-Z a-z 0-9 - _ . ! ~ * ' ( ) and percent escapes (MDN), all of which are in pchar, so yes, as long as the input is not an unpaired surrogate (which throws).

What is a dot-segment?

A path segment that is "." (current level) or ".." (parent level). RFC 3986 says they are interpreted and removed while resolving a relative reference.

Sources

  1. IETF: RFC 3986: URI Generic Syntax Used for: dec-octet and IPv4address, segment and pchar grammar, pct-encoded, unreserved, sub-delims.
  2. MDN: URL Used for: The URL interface parses, constructs, normalizes and encodes URLs; available in Web Workers.
  3. MDN: encodeURIComponent() Used for: The characters encodeURIComponent leaves unescaped: A-Z a-z 0-9 - _ . ! ~ * ' ( ).
  4. MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
  5. PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
  6. Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
  7. Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
  8. Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
  9. Google RE2: RE2 Syntax (wiki) Used for: RE2 syntax with explicit NOT SUPPORTED markers: lookaround, backreferences, possessive quantifiers, \Z, atomic groups.

Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.