Regex for a URL slug (lowercase words and hyphens)
A slug is the readable last part of a web address. No standard defines one, so this page states its rule openly and tests two versions of it: lowercase ASCII, and lowercase Unicode letters and digits.
ASCII slug
Rule: This page's convention (no standard defines it): one or more groups of lowercase ASCII letters and digits, separated by single hyphens, with no hyphen at either end.
/^[a-z0-9]+(?:-[a-z0-9]+)*$/ Open in tester Loads the pattern with sample text, JavaScript flavor.
Parts of the pattern
[a-z0-9]+- The first word: one or more lowercase ASCII letters or digits.
(?:-[a-z0-9]+)*- Zero or more further words, each introduced by exactly one hyphen and containing at least one character. This is what rules out "a--b", "-a" and "a-".
^ … $- The whole string is the slug.
Test cases
14 of 14 rows agree with the rule.
| Input | Rule says | Pattern says | Result | Note |
|---|---|---|---|---|
hello-world | valid | match | pass | Two words. |
a | valid | match | pass | A single character. |
2026-roadmap | valid | match | pass | Digits are allowed, including at the start. |
regex-cheat-sheet-2026 | valid | match | pass | Several words. |
hello--world | invalid | no match | pass | Doubled hyphen. |
-hello | invalid | no match | pass | Leading hyphen. |
hello- | invalid | no match | pass | Trailing hyphen. |
Hello-World | invalid | no match | pass | Upper case. |
hello␣world | invalid | no match | pass | Space. |
hello_world | invalid | no match | pass | Underscore is not in this convention. |
(empty string) | invalid | no match | pass | An empty slug is not a slug. |
café | invalid | no match | pass | Non-ASCII letter: rejected by the ASCII version, accepted by the Unicode one. |
hello-world\n | invalid | no match | pass | Trailing newline: JavaScript's $ rejects it; PCRE2, Python and Java's $ would accept it. |
a-b-c | valid | match | pass | Single-character words. |
Unicode slug (lowercase letters and digits of any script)
Rule: Same as the ASCII convention, but words may contain any lowercase letter (Unicode category Ll) or decimal digit (Nd).
/^[\p{Ll}\p{Nd}]+(?:-[\p{Ll}\p{Nd}]+)*$/u Open in tester Loads the pattern with sample text, JavaScript flavor.
Parts of the pattern
[\p{Ll}\p{Nd}]- A lowercase letter or decimal digit from any script. \p{...} needs the u flag in JavaScript; without it the sequence is an identity escape for the letter p (MDN).
rest- Same structure as the ASCII pattern: words separated by single hyphens.
Test cases
10 of 10 rows agree with the rule.
| Input | Rule says | Pattern says | Result | Note |
|---|---|---|---|---|
café | valid | match | pass | é (U+00E9) is a lowercase letter. |
hello-wörld | valid | match | pass | ö is a lowercase letter. |
привет-мир | valid | match | pass | Lowercase Cyrillic. |
hello-world-2026 | valid | match | pass | ASCII still works. |
CAFÉ | invalid | no match | pass | Upper case. |
Café | invalid | no match | pass | Capital C. |
hello--world | invalid | no match | pass | Doubled hyphen. |
café | invalid | no match | pass | KNOWN LIMIT: "e" followed by U+0301 (combining acute) is the decomposed spelling of é. The combining mark is category Mn, not Ll, so it is rejected. Convert the text to its composed (NFC) form before matching. |
日本語 | invalid | no match | pass | CJK ideographs are category Lo (other letter), not Lowercase Letter, so the pattern rejects them even though a slug in Japanese is reasonable. Use \p{L} if you want to accept them and handle case another way. |
hello␣world | invalid | no match | pass | Space. |
How it works
No specification defines "slug". MDN's glossary describes it as the unique identifying part of a web address, typically at the end of the URL, so the pattern encodes a convention instead: words made of lowercase letters and digits, separated by single hyphens.
The structure is the important part. A first word, then zero or more groups of a hyphen followed by another word. Because each hyphen is bound to a following word, a hyphen can never be first, last or doubled, with no lookahead needed. The alternative that people often write, ^[a-z0-9-]+$, accepts all of those mistakes.
The Unicode version swaps the class [a-z0-9] for [\p{Ll}\p{Nd}]. In JavaScript the u flag is required for \p; PCRE2 supports \p when UTF mode is on (this site always turns it on) and Go supports it natively.
Known false positives
- A slug that matches is only well-formed. It can still be misleading, a duplicate of another slug, or a reserved route in your application (admin, api).
- Both patterns accept a slug consisting only of digits, such as 2026. Many routers would treat that as an ID.
Known false negatives
- Underscores, dots and upper-case letters are rejected because this convention does not include them. If your platform allows underscores, add _ to the class (and then consider whether a-_-b should be allowed).
- The Unicode version rejects scripts without case (Chinese, Japanese, Arabic and so on use categories such as Lo, which is not Ll), and rejects decomposed accents. Both are shown in the table.
Flavor notes
- The ASCII pattern is portable to every flavor on this site, with the usual trailing-newline caveat: JavaScript and Go's $ match only at the very end, while PCRE2 (default), Python and Java also match before a final newline. Use \z, \Z, fullmatch() or matches() when validating.
- The Unicode pattern: JavaScript needs the u flag; PCRE2 needs UTF mode (always on here); Go accepts \p{Ll} directly. Python's re module documentation does not list \p (observed on 3.13.5: \p{L} is a "bad escape"), so in Python you cannot port this pattern as written; use str.islower() and str.isdigit() per character, or the third-party regex module (not read for this page).
Why not regex here?
Do not use a regex to generate a slug from a title. Lower-casing, transliteration (é to e), collapsing punctuation and truncation all matter, and a regex substitution can do each step but cannot decide what your site wants. Use the regex only to check a slug that already exists.
To percent-encode arbitrary text for a path, use encodeURIComponent (MDN lists the characters it leaves unescaped). That is for path segments, and is different from a slug.
How to use
- Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
- Check how it is anchored: the patterns use
^and$to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases. - If your language is not JavaScript, read Flavor notes and change
$to\zor use a full-match function. - Open it in the tester to see the explanation of each token and try your own inputs.
Worked examples
Check a slug in a form
Use the ASCII pattern on the field and show "lowercase letters, digits and single hyphens only" when it fails. The tester highlights exactly which part fails if you paste the slug with the pattern.
Normalise before checking
Convert the text to its composed Unicode form (NFC) and lower-case it before the Unicode pattern, so decomposed accents become single letters. The table shows the decomposed "café" being rejected.
Limits & gotchas
- The rule is a convention. If your CMS has a different rule (allowing underscores, or limiting length), change the pattern and re-run the table.
- Length is not checked. Add a lookahead such as (?=.{1,60}$) in JavaScript or PCRE2 for a length limit; Go has no lookahead, so check the length in code there.
FAQ
Why not use ^[a-z0-9-]+$?
It accepts "-", "--", "-a", "a-" and "a--b". The structure word(-word)* makes every hyphen sit between two words.
Is there a standard for slugs?
Not one that this page found. MDN's glossary defines a slug as the identifying part of a web address; the character rules are a convention of each site or framework.
Why does the Unicode variant reject Japanese?
Japanese ideographs are in the Unicode category Lo (other letter), which has no case and so is not "Ll". Use \p{L} instead of \p{Ll} if you want to accept them, and decide separately how to treat case.
Do I need the u flag?
Only for the Unicode variant. In JavaScript \p{...} is an identity escape (it matches the letter p) without u or v, so the pattern would silently mean something else.
Sources
- MDN: Slug (glossary) Used for: A slug is the identifying part of a web address, typically at the end of the URL.
- MDN: Unicode character class escape: \p{...}, \P{...} Used for: \p and \P need the u or v flag.
- MDN: encodeURIComponent() Used for: The characters encodeURIComponent leaves unescaped: A-Z a-z 0-9 - _ . ! ~ * ' ( ).
- MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
- PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
- Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
- Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
- Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
- MDN: Regular expressions (reference) Used for: List of flags (d g i m s u v y), the groups of syntax, and which syntaxes are assertions.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.