Regex to trim leading and trailing whitespace
Removing whitespace from both ends of a string is the classic regex task, and one where the built-in is almost always the better tool. This page tests the pattern anyway, shows what \s means in each flavor, and measures the cost.
Using \s (same set as JavaScript trim())
Rule: The result must equal String.prototype.trim(): MDN defines the whitespace removed as white space characters plus line terminators.
/^\s+|\s+$/g Open in replace tester Loads the pattern with sample text; leave the replacement empty to delete every match.
Parts of the pattern
^\s+- One or more whitespace characters at the very start.
|- Or.
\s+$- One or more whitespace characters at the very end. With the g flag both alternatives are found and deleted; replacing every match with an empty string gives the trimmed text.
Test cases
13 of 13 rows agree with the rule.
| Input | Output after deleting matches | Pattern matches? | Result | Note |
|---|---|---|---|---|
␣␣hello␣␣ | hello | yes | pass | Spaces at both ends. |
hello | hello | no | pass | Nothing to trim: the pattern finds no match. |
(empty string) | (empty string) | no | pass | Empty input. |
␣␣␣ | (empty string) | yes | pass | All whitespace: the first alternative removes all of it. |
\t\n␣hello␣world␣\r\n | hello␣world | yes | pass | Tabs and line breaks count as whitespace. |
␣a␣␣b␣ | a␣␣b | yes | pass | Inner spaces are untouched. |
a\n | a | yes | pass | Trailing newline removed. |
\u{200B}hello | \u{200B}hello | no | pass | ZERO WIDTH SPACE (U+200B) is not whitespace in JavaScript, so neither \s nor trim() removes it. |
\u{0085}hello | \u{0085}hello | no | pass | NEXT LINE (U+0085) is not matched by JavaScript's \s (observed in Node 20), and trim() leaves it too. |
\u{00A0}hello\u{00A0} | hello | yes | pass | No-break spaces are whitespace in JavaScript (MDN lists U+00A0). Go, PCRE2 without UCP and Java do not treat them as \s. |
\u{FEFF}hello | hello | yes | pass | The byte-order mark U+FEFF is on MDN's list of JavaScript white space. |
\u{2028}hello\u{2029} | hello | yes | pass | Line separator and paragraph separator are line terminators, covered by \s in JavaScript. |
\u{000B}hello\u{000C} | hello | yes | pass | Vertical tab and form feed. Go's \s does not include the vertical tab. |
Portable ASCII whitespace
Rule: Remove only the ASCII white-space characters U+0009 to U+000D and U+0020. It does NOT always equal trim(): it leaves no-break spaces and other Unicode spaces in place.
/^[\x09-\x0d\x20]+|[\x09-\x0d\x20]+$/g Open in replace tester Loads the pattern with sample text; leave the replacement empty to delete every match.
Parts of the pattern
[\x09-\x0d\x20]- Tab (09), line feed (0A), vertical tab (0B), form feed (0C), carriage return (0D) and space (20), written as code points so the class is identical in every flavor.
rest- Same two anchored alternatives as the \s version.
Test cases
7 of 10 rows agree with the rule. The other 3 are known limits of this pattern, explained in their notes.
| Input | Output after deleting matches | Pattern matches? | Result | Note |
|---|---|---|---|---|
␣␣hello␣␣ | hello | yes | pass | Spaces at both ends. |
\t\n␣hello␣\r\n | hello | yes | pass | ASCII tabs and line breaks. |
\u{000B}hello\u{000C} | hello | yes | pass | Vertical tab and form feed are in U+0009 to U+000D. |
␣␣␣ | (empty string) | yes | pass | All whitespace. |
(empty string) | (empty string) | no | pass | Empty input. |
␣a␣b␣ | a␣b | yes | pass | Inner space untouched. |
\u{00A0}hello\u{00A0} | \u{00A0}hello\u{00A0} | no | differs from rule | DIFFERS FROM trim(): no-break spaces stay. This is intended for data where NBSP is meaningful. |
\u{2003}hello | \u{2003}hello | no | differs from rule | DIFFERS FROM trim(): EM SPACE (a Unicode Space_Separator, which MDN lists as JavaScript white space) stays. |
\u{FEFF}hello | \u{FEFF}hello | no | differs from rule | DIFFERS FROM trim(): the byte-order mark stays. |
\u{0085}hello | \u{0085}hello | no | pass | Same as trim(): U+0085 is not JavaScript whitespace either. |
How it works
The pattern has two anchored alternatives, one at the start of the string and one at the end, and the g flag makes the engine find both. Replacing every match with an empty string removes exactly the whitespace at the two ends and leaves the middle alone. The test table here is of that kind: each row shows the text after the deletion.
Which characters count as whitespace is the whole difference between flavors. In JavaScript, MDN lists U+0009, U+000B, U+000C, U+0020, U+00A0, U+FEFF and any other Unicode Space_Separator character as white space, and \s covers white space plus line terminators; trim() removes the same set. In Go, \s is [\t\n\f\r ] only. PCRE2 without the UCP option never matches characters above 127 with \s. Java's \s is [ \t\n\x0B\f\r] unless UNICODE_CHARACTER_CLASS is set. Python 3 str patterns match Unicode whitespace. Those are the documented rules; the page cross-checks the ASCII-only cases on the real PCRE2 and Go engines.
The second variant avoids the whole question by naming the characters. It gives the same answer in every flavor, and the table marks the rows where that answer differs from trim().
Known false positives
- Both patterns delete whitespace the text may need: Markdown relies on two trailing spaces for a line break, and fixed-width data may use padding on purpose.
- The ASCII variant leaves a no-break space or a byte-order mark at the start of a string. If your input comes from a file, a leading BOM is a classic cause of "this key looks identical but does not match". The \s version removes it in JavaScript.
Known false negatives
- Zero-width space (U+200B) and next-line (U+0085) are not trimmed by either pattern or by trim() in JavaScript (the table has rows for both). Characters that look blank but are not whitespace need an explicit class.
- In flavors where \s is ASCII only (Go; PCRE2 without UCP; Java by default) the \s version leaves no-break spaces and other Unicode spaces behind.
Flavor notes
- \s is the main difference: JavaScript white space plus line terminators; Go ASCII only [\t\n\f\r ]; PCRE2 ASCII unless the UCP option is set (flag U on this site); Python Unicode for str patterns; Java ASCII unless UNICODE_CHARACTER_CLASS. The ASCII variant sidesteps the difference.
- Alternatives with ^ and $ and the g flag work in JavaScript, PCRE2 and Go. Without the multiline flag, ^ and $ are the ends of the whole input in JavaScript and Go. In PCRE2 (default), Python and Java, $ also matches before a final newline, but because the alternative \s+$ already consumes trailing newlines, the result is the same here.
- Python and Java are notes only on this site: nothing is executed on them.
Why not regex here?
Use the built-in. String.prototype.trim() removes whitespace from both ends (MDN), trimStart() and trimEnd() do one end. They are shorter and immune to the problem below. Other languages have their own trim functions; check the documentation of yours (only the JavaScript ones were read for this page).
Performance: \s+$ is quadratic on a long run of whitespace that is NOT at the end of the string. The engine tries each starting position in the run and scans to the run's end before discovering that no end-of-string follows. Measured in Node 20 on this box: 10,000 spaces then "x" took about 68 ms, 20,000 took about 269 ms, 40,000 about 1,053 ms. Doubling the run quadrupled the time. trim() has no such case.
How to use
- Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
- Check how it is anchored: the patterns use
^and$to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases. - If your language is not JavaScript, read Flavor notes and change
$to\zor use a full-match function. - Open it in the tester to see the explanation of each token and try your own inputs.
Worked examples
Normalise form input
Call value.trim() on the field. Only reach for the regex if you need to trim a different character set, for example just spaces and not tabs: replace ^ +| +$ with an empty string.
Collapse inner runs too
Replace \s+ with a single space after trimming to collapse repeated whitespace inside the string. Try it in the replace tester.
Limits & gotchas
- The patterns remove whitespace, not "invisible" characters in general (zero-width spaces, soft hyphens, bidirectional marks).
- The quadratic behavior is real but only matters for very long runs of whitespace in untrusted input. Limit input length first.
FAQ
Is \s the same as trim()?
In JavaScript they cover the same set: white space plus line terminators. In Go, and in PCRE2 and Java by default, \s is ASCII only, so a no-break space is not removed there.
Why is /\s+$/ slow on some strings?
On a long run of whitespace followed by a non-space, the engine tries the pattern from every position in the run and scans to its end each time, so the total work grows with the square of the run length. Our measurements in Node 20 show it: doubling the run roughly quadruples the time.
Should I use ^\s+|\s+$ or two replaces?
They behave the same. The combined form needs the g flag so that both ends are found.
What about trimming only one end?
Use trimStart() or trimEnd() in JavaScript, or just the relevant alternative: ^\s+ for the start and \s+$ for the end.
Sources
- MDN: String.prototype.trim() Used for: trim() removes whitespace and line terminators from both ends.
- MDN: Whitespace (glossary) Used for: JavaScript white space: tab, vertical tab, form feed, space, no-break space, U+FEFF and other Space_Separator code points.
- MDN: Character class escape: \d, \D, \w, \W, \s, \S Used for: \d is [0-9]; \w is letters, digits and underscore; \s is whitespace plus line terminators.
- PCRE2: pcre2pattern Used for: Syntax and semantics: groups, named groups, lookbehind rules, atomic groups, possessive quantifiers, \d \s \w with and without UCP, dollar and newline handling, \A \Z \z.
- Go: regexp/syntax Used for: Full syntax table; no lookaround or backreferences; \d \s \w are ASCII-only; $ is \z; (?P<name>) and (?<name>); repetition limit 1000; \A and \z.
- Google RE2: RE2 Syntax (wiki) Used for: RE2 syntax with explicit NOT SUPPORTED markers: lookaround, backreferences, possessive quantifiers, \Z, atomic groups.
- Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.
- Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
- MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.