Regex for a double-quoted string with escapes

Matching "text with \"escaped\" quotes" looks easy and is easy to get wrong. This page tests a pattern that follows the JSON string grammar (RFC 8259) and shows the popular shortcut next to it, with the inputs that break the shortcut.

JSON string (RFC 8259), unrolled

Rule: RFC 8259 string: a quotation mark, then any number of characters that are either unescaped (anything except the quotation mark, the backslash and the control characters U+0000 to U+001F) or one of the escape sequences \" \\ \/ \b \f \n \r \t \uXXXX, then a closing quotation mark.

/^"(?:[^"\\\x00-\x1f]|\\(?:["\\/bfnrt]|u[0-9a-fA-F]{4}))*"$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

^" … "$
An opening quote at the start and a closing quote at the very end.
[^"\\\x00-\x1f]
One unescaped character: not a quote, not a backslash, not a control character U+0000 to U+001F (the characters RFC 8259 says MUST be escaped).
\\(?:["\\/bfnrt]|u[0-9a-fA-F]{4})
A backslash followed by one of the eight single-letter escapes, or u and exactly four hexadecimal digits.
(?: A | B )*
Zero or more of those two kinds of item. The two alternatives can never start with the same character (one starts with a backslash, the other cannot), so at every point there is only one way forward and the engine never has to retry.

Test cases

17 of 17 rows agree with the rule.

Test cases for JSON string (RFC 8259), unrolled
Input Rule says Pattern says Result Note
"" valid match pass The empty string.
"hello" valid match pass Plain.
"a\"b" valid match pass Escaped quote.
"back\\slash" valid match pass Escaped backslash.
"tab\there" valid match pass Backslash-t escape.
"\u00e9" valid match pass Unicode escape with four hex digits.
"a\/b" valid match pass Escaped solidus is allowed by RFC 8259.
"café␣😀" valid match pass Non-ASCII characters need no escape.
hello invalid no match pass No quotes.
"unterminated invalid no match pass No closing quote.
"a"b" invalid no match pass An unescaped quote in the middle.
"bad\x" invalid no match pass \x is not a JSON escape.
"\u12" invalid no match pass \u needs four hex digits.
"line\nbreak" invalid no match pass A raw line feed: control characters must be escaped in JSON.
"ends␣with␣backslash\" invalid no match pass The final backslash escapes the closing quote, so the string never ends.
'single' invalid no match pass Single quotes are not JSON.
"\ud800" valid match pass KNOWN LIMIT: RFC 8259 allows this text but warns that an unpaired surrogate escape gives unpredictable results in receiving software. The pattern checks syntax, not meaning.

Naive shortcut: "(?:[^"]|\.)*"

Rule: Same JSON string rule as above (this variant is shown to demonstrate what the shortcut gets wrong).

/^"(?:[^"]|\\.)*"$/

Open in tester Loads the pattern with sample text, JavaScript flavor.

Parts of the pattern

[^"]
Any character except a quote. This includes the backslash and control characters, which is the source of the false positives.
\\.
A backslash and any one character (but not a line terminator, since . does not match one without the s flag). This is the second alternative.
(?:[^"]|\\.)*
The two alternatives overlap: a backslash can be consumed by either. When the overall match fails, the engine tries every way of splitting the backslashes between them, which grows exponentially. See the backtracking page.

Test cases

7 of 11 rows agree with the rule. The other 4 are known limits of this pattern, explained in their notes.

Test cases for Naive shortcut: "(?:[^"]|\.)*"
Input Rule says Pattern says Result Note
"" valid match pass Empty string.
"hello" valid match pass Plain.
"a\"b" valid match pass Escaped quote.
"back\\slash" valid match pass Escaped backslash.
hello invalid no match pass No quotes.
"unterminated invalid no match pass No closing quote.
"a"b" invalid no match pass An unescaped quote in the middle.
"bad\x" invalid match false positive KNOWN FALSE POSITIVE: \x is not a JSON escape but "backslash plus any character" accepts it.
"\u12" invalid match false positive KNOWN FALSE POSITIVE: a malformed \u escape is accepted.
"line\nbreak" invalid match false positive KNOWN FALSE POSITIVE: a raw line feed is matched by [^"].
"ends␣with␣backslash\" invalid match false positive KNOWN FALSE POSITIVE: the engine can read the backslash as an ordinary character via [^"], and then the final quote closes the string. A correct reader sees an escaped quote and an unterminated string.

How it works

RFC 8259 section 7 defines string = quotation-mark *char quotation-mark, where a char is an unescaped character (U+0020-21, U+0023-5B, U+005D-10FFFF) or an escape: a backslash followed by one of the characters " \ / b f n r t, or u and four hexadecimal digits. The first pattern is that grammar, written as a loop.

The "unrolled loop" shape matters. The pattern repeats (normal-character | escape-sequence) where the two alternatives start with different characters: a backslash starts only the escape, and the normal class excludes the backslash. With no overlap, the engine has exactly one choice at each position, so a failing match fails in time proportional to the input, not exponential in it.

The naive pattern lets the backslash be consumed by [^"] and by \\. . Every run of n backslashes can be divided between the two alternatives in many ways. Our tests measure this with the real PCRE2 engine's own match counter, on a string that starts with a quote, then n backslashes, then a quote and a stray x. For n = 4, 8, 12, 16 and 20 backslashes the naive pattern takes 38, 262, 1,796, 12,310 and 84,374 steps (about 7 times more for every 4 added), while the unrolled pattern needs 10, 16, 22, 28 and 34 (6 more per extra backslash). The tests assert that growth shape, not the exact numbers, and the backtracking page shows the effect live.

Known false positives

  • The JSON pattern accepts an escaped unpaired surrogate such as "\ud800". RFC 8259 notes that the behavior of software receiving such values is unpredictable. The pattern checks the text grammar, not what the string means.
  • The naive pattern accepts invalid escapes, malformed \u sequences, raw control characters, and strings whose closing quote is escaped. They are all listed in its table.

Known false negatives

  • Single-quoted strings, JavaScript template literals and strings with a raw line break are rejected: they are not JSON strings. Most programming languages use similar but different rules (for example, they allow \x escapes); change the escape alternative to the grammar of your language before reusing the pattern.
  • This pattern only matches a whole string. To find strings inside a larger text, remove the anchors and add the g flag, but note that a quote preceded by a backslash outside any string would confuse it.

Flavor notes

  • Both patterns use no lookaround or backreferences, so they compile in JavaScript, PCRE2 and Go. Python and Java accept this syntax too, but nothing is executed on them here.
  • In a Python raw string or a Java string literal the backslashes need quoting: this page shows the pattern as the regex engine sees it, not as you would type it in source code.
  • Trailing newline: PCRE2 (default), Python and Java let $ match before a final newline, so a quoted string followed by a line feed would pass there. JavaScript and Go's $ do not. The cross-engine tests use \z for PCRE2.
  • The unrolled pattern is safe on a backtracking engine because it has no overlapping alternatives. In Go (RE2 syntax) even the naive pattern runs in linear time, because the engine does not backtrack; this site's Go engine is a real Go regexp, and the "linear time" claim comes from the Go documentation.

Why not regex here?

If the text is JSON, parse it: JSON.parse reports malformed strings and returns the decoded value, which a regex cannot do (it only tells you the text is well-formed). Use the regex when you need to find or highlight string literals in a language you are not parsing, such as in a syntax highlighter.

Regular expressions cannot count nesting, so a quoted string inside a quoted string (as in JSON embedded in JSON) needs a parser or a second pass.

How to use

  1. Copy the pattern for the variant that fits your rule. Variants differ in strictness, and the rule line says exactly what each accepts.
  2. Check how it is anchored: the patterns use ^ and $ to test a whole string. To find the same thing inside longer text, remove the anchors (and add word-boundary or lookaround checks) and re-run the cases.
  3. If your language is not JavaScript, read Flavor notes and change $ to \z or use a full-match function.
  4. Open it in the tester to see the explanation of each token and try your own inputs.

Worked examples

Watch the false positive

Open the naive pattern in the tester with the text "ends with backslash\" and see the whole line highlighted; then switch to the JSON pattern and see no match.

Find strings in a log

Remove the anchors from the JSON pattern, add the g flag, and run it on a log line with several quoted fields to highlight each one.

Limits & gotchas

  • The pattern knows nothing about the content of the string, only its escapes.
  • The JSON pattern allows U+007F (DEL) and any character above U+001F unescaped, because RFC 8259 only requires control characters U+0000 to U+001F to be escaped.

FAQ

Why not just use "[^"]*"?

It cannot represent an escaped quote: "a\"b" would match as "a\" and stop. If your strings never contain escapes, "[^"]*" is correct and the simplest.

Why does the naive pattern accept "ends with backslash\"?

Because [^"] also matches the backslash, the engine can treat the final backslash as a plain character and let the last quote close the string. Excluding the backslash from the normal-character class forces every backslash to start an escape.

Is this pattern safe against ReDoS?

The JSON pattern has no overlapping alternatives inside the loop, and the tests measure its step count growing in a straight line. That is evidence for these inputs and this engine, not a proof for every engine; test your own flavor.

Can I use it in a language that allows single quotes?

Change both quote characters, and change the list of allowed escapes to match your language. The structure (normal-class excluding the quote and backslash, or a backslash escape) stays the same.

Sources

  1. IETF: RFC 8259: JSON Used for: String grammar: unescaped characters and escape sequences.
  2. OWASP: Regular expression Denial of Service - ReDoS Used for: Explanation of backtracking, "evil" patterns such as (a+)+$, and the doubling of paths per extra character.
  3. PCRE2: pcre2api Used for: Match limit (PCRE2_ERROR_MATCHLIMIT), default newline, DOLLAR_ENDONLY, pcre2_substitute replacement syntax ($1, ${1}, $<name>, $0 or $&).
  4. Go: regexp package Used for: Linear-time guarantee, leftmost-first semantics, Expand template syntax ($1, ${name}, $name longest-name rule).
  5. MDN: Wildcard: . Used for: . excludes line terminators unless the s flag is set; code units vs code points with the u flag.
  6. MDN: Input boundary assertion: ^, $ Used for: ^ and $ are the start and end of input, or of each line with the m flag.
  7. Python docs: re: Regular expression operations (3.13) Used for: Syntax, (?P<name>), atomic groups and possessive quantifiers (3.11+), fixed-length lookbehind, Unicode \d \s \w, $ before trailing newline, \A \Z, re.sub replacement syntax, inline flags at start only (3.11+).
  8. Oracle (Java SE 21): java.util.regex.Pattern Used for: Construct table, \d \s \w without UNICODE_CHARACTER_CLASS, possessive and atomic constructs, named groups, line terminators and $, \A \Z \z.

Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.