Fix infinite loop and memory exhaustion on an over-long charset name - #34
Merged
alecpl merged 1 commit intoAug 2, 2026
Merged
Conversation
encodeHeaderValue() splits a value into chunks of
75 - strlen("=?<charset>?B?" . "?=") characters and never clamped that size.
A charset name of 65 characters or more drove it to zero, so
substr($value, 0, 0) returned '' and substr($value, 0) returned the value
unchanged: the loop consumed nothing while appending an encoded-word prefix
and suffix to the output on every pass, until PHP ran out of memory.
This is the defect fixed for buildRFC2047Param() in 2001739 (pear#32), surviving
in the header-value path that commit did not touch. Apply the same 48-character
charset cap here, and extract it into Mail_mimePart::MAX_CHARSET_LENGTH so the
two call sites cannot drift. The longest charset name mbstring recognises is
23 characters, so a valid name is never truncated.
Additionally clamp the derived chunk lengths. A large $prefix_len could already
drive the first-line budget negative independently of the charset, which fed an
invalid ".{0,-n}" pattern to preg_match() in the quoted-printable branch.
The buggy path is only reached when encodeMB() declines, i.e. without ext/mbstring,
or with it on PHP < 8.0 where mb_strlen() returns false for an unknown encoding
rather than throwing.
Not addressed here: the quoted-printable splitter can still emit encoded-words
of up to 77 characters, because "(.{0,$maxLength}[^\=][^\=])" matches up to two
characters past its bound. That is independent of the charset length (it also
happens with a one-character charset name) so it is left for a separate fix.
alecpl
pushed a commit
that referenced
this pull request
Aug 4, 2026
…ng (#35) * Fix quoted-printable encoded-words exceeding the RFC 2047 75-character limit encodeHeaderValue() cuts quoted-printable text with "(.{0,$maxLength}[^\=][^\=])". The two trailing atoms each match one further character, so every chunk could run two characters past its bound: a subject of "Lorem ipsum ... magna aliqua" plus an ellipsis produced a 77-character encoded-word on a 78-character line, against RFC 2047 limits of 75 and 76. The [^\=][^\=] pair is what stops a cut landing inside a "=XX" escape, so the repetition count is reduced by two rather than the atoms removed. The chunk length floors added in 4216044 (#34) guarantee the count cannot go negative. Only the encodeHeaderValue() fallback is affected, i.e. builds without ext/mbstring; encodeMB() already checked its limit correctly and produces byte-identical output for the same input after this change. tests/headers_without_mbstring.phpt pinned the old chunking, so its expectations are regenerated. Case [07] now emits the same two lines as headers_with_mbstring.phpt, where the two paths previously disagreed. * Run headers_without_mbstring.phpt instead of always skipping it The test skips whenever ext/mbstring is present, which is every job in the CI matrix, so the non-mbstring header encoding path had no coverage at all. That is how its expectations came to encode a chunking bug unnoticed. Add an --INI-- section disabling mb_substr()/mb_strlen(). Unlike the -i command line flag, --INI-- also applies to --SKIPIF--, so the existing skip condition now acts as a fallback: the test runs when the functions can be disabled, and still skips if they cannot.
alecpl
pushed a commit
that referenced
this pull request
Aug 5, 2026
…37) The "Q" character class escaped printable punctuation and every 8-bit byte but neither \x00-\x1F nor \x7F, so C0 controls and DEL were written into the header verbatim. RFC 2047 defines encoded-text as printable ASCII other than "?" and SPACE, so any of them makes the encoded-word invalid, and a raw CR or LF in a header is a header injection vector. A literal LF failed in a second way with ext/mbstring. encodeMB() uses "\n" as its own chunk separator and expands it into a fold when assembling the result, so an LF in the value was indistinguishable from a chunk boundary: the value was split into two encoded-words and the byte silently disappeared from the decoded header. Add the two missing ranges, which lets \x7B-\x7E, \x7F and \x80-\xFF collapse into \x7B-\xFF. The class was duplicated verbatim in encodeQP() and encodeMB(), the second copy carrying a "see encodeQP()" comment, so it moves into Mail_mimePart::QP_ESCAPE_REGEXP as 4216044 (#34) did for MAX_CHARSET_LENGTH. tests/headers_with_mbstring.phpt pinned the defect as expected output and is regenerated. Case [31] encodes a Japanese subject to ISO-2022-JP, whose charset-switching escapes are ESC, and one of the JIS X 0208 bytes involved is \x0D: the expectations held ten raw ESC bytes and a bare carriage return inside a Subject header, invisible unless viewed with cat -v. Unlike #33, #34 and #35 this reproduces with ext/mbstring present, so tests/rfc2047_control_chars.phpt needs no --INI-- section. It sweeps all 32 C0 control characters plus DEL across both encodings, checking that the encoded-text holds only printable ASCII and that the byte survives a round trip. It reports 34 failures against the previous code.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hello again,
This PR addresses another finding of the LLM, related to #32, which also leads to an infinite loop and eventually to an OOM error on bogus charsets.
Mail_mimePart::encodeHeaderValue()can loop forever and exhaust memory when the charset name is long enough that no room is left for the value.This is the defect fixed for
buildRFC2047Param()in 2001739 (#32), surviving in the header-value path that commit did not touch.How to reproduce
poc.php
Before:
After:
Root cause
In
encodeHeaderValue():At 65 charset characters the budget is 3, which base64 rounds down to a multiple of 4, zero.
substr($value, 0, 0) returns ''andsubstr($value, 0)returns the value unchanged, so the loop consumes nothing while appending a prefix and suffix to$outputon every pass.Fix
Apply the same 48-character charset cap used by
buildRFC2047Param(), extracted intoMail_mimePart::MAX_CHARSET_LENGTHso the two call sites cannot drift. The longest charset name mbstring recognises is ISO-2022-JP-MOBILE#KDDI at 23 characters, so a valid name is never truncated.Also clamp the derived chunk lengths. A large
$prefix_lencould already drive the first-line budget negative independently of the charset, feeding an invalid.{0,-n}pattern topreg_match()in the quoted-printable branch.Scope
The buggy path is only reached when
encodeMB()declines (without ext/mbstring), or with it on PHP < 8.0 wheremb_strlen()returnsfalsefor an unknown encoding rather than throwing. On PHP ≥ 8.0 with mbstring you get an uncaughtValueErrorinstead of a hang.Not addressed here: the quoted-printable splitter can still emit encoded-words of up to 77 characters, because
(.{0,$maxLength}[^\=][^\=])matches up to two characters past its bound. That is independent of the charset length: it also happens with a one-character charset name, so it is left for a separate fix.Test
tests/rfc2047_long_charset.phptsweeps 64 combinations (both encodings × 8 charset lengths × 4 prefix lengths) plus one end-to-end case throughMail_mime. Two notes:--INI-- disable_functions=mb_substr,mb_strlenis what makes the test meaningful in CI: without itencodeMB()handles the value and the fixed code never runs.memory_limit=64Mturns a regression into a failing test rather than a hung CI job.AI use disclosure
I used Anthropic's Claude LLM with Opus 5 to find this bug. It also suggested a fix, which I have reviewed and tested.