Fix base64 encoded-words exceeding the RFC 2047 75-character limit - #33
Merged
alecpl merged 1 commit intoJul 30, 2026
Merged
Conversation
encodeMB() flushed the last base64 chunk without checking its length: the $i == $length test short-circuited the size check, so the final encoded-word could reach 64 payload characters instead of 63. Because base64 output is always a multiple of 4 while $maxLength is 63 for UTF-8, the exact-fit branch never fired and every split was handled by the overflow branch, which the last iteration skipped. The result was a 76-character encoded-word on a 77-character line, e.g. for a Subject of 42 x U+00E4, exceeding RFC 2047's limits of 75 and 76 respectively. Only base64 header encoding is affected; the quoted-printable branch checks its limit correctly. Run the overflow check before the end-of-string test, flushing the previous (fitting) chunk and re-cutting the remainder onto a new encoded-word. Also reset $prev on both flush paths, which the original left stale.
alecpl
pushed a commit
that referenced
this pull request
Aug 5, 2026
…37) The "Q" character class escaped printable punctuation and every 8-bit byte but neither \x00-\x1F nor \x7F, so C0 controls and DEL were written into the header verbatim. RFC 2047 defines encoded-text as printable ASCII other than "?" and SPACE, so any of them makes the encoded-word invalid, and a raw CR or LF in a header is a header injection vector. A literal LF failed in a second way with ext/mbstring. encodeMB() uses "\n" as its own chunk separator and expands it into a fold when assembling the result, so an LF in the value was indistinguishable from a chunk boundary: the value was split into two encoded-words and the byte silently disappeared from the decoded header. Add the two missing ranges, which lets \x7B-\x7E, \x7F and \x80-\xFF collapse into \x7B-\xFF. The class was duplicated verbatim in encodeQP() and encodeMB(), the second copy carrying a "see encodeQP()" comment, so it moves into Mail_mimePart::QP_ESCAPE_REGEXP as 4216044 (#34) did for MAX_CHARSET_LENGTH. tests/headers_with_mbstring.phpt pinned the defect as expected output and is regenerated. Case [31] encodes a Japanese subject to ISO-2022-JP, whose charset-switching escapes are ESC, and one of the JIS X 0208 bytes involved is \x0D: the expectations held ten raw ESC bytes and a bare carriage return inside a Subject header, invisible unless viewed with cat -v. Unlike #33, #34 and #35 this reproduces with ext/mbstring present, so tests/rfc2047_control_chars.phpt needs no --INI-- section. It sweeps all 32 C0 control characters plus DEL across both encodings, checking that the encoded-text holds only printable ASCII and that the byte survives a round trip. It reports 34 failures against the previous code.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hello again,
While I was at it, I asked the LLM to find other RFC2047 violations and it did find a few. Here is the first one.
RFC 2047 §2 sets two hard limits:
Mail_mimePart::encodeMB()breaks both when it splits a base64-encoded header, because the final encoded-word is emitted without a length check.How to reproduce
poc.php
Before:
The last Subject line is 77 characters and its encoded-word is 76:
After
The trailing "…" moves into its own encoded-word:
Root cause
In the base64 branch of encodeMB():
$i == $lengthshort-circuits the size check, so the last chunk is flushed whatever its length. Base64 output is always a multiple of 4 while$maxLengthis 63 for UTF-8, so the exact-fit arm never matches and all real splitting is done by the> $maxLengtharm below, which the final iteration never reaches. The tail therefore lands at 64 payload characters, giving a 76-character encoded-word on a 77-character line.Fix
Run the overflow check before the end-of-string test: flush the previous (fitting) chunk, then re-cut the remainder onto a new encoded-word.
$previs also reset on both flush paths, which the original left stale after a flush.Scope
Only base64 header encoding is affected.
head_encodingdefaults toquoted-printable, and that branch tests its limit with>, so it was already correct.Test
tests/rfc2047_encoded_word_length.phptsweeps 1-120 repetitions of 2-, 3- and 4-byte UTF-8 characters (360 subjects), asserting every encoded-word ≤ 75, every line ≤ 76, and a byte-identical base64 round-trip. It fails onmaster(45 violations).AI use disclosure
I used Anthropic's Claude LLM with Opus 5 to find the RFC violation. It also suggested a fix, which I have tested and reviewed.