Skip to content

Fix base64 encoded-words exceeding the RFC 2047 75-character limit - #33

Merged
alecpl merged 1 commit into
pear:masterfrom
sebastka:fix/rfc2047-base64-encoded-word-length
Jul 30, 2026
Merged

alecpl merged 1 commit into
pear:masterfrom
sebastka:fix/rfc2047-base64-encoded-word-length

Conversation

@sebastka

Copy link
Copy Markdown
Contributor

Hello again,

While I was at it, I asked the LLM to find other RFC2047 violations and it did find a few. Here is the first one.


RFC 2047 §2 sets two hard limits:

An 'encoded-word' may not be more than 75 characters long, including 'charset', 'encoding', 'encoded-text', and delimiters.
... each line of a header field that contains one or more 'encoded-word's is limited to 76 characters.

Mail_mimePart::encodeMB() breaks both when it splits a base64-encoded header, because the final encoded-word is emitted without a length check.

How to reproduce

poc.php
<?php declare(strict_types=1);
require_once 'Mail/mime.php';

// The trailing "…" (U+2026, 3 bytes in UTF-8) is what makes the value non-ASCII
$lorem = 'Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod '
    . 'tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, '
    . 'quis nostrud exercitation…';

// Using the public Mail_mime API instead of Mail_mimePart::encodeHeader() directly
// to show the defect is reachable from documented usage.
// Also, the body encoding ('text_encoding') is separate and unaffected by this bug.
$mime = new Mail_mime([
    'eol'           => "\r\n",
    'head_charset'  => 'UTF-8',
    'head_encoding' => 'base64',
    'text_charset'  => 'UTF-8',
]);

$mime->setTXTBody($lorem);
$mime->setSubject($lorem);
$mime->headers(['From' => 'sender@example.com', 'To' => 'recipient@example.com']);

$body    = $mime->get();
$message = $mime->txtHeaders() . "\r\n" . $body;

// The whole message, exactly as it goes on the wire (CRLF folding included)
echo $message . PHP_EOL;

Before:

The last Subject line is 77 characters and its encoded-word is 76:

MIME-Version: 1.0
From: sender@example.com
To: recipient@example.com
Subject: =?UTF-8?B?TG9yZW0gaXBzdW0gZG9sb3Igc2l0IGFtZXQsIGNvbnNlY3RldHVy?=
 =?UTF-8?B?IGFkaXBpc2NpbmcgZWxpdCwgc2VkIGRvIGVpdXNtb2QgdGVtcG9yIGluY2lk?=
 =?UTF-8?B?aWR1bnQgdXQgbGFib3JlIGV0IGRvbG9yZSBtYWduYSBhbGlxdWEuIFV0IGVu?=
 =?UTF-8?B?aW0gYWQgbWluaW0gdmVuaWFtLCBxdWlzIG5vc3RydWQgZXhlcmNpdGF0aW9u4oCm?=
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tem=
por incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, q=
uis nostrud exercitation=E2=80=A6

After

The trailing "…" moves into its own encoded-word:

MIME-Version: 1.0
From: sender@example.com
To: recipient@example.com
Subject: =?UTF-8?B?TG9yZW0gaXBzdW0gZG9sb3Igc2l0IGFtZXQsIGNvbnNlY3RldHVy?=
 =?UTF-8?B?IGFkaXBpc2NpbmcgZWxpdCwgc2VkIGRvIGVpdXNtb2QgdGVtcG9yIGluY2lk?=
 =?UTF-8?B?aWR1bnQgdXQgbGFib3JlIGV0IGRvbG9yZSBtYWduYSBhbGlxdWEuIFV0IGVu?=
 =?UTF-8?B?aW0gYWQgbWluaW0gdmVuaWFtLCBxdWlzIG5vc3RydWQgZXhlcmNpdGF0aW9u?=
 =?UTF-8?B?4oCm?=
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tem=
por incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, q=
uis nostrud exercitation=E2=80=A6

Root cause

In the base64 branch of encodeMB():

if ($line_length + $chunk_len == $maxLength || $i == $length) {

$i == $length short-circuits the size check, so the last chunk is flushed whatever its length. Base64 output is always a multiple of 4 while $maxLength is 63 for UTF-8, so the exact-fit arm never matches and all real splitting is done by the > $maxLength arm below, which the final iteration never reaches. The tail therefore lands at 64 payload characters, giving a 76-character encoded-word on a 77-character line.

Fix

Run the overflow check before the end-of-string test: flush the previous (fitting) chunk, then re-cut the remainder onto a new encoded-word. $prev is also reset on both flush paths, which the original left stale after a flush.

Scope

Only base64 header encoding is affected. head_encoding defaults to quoted-printable, and that branch tests its limit with >, so it was already correct.

Test

tests/rfc2047_encoded_word_length.phpt sweeps 1-120 repetitions of 2-, 3- and 4-byte UTF-8 characters (360 subjects), asserting every encoded-word ≤ 75, every line ≤ 76, and a byte-identical base64 round-trip. It fails on master (45 violations).

AI use disclosure

I used Anthropic's Claude LLM with Opus 5 to find the RFC violation. It also suggested a fix, which I have tested and reviewed.

encodeMB() flushed the last base64 chunk without checking its length: the
$i == $length test short-circuited the size check, so the final encoded-word
could reach 64 payload characters instead of 63. Because base64 output is
always a multiple of 4 while $maxLength is 63 for UTF-8, the exact-fit branch
never fired and every split was handled by the overflow branch, which the
last iteration skipped.

The result was a 76-character encoded-word on a 77-character line, e.g. for a
Subject of 42 x U+00E4, exceeding RFC 2047's limits of 75 and 76 respectively.
Only base64 header encoding is affected; the quoted-printable branch checks
its limit correctly.

Run the overflow check before the end-of-string test, flushing the previous
(fitting) chunk and re-cutting the remainder onto a new encoded-word. Also
reset $prev on both flush paths, which the original left stale.
@alecpl
alecpl merged commit cc3a92d into pear:master Jul 30, 2026
12 checks passed
alecpl pushed a commit that referenced this pull request Aug 5, 2026
…37)

The "Q" character class escaped printable punctuation and every 8-bit byte but
neither \x00-\x1F nor \x7F, so C0 controls and DEL were written into the header
verbatim. RFC 2047 defines encoded-text as printable ASCII other than "?" and
SPACE, so any of them makes the encoded-word invalid, and a raw CR or LF in a
header is a header injection vector.

A literal LF failed in a second way with ext/mbstring. encodeMB() uses "\n" as
its own chunk separator and expands it into a fold when assembling the result,
so an LF in the value was indistinguishable from a chunk boundary: the value
was split into two encoded-words and the byte silently disappeared from the
decoded header.

Add the two missing ranges, which lets \x7B-\x7E, \x7F and \x80-\xFF collapse
into \x7B-\xFF. The class was duplicated verbatim in encodeQP() and encodeMB(),
the second copy carrying a "see encodeQP()" comment, so it moves into
Mail_mimePart::QP_ESCAPE_REGEXP as 4216044 (#34) did for MAX_CHARSET_LENGTH.

tests/headers_with_mbstring.phpt pinned the defect as expected output and is
regenerated. Case [31] encodes a Japanese subject to ISO-2022-JP, whose
charset-switching escapes are ESC, and one of the JIS X 0208 bytes involved is
\x0D: the expectations held ten raw ESC bytes and a bare carriage return inside
a Subject header, invisible unless viewed with cat -v.

Unlike #33, #34 and #35 this reproduces with ext/mbstring present, so
tests/rfc2047_control_chars.phpt needs no --INI-- section. It sweeps all 32 C0
control characters plus DEL across both encodings, checking that the
encoded-text holds only printable ASCII and that the byte survives a round
trip. It reports 34 failures against the previous code.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants