Skip to content

Fix infinite loop and memory exhaustion on an over-long charset name - #34

Merged
alecpl merged 1 commit into
pear:masterfrom
sebastka:fix/rfc2047-encode-header-value-infinite-loop
Aug 2, 2026
Merged

alecpl merged 1 commit into
pear:masterfrom
sebastka:fix/rfc2047-encode-header-value-infinite-loop

Conversation

@sebastka

@sebastka sebastka commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Hello again,

This PR addresses another finding of the LLM, related to #32, which also leads to an infinite loop and eventually to an OOM error on bogus charsets.


Mail_mimePart::encodeHeaderValue() can loop forever and exhaust memory when the charset name is long enough that no room is left for the value.

This is the defect fixed for buildRFC2047Param() in 2001739 (#32), surviving in the header-value path that commit did not touch.

How to reproduce

poc.php
#!/usr/bin/env -S php -d disable_functions=mb_substr,mb_strlen
<?php declare(strict_types=1);
require_once 'Mail/mime.php';

// Run as:
// - ./snippet.php: the shebang disables ext/mbstring, without which encodeMB() bails out early and encodeHeaderValue() is never reached.
// - "php snippet.php" ignores the shebang and, on PHP >= 8.0, shows the ValueError path handled at the bottom instead

// The loop below never terminates, so cap memory and let PHP abort rather than hang the terminal
ini_set('memory_limit', '32M');

// At 65 the budget rounds down to 0, so in the base64 branch substr($value, 0, 0) returns '' and substr($value, 0) returns $value unchanged:
// while ($value) spins forever, appending a prefix and a suffix to $output on every pass
$charset = str_repeat('X', 65);

// The subject must contain non-ASCII, otherwise it is emitted verbatim and never reaches the RFC 2047 encoder at all
$mime = new Mail_mime([
    'eol'           => "\r\n",
    'head_charset'  => $charset,
    'head_encoding' => 'base64',
]);

$mime->setTXTBody('body');
$mime->setSubject('söme välue héré');

printf('building a message with a %d-char head_charset...' . PHP_EOL . PHP_EOL, strlen($charset));

try {
    $mime->get();
    printf('%s' . PHP_EOL, $mime->txtHeaders());
    printf('returned normally' . PHP_EOL);
} catch (Throwable $e) {
    // PHP >= 8.0 with ext/mbstring stops earlier, inside encodeMB()
    printf('%s: %s' . PHP_EOL, get_class($e), $e->getMessage());
}

Before:

building a message with a 65-char head_charset...

PHP Fatal error:  Allowed memory size of 33554432 bytes exhausted (tried to allocate 14680096 bytes) in ./Mail_Mime/Mail/mimePart.php on line 1091
Stack trace:
#0 ./Mail_Mime/Mail/mimePart.php(993): Mail_mimePart::encodeHeaderValue()
#1 ./Mail_Mime/Mail/mime.php(1328): Mail_mimePart::encodeHeader()
#2 ./Mail_Mime/Mail/mime.php(1303): Mail_mime->encodeHeader()
#3 ./Mail_Mime/Mail/mime.php(1096): Mail_mime->encodeHeaders()
#4 ./Mail_Mime/Mail/mime.php(1113): Mail_mime->headers()
#5 ./Mail_Mime/snippet.php(30): Mail_mime->txtHeaders()
#6 {main}

After:

building a message with a 65-char head_charset...

MIME-Version: 1.0
Subject: =?XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX?B?c8O2bWUg?=
 =?XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX?B?dsOkbHVlIGjDqXLDqQ==?=
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

returned normally

Root cause

In encodeHeaderValue():

$maxLength = 75 - strlen($prefix . $suffix);
$maxLength1stLine = $maxLength - $prefix_len;

At 65 charset characters the budget is 3, which base64 rounds down to a multiple of 4, zero. substr($value, 0, 0) returns '' and substr($value, 0) returns the value unchanged, so the loop consumes nothing while appending a prefix and suffix to $output on every pass.

Fix

Apply the same 48-character charset cap used by buildRFC2047Param(), extracted into Mail_mimePart::MAX_CHARSET_LENGTH so the two call sites cannot drift. The longest charset name mbstring recognises is ISO-2022-JP-MOBILE#KDDI at 23 characters, so a valid name is never truncated.

Also clamp the derived chunk lengths. A large $prefix_len could already drive the first-line budget negative independently of the charset, feeding an invalid .{0,-n} pattern to preg_match() in the quoted-printable branch.

Scope

The buggy path is only reached when encodeMB() declines (without ext/mbstring), or with it on PHP < 8.0 where mb_strlen() returns false for an unknown encoding rather than throwing. On PHP ≥ 8.0 with mbstring you get an uncaught ValueError instead of a hang.

Not addressed here: the quoted-printable splitter can still emit encoded-words of up to 77 characters, because (.{0,$maxLength}[^\=][^\=]) matches up to two characters past its bound. That is independent of the charset length: it also happens with a one-character charset name, so it is left for a separate fix.

Test

tests/rfc2047_long_charset.phpt sweeps 64 combinations (both encodings × 8 charset lengths × 4 prefix lengths) plus one end-to-end case through Mail_mime. Two notes:

  1. --INI-- disable_functions=mb_substr,mb_strlen is what makes the test meaningful in CI: without it encodeMB() handles the value and the fixed code never runs.
  2. memory_limit=64M turns a regression into a failing test rather than a hung CI job.

AI use disclosure

I used Anthropic's Claude LLM with Opus 5 to find this bug. It also suggested a fix, which I have reviewed and tested.

encodeHeaderValue() splits a value into chunks of
75 - strlen("=?<charset>?B?" . "?=") characters and never clamped that size.
A charset name of 65 characters or more drove it to zero, so
substr($value, 0, 0) returned '' and substr($value, 0) returned the value
unchanged: the loop consumed nothing while appending an encoded-word prefix
and suffix to the output on every pass, until PHP ran out of memory.

This is the defect fixed for buildRFC2047Param() in 2001739 (pear#32), surviving
in the header-value path that commit did not touch. Apply the same 48-character
charset cap here, and extract it into Mail_mimePart::MAX_CHARSET_LENGTH so the
two call sites cannot drift. The longest charset name mbstring recognises is
23 characters, so a valid name is never truncated.

Additionally clamp the derived chunk lengths. A large $prefix_len could already
drive the first-line budget negative independently of the charset, which fed an
invalid ".{0,-n}" pattern to preg_match() in the quoted-printable branch.

The buggy path is only reached when encodeMB() declines, i.e. without ext/mbstring,
or with it on PHP < 8.0 where mb_strlen() returns false for an unknown encoding
rather than throwing.

Not addressed here: the quoted-printable splitter can still emit encoded-words
of up to 77 characters, because "(.{0,$maxLength}[^\=][^\=])" matches up to two
characters past its bound. That is independent of the charset length (it also
happens with a one-character charset name) so it is left for a separate fix.
@alecpl
alecpl merged commit 4216044 into pear:master Aug 2, 2026
12 checks passed
alecpl pushed a commit that referenced this pull request Aug 4, 2026
…ng (#35)

* Fix quoted-printable encoded-words exceeding the RFC 2047 75-character limit

encodeHeaderValue() cuts quoted-printable text with
"(.{0,$maxLength}[^\=][^\=])". The two trailing atoms each match one further
character, so every chunk could run two characters past its bound: a subject
of "Lorem ipsum ... magna aliqua" plus an ellipsis produced a 77-character
encoded-word on a 78-character line, against RFC 2047 limits of 75 and 76.

The [^\=][^\=] pair is what stops a cut landing inside a "=XX" escape, so the
repetition count is reduced by two rather than the atoms removed. The chunk
length floors added in 4216044 (#34) guarantee the count cannot go negative.

Only the encodeHeaderValue() fallback is affected, i.e. builds without
ext/mbstring; encodeMB() already checked its limit correctly and produces
byte-identical output for the same input after this change.

tests/headers_without_mbstring.phpt pinned the old chunking, so its
expectations are regenerated. Case [07] now emits the same two lines as
headers_with_mbstring.phpt, where the two paths previously disagreed.

* Run headers_without_mbstring.phpt instead of always skipping it

The test skips whenever ext/mbstring is present, which is every job in the CI
matrix, so the non-mbstring header encoding path had no coverage at all. That
is how its expectations came to encode a chunking bug unnoticed.

Add an --INI-- section disabling mb_substr()/mb_strlen(). Unlike the -i
command line flag, --INI-- also applies to --SKIPIF--, so the existing skip
condition now acts as a fallback: the test runs when the functions can be
disabled, and still skips if they cannot.
alecpl pushed a commit that referenced this pull request Aug 5, 2026
…37)

The "Q" character class escaped printable punctuation and every 8-bit byte but
neither \x00-\x1F nor \x7F, so C0 controls and DEL were written into the header
verbatim. RFC 2047 defines encoded-text as printable ASCII other than "?" and
SPACE, so any of them makes the encoded-word invalid, and a raw CR or LF in a
header is a header injection vector.

A literal LF failed in a second way with ext/mbstring. encodeMB() uses "\n" as
its own chunk separator and expands it into a fold when assembling the result,
so an LF in the value was indistinguishable from a chunk boundary: the value
was split into two encoded-words and the byte silently disappeared from the
decoded header.

Add the two missing ranges, which lets \x7B-\x7E, \x7F and \x80-\xFF collapse
into \x7B-\xFF. The class was duplicated verbatim in encodeQP() and encodeMB(),
the second copy carrying a "see encodeQP()" comment, so it moves into
Mail_mimePart::QP_ESCAPE_REGEXP as 4216044 (#34) did for MAX_CHARSET_LENGTH.

tests/headers_with_mbstring.phpt pinned the defect as expected output and is
regenerated. Case [31] encodes a Japanese subject to ISO-2022-JP, whose
charset-switching escapes are ESC, and one of the JIS X 0208 bytes involved is
\x0D: the expectations held ten raw ESC bytes and a bare carriage return inside
a Subject header, invisible unless viewed with cat -v.

Unlike #33, #34 and #35 this reproduces with ext/mbstring present, so
tests/rfc2047_control_chars.phpt needs no --INI-- section. It sweeps all 32 C0
control characters plus DEL across both encodings, checking that the
encoded-text holds only printable ASCII and that the byte survives a round
trip. It reports 34 failures against the previous code.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants