Skip to content

Fix/rfc2047 qp chunk off by two - #35

Merged
alecpl merged 2 commits into
pear:masterfrom
sebastka:fix/rfc2047-qp-chunk-off-by-two
Aug 4, 2026
Merged

alecpl merged 2 commits into
pear:masterfrom
sebastka:fix/rfc2047-qp-chunk-off-by-two

Conversation

@sebastka

@sebastka sebastka commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Hello again,

This PR addresses another finding of the LLM that was mentioned in #34: encoded-words that overshoot the length limits set by RFC 2047.


Mail_mimePart::encodeHeaderValue() cuts quoted-printable text two characters later than it should, emitting encoded-words of up to 77 characters on lines of up to 78. RFC 2047 allows 75 and 76 respectively.

How to reproduce

poc.php
#!/usr/bin/env -S php -d disable_functions=mb_substr,mb_strlen
<?php declare(strict_types=1);
require_once 'Mail/mime.php';

// Run as ./snippet.php to disable mb_string, since encodeMB() splits correctly:
// the off-by-two lives in the encodeHeaderValue() fallback that only runs when mbstring is unavailable.

// The trailing "…" (U+2026, 3 bytes in UTF-8) is what makes the value non-ASCII
$lorem = 'Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod '
    . 'tempor incididunt ut labore et dolore magna aliqua…';

$mime = new Mail_mime([
    'eol'           => "\r\n",
    'head_charset'  => 'UTF-8',
    'head_encoding' => 'quoted-printable',
]);

$mime->setTXTBody('body');
$mime->setSubject($lorem);
$mime->get();

$headers = $mime->headers();
$header  = 'Subject: ' . $headers['Subject'];

// RFC 2047 chunks the value with "(.{0,$maxLength}[^\=][^\=])".
// The two trailing atoms each match one further character, so every chunk may run two characters past its bound.

printf('%s' . PHP_EOL, $mime->txtHeaders());

// RFC 2047: "each line of a header field that contains one or more 'encoded-word's is limited to 76 characters"
foreach (explode("\r\n", $header) as $line)
    printf('line %3d | %s%s' . PHP_EOL, strlen($line), $line, (strlen($line) > 76 ? '  <-- exceeds the RFC 2047 limit of 76' : ''));

printf(PHP_EOL);

// RFC 2047: "An 'encoded-word' may not be more than 75 characters long"
preg_match_all('/=\?[^?]+\?Q\?[^?]*\?=/', $header, $matches);

foreach ($matches[0] as $word)
    printf('word %3d | %s%s' . PHP_EOL, strlen($word), $word, (strlen($word) > 75 ? '  <-- exceeds the RFC 2047 limit of 75' : ''));

Before

MIME-Version: 1.0
Subject: =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_eli?=
 =?UTF-8?Q?t=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_magna_a?=
 =?UTF-8?Q?liqua=E2=80=A6?=
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

line  77 | Subject: =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_eli?=  <-- exceeds the RFC 2047 limit of 76
line  78 |  =?UTF-8?Q?t=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_magna_a?=  <-- exceeds the RFC 2047 limit of 76
line  27 |  =?UTF-8?Q?liqua=E2=80=A6?=

word  68 | =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_eli?=
word  77 | =?UTF-8?Q?t=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_magna_a?=  <-- exceeds the RFC 2047 limit of 75
word  26 | =?UTF-8?Q?liqua=E2=80=A6?=

After

MIME-Version: 1.0
Subject: =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_e?=
 =?UTF-8?Q?lit=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_mag?=
 =?UTF-8?Q?na_aliqua=E2=80=A6?=
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

line  75 | Subject: =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_e?=
line  76 |  =?UTF-8?Q?lit=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_mag?=
line  31 |  =?UTF-8?Q?na_aliqua=E2=80=A6?=

word  66 | =?UTF-8?Q?Lorem_ipsum_dolor_sit_amet=2C_consectetur_adipiscing_e?=
word  75 | =?UTF-8?Q?lit=2C_sed_do_eiusmod_tempor_incididunt_ut_labore_et_dolore_mag?=
word  30 | =?UTF-8?Q?na_aliqua=E2=80=A6?=

This is byte-identical to what encodeMB() already produced for the same input, so the new chunking is not an arbitrary choice: the two code paths now agree.

Root cause

In encodeHeaderValue():

$reg1st = "|(.{0,$maxLength1stLine}[^\=][^\=])|";
$reg2nd = "|(.{0,$maxLength}[^\=][^\=])|";

The repetition is bounded at $maxLength, but the two trailing atoms each match one further character, so the whole match can run to $maxLength + 2.

Fix

The [^\=][^\=] pair is what stops a cut landing inside a =XX escape, so it has to stay. The repetition count is reduced by two instead. The chunk length floors added in 4216044 (#34) guarantee the count cannot go negative.

Scope

Only the encodeHeaderValue() fallback is affected, i.e. builds without ext/mbstring. encodeMB() already checked its limit correctly.

tests/headers_without_mbstring.phpt pinned the old chunking, so its expectations are regenerated. Case [07] now emits the same two lines as headers_with_mbstring.phpt, where the two paths previously disagreed.

Test

tests/rfc2047_qp_chunk_length.phpt sweeps the cut point across the value, 122 truncations for each of the two encodings, asserting that every encoded-word stays within 75 characters and every line within 76. It reports 98 failures against the current code and none with the fix.

As in #34, an --INI-- section disables mb_substr()/mb_strlen() so the test actually exercises the fallback in CI.

Test fix carried along

headers_without_mbstring.phpt never ran: it skips whenever ext/mbstring is present, which is every job in the CI matrix, so the non-mbstring path had no coverage at all, which is how its expectations came to encode this bug unnoticed. Adding an --INI-- section disabling the two functions makes it run. Unlike the -i command line flag, --INI-- also applies to --SKIPIF--, so the existing skip condition still acts as a fallback if the functions cannot be disabled.

The suite goes from 52 passed / 2 skipped to 54 passed / 1 skipped.

AI use disclosure

I used Anthropic's Claude LLM with Opus 5 to find this bug. It also suggested a fix, which I have reviewed and tested.

…r limit

encodeHeaderValue() cuts quoted-printable text with
"(.{0,$maxLength}[^\=][^\=])". The two trailing atoms each match one further
character, so every chunk could run two characters past its bound: a subject
of "Lorem ipsum ... magna aliqua" plus an ellipsis produced a 77-character
encoded-word on a 78-character line, against RFC 2047 limits of 75 and 76.

The [^\=][^\=] pair is what stops a cut landing inside a "=XX" escape, so the
repetition count is reduced by two rather than the atoms removed. The chunk
length floors added in 4216044 (pear#34) guarantee the count cannot go negative.

Only the encodeHeaderValue() fallback is affected, i.e. builds without
ext/mbstring; encodeMB() already checked its limit correctly and produces
byte-identical output for the same input after this change.

tests/headers_without_mbstring.phpt pinned the old chunking, so its
expectations are regenerated. Case [07] now emits the same two lines as
headers_with_mbstring.phpt, where the two paths previously disagreed.
The test skips whenever ext/mbstring is present, which is every job in the CI
matrix, so the non-mbstring header encoding path had no coverage at all. That
is how its expectations came to encode a chunking bug unnoticed.

Add an --INI-- section disabling mb_substr()/mb_strlen(). Unlike the -i
command line flag, --INI-- also applies to --SKIPIF--, so the existing skip
condition now acts as a fallback: the test runs when the functions can be
disabled, and still skips if they cannot.
@alecpl
alecpl merged commit 28e3e84 into pear:master Aug 4, 2026
12 checks passed
alecpl pushed a commit that referenced this pull request Aug 5, 2026
…37)

The "Q" character class escaped printable punctuation and every 8-bit byte but
neither \x00-\x1F nor \x7F, so C0 controls and DEL were written into the header
verbatim. RFC 2047 defines encoded-text as printable ASCII other than "?" and
SPACE, so any of them makes the encoded-word invalid, and a raw CR or LF in a
header is a header injection vector.

A literal LF failed in a second way with ext/mbstring. encodeMB() uses "\n" as
its own chunk separator and expands it into a fold when assembling the result,
so an LF in the value was indistinguishable from a chunk boundary: the value
was split into two encoded-words and the byte silently disappeared from the
decoded header.

Add the two missing ranges, which lets \x7B-\x7E, \x7F and \x80-\xFF collapse
into \x7B-\xFF. The class was duplicated verbatim in encodeQP() and encodeMB(),
the second copy carrying a "see encodeQP()" comment, so it moves into
Mail_mimePart::QP_ESCAPE_REGEXP as 4216044 (#34) did for MAX_CHARSET_LENGTH.

tests/headers_with_mbstring.phpt pinned the defect as expected output and is
regenerated. Case [31] encodes a Japanese subject to ISO-2022-JP, whose
charset-switching escapes are ESC, and one of the JIS X 0208 bytes involved is
\x0D: the expectations held ten raw ESC bytes and a bare carriage return inside
a Subject header, invisible unless viewed with cat -v.

Unlike #33, #34 and #35 this reproduces with ext/mbstring present, so
tests/rfc2047_control_chars.phpt needs no --INI-- section. It sweeps all 32 C0
control characters plus DEL across both encodings, checking that the
encoded-text holds only printable ASCII and that the byte survives a round
trip. It reports 34 failures against the previous code.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants