Skip to content

fix(stdlib): enforce max_string_bytes in concat, gsub, format, pack and os.date - #426

Open
MrYawe wants to merge 1 commit into
tv-labs:mainfrom
MrYawe:fix/max-string-bytes-stdlib
Open

MrYawe wants to merge 1 commit into
tv-labs:mainfrom
MrYawe:fix/max-string-bytes-stdlib

Conversation

@MrYawe

@MrYawe MrYawe commented Oct 1, 2026

Copy link
Copy Markdown

Closes #425.

Why

The sandboxing guide tells embedders to set :max_string_bytes below their :max_heap_size cap, so that string bombs are refused deterministically and the heap cap stays a backstop. That only held for string.rep, .. and the load reader. table.concat, string.gsub, string.format, string.pack and os.date built their result without consulting the ceiling, so each could still produce an oversized string in one allocation, which is the case max_heap_size handles worst because it is only checked at garbage collection. Under the default configuration, string.pack("c2147483647", "") asked for about 2 GiB.

What changes

All five now work out the size of their result before building it and raise the same catchable resulting string too large error as string.rep. A result exactly at the ceiling still builds, and :infinity still disables the check.

  • table.concat and string.format total their pieces before the final join. The pieces only reference the caller's strings, so nothing large exists before the check.
  • string.pack, os.date and string.gsub build incrementally, so they check as they go and stop partway.

The :max_string_bytes option doc and the guide's resource-limits table now list the five functions, and there is a changelog entry.

Notes for review

Four things go beyond the fix sketched in the issue:

  • gsub capture expansion is sized before it is built. A running count over finished replacements is not enough: a replacement string made of repeated %0 can expand a single match far past the ceiling before the count is consulted. The expansion is now assembled as iodata (captures and literal runs are referenced, not copied), measured, then flattened. This is also where the gsub speedups below come from, since the old expansion rebuilt the replacement byte by byte on every match.
  • The error keeps side effects made before it. In gsub (replacement callbacks) and table.concat (__index, __len), Lua code runs before the size is known. The error carries the VM state at the raise, so pcall keeps those effects, the same way gsub's existing "invalid replacement value" error does. Limits.check_string_size! takes an optional state for this. table.concat's existing "invalid value" error still drops them, as on main; I left that alone.
  • os.date no longer goes through a regex. Walking the format directly lets it stop as soon as the ceiling is hit, without tokenising the whole format first. Output is unchanged: a new test covers every directive plus %%, unknown directives, a % before a newline and a trailing %, and it passes against the old implementation too.
  • string.char and utf8.char are not covered. They can exceed a very small ceiling, but emit at most 1 and 6 bytes per argument, so they cannot amplify. I can add them if you would rather the ceiling be absolute.

Unrelated to this change: the reset/0 tests in Lua.VM.BootstrapTest fail intermittently for me on unmodified main (4 of 25 runs of mix test --include lua53 at 3f2aead), in case CI trips on them.

Benchmarks

main at 3f2aead against this branch at 840ead7, each built in its own worktree. Apple M2 (8 cores), macOS 26.6.2, Elixir 1.20.0, Erlang/OTP 29.0.5 with the JIT.

Four alternating passes (main, branch, branch, main, main, branch, branch, main), each started once the 1-minute load average was below 3. Every figure is the median of the four per-pass medians. Lower is better, and the ratio is branch / main.

Project workloads

MIX_ENV=benchmark mix run benchmarks/<workload>.exs, quick mode, lua (chunk) job. The Luerl job runs identical code on both sides, so its ratio shows the noise floor for that row. luaport is not installed on this machine, so there is no C Lua job.

Workload main branch ratio Luerl ratio (control)
patterns: find/match field extraction (not touched) 1.39 ms 1.40 ms 1.01× 0.99×
patterns: find-based tokenizer (not touched) 2.15 ms 2.19 ms 1.02× 1.01×
patterns: gsub template substitution 2.75 ms 2.94 ms 1.07× 1.01×
string_ops: table.concat (n=100) 42.6 μs 42.6 μs 1.00× 0.99×
string_ops: string.format (n=100) 101.2 μs 100.6 μs 0.99× 0.97×
string_format: long literal-heavy (n=1000) 1.15 ms 1.11 ms 0.97× 0.99×
string_format: width-flagged (n=1000) 1.89 ms 1.86 ms 0.99× 1.00×
string_format: many specifiers (n=1000) 2.58 ms 2.52 ms 0.98× 1.00×

The gsub row includes one disturbed branch pass (3.78 ms, with Luerl also jumping to 3.95 ms in that pass). Over the other three passes it is 1.04×, and an earlier run of the same protocol on the same gsub code measured 1.02×, so I read it as 2–7% slower. It is the one place the per-replacement check shows up in the project's suite.

Paths the project workloads do not reach

A supplementary script (below), run with MIX_ENV=prod. Each case is a pre-compiled chunk that loops over the call under test, timed 25 times per pass.

Case main branch ratio
os.date, six directives (×20 000) 100.0 ms 52.0 ms 0.52×
string.pack, mixed fields (×100 000) 106.8 ms 108.1 ms 1.01×
string.pack, integers (×100 000) 118.3 ms 119.3 ms 1.01×
string.format, three specifiers (×100 000) 106.1 ms 111.0 ms 1.05×
table.concat, 1000 strings, with separator (×2000) 108.2 ms 115.5 ms 1.07×
table.concat, 1000 strings, no separator (×2000) 102.4 ms 90.9 ms 0.89×
gsub, 1-byte string replacement (80k matches) 25.8 ms 24.9 ms 0.97×
gsub, capture template %2=%1 (40k matches) 32.2 ms 27.5 ms 0.85×
gsub, capture template with markup (40k matches) 38.3 ms 29.8 ms 0.78×
gsub, 200-byte string replacement (25k matches) 426.0 ms 14.9 ms 0.035×
gsub, function replacement (30k matches) 8.46 ms 8.61 ms 1.02×
gsub, table replacement (60k matches) 14.60 ms 14.54 ms 1.00×

Where the check costs something: table.concat with a separator is about 7% slower at 1000 elements (one extra pass to total the sizes; not visible at n=100), and a tight three-specifier string.format loop is about 5% slower (not visible in the project's format workloads). Where it got faster: os.date about 1.9×, gsub with capture templates 15–22%, gsub with a long literal replacement about 29×, and table.concat without a separator about 11%.

Supplementary script
# Supplementary timings for the stdlib paths touched by the max_string_bytes
# fix that the Benchee workloads under benchmarks/ do not exercise.
#
#   MIX_ENV=prod mix run supplementary_bench.exs
#
# Each case is one pre-compiled chunk that loops over the call under test.
# It is warmed up, then timed 25 times; the median and minimum are reported.

cases = [
  {"os.date, 6 directives (x20000)",
   ~s|local n = 0 for i = 1, 20000 do n = n + #os.date("%Y-%m-%d %H:%M:%S", i) end return n|},
  {"string.pack, mixed fields (x100000)",
   ~s|local n = 0 for i = 1, 100000 do n = n + #string.pack("<i4 z s2 d", i, "abc", "defg", 1.5) end return n|},
  {"string.pack, integers (x100000)",
   ~s|local n = 0 for i = 1, 100000 do n = n + #string.pack("<i4 i4 i8 i2 d", i, i, i, 7, 1.5) end return n|},
  {"string.format, 3 specifiers (x100000)",
   ~s|local n = 0 for i = 1, 100000 do n = n + #string.format("%d:%s:%5.2f", i, "abc", 1.5) end return n|},
  {"table.concat, 1000 strings, separator (x2000)",
   ~s|local t = {} for i = 1, 1000 do t[i] = "item" .. i end local n = 0 for i = 1, 2000 do n = n + #table.concat(t, ",") end return n|},
  {"table.concat, 1000 strings, no separator (x2000)",
   ~s|local t = {} for i = 1, 1000 do t[i] = "item" .. i end local n = 0 for i = 1, 2000 do n = n + #table.concat(t) end return n|},
  {"gsub, 1-byte string replacement (80k matches)",
   ~s|local s = string.rep("hello world ", 200) local n = 0 for i = 1, 200 do local r = string.gsub(s, "o", "0") n = n + #r end return n|},
  {"gsub, capture template %2=%1 (40k matches)",
   ~s|local s = string.rep("key=value ", 200) local n = 0 for i = 1, 200 do local r = string.gsub(s, "(%w+)=(%w+)", "%2=%1") n = n + #r end return n|},
  {"gsub, capture template with markup (40k matches)",
   ~s|local s = string.rep("key=value ", 200) local n = 0 for i = 1, 200 do local r = string.gsub(s, "(%w+)=(%w+)", "<td>%1</td><td>%2</td>") n = n + #r end return n|},
  {"gsub, 200-byte string replacement (25k matches)",
   ~s|local s = string.rep("a,", 500) local rep = string.rep("-", 200) local n = 0 for i = 1, 50 do local r = string.gsub(s, ",", rep) n = n + #r end return n|},
  {"gsub, function replacement (30k matches)",
   ~s|local s = string.rep("abc ", 300) local n = 0 for i = 1, 100 do local r = string.gsub(s, "%w+", string.upper) n = n + #r end return n|},
  {"gsub, table replacement (60k matches)",
   ~s|local s = string.rep("abc def ", 300) local t = {abc = "x", def = "yy"} local n = 0 for i = 1, 100 do local r = string.gsub(s, "%w+", t) n = n + #r end return n|}
]

lua = Lua.new(sandboxed: [])

for {name, code} <- cases do
  {chunk, lua} = Lua.load_chunk!(lua, code)
  for _ <- 1..3, do: Lua.eval!(lua, chunk)
  times = for _ <- 1..25, do: elem(:timer.tc(fn -> Lua.eval!(lua, chunk) end), 0)
  sorted = Enum.sort(times)
  median = Enum.at(sorted, 12) / 1000
  min = hd(sorted) / 1000

  IO.puts(
    "RESULT|#{name}|#{:erlang.float_to_binary(median, decimals: 2)}|#{:erlang.float_to_binary(min, decimals: 2)}"
  )
end

Tests

  • 13 new tests for the ceiling at 1 KiB: each function over the ceiling and exactly at it, and for gsub every replacement kind, anchored patterns, capture expansion, the unmatched remainder, and side effects surviving pcall. One more test pins os.date output.
  • mix test --include lua53: 2904 passed. mix format --check-formatted, mix docs --warnings-as-errors and mix dialyzer are clean.
  • The reproduction script from the issue now refuses all seven cases.

…nd os.date

`:max_string_bytes` was checked by `string.rep`, the `..` operator and the
`load` reader, but `table.concat`, `string.gsub`, `string.format`,
`string.pack` and `os.date` built their result without consulting it. With
the ceiling set below a process heap cap, as the sandboxing guide
recommends, each of them could still produce a string past the ceiling in
one allocation -- `string.pack("c2147483647", "")` asked for 2 GiB.

Each now sizes its result before building it and raises the same catchable
"resulting string too large" error:

- `table.concat` sums the elements and separators before joining.
- `string.format` counts bytes as it renders and checks before flattening.
- `string.pack` checks before each field is appended.
- `os.date` and `string.gsub` keep a running size and stop partway.

`string.gsub` expands `%0`-`%9` into iodata so a replacement is sized
before it is built, and its error carries the VM state so a replacement
callback's side effects survive `pcall`. `table.concat` does the same for
`__index` side effects on this error.

`os.date` walks the format directly instead of through a regex; output is
unchanged and pinned by a new test.

Closes tv-labs#425
@MrYawe
MrYawe marked this pull request as ready for review October 1, 2026 12:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

max_string_bytes is not enforced by table.concat, string.gsub, string.format, string.pack or os.date

1 participant