Follow-up to #536, raised by the architecture review of its fix (merged as
#538, v0.94.1).
#536 made the per-eval heap budget relative to the state the callback was
handed, which is correct: the old absolute cap bounded the persistent Luerl
state instead of the callback, and killed handlers that had allocated nothing.
But it is worth being precise about what that did and did not remove. There
was never a hard ceiling on the persistent state. The old cap killed the
eval worker; the state lived in the parent gen_server and was untouched, so
a 690 MB state was 690 MB before the kill and 690 MB after. What #536 removed
was a mechanism that produced a denial of service on the callback without
reclaiming a byte.
So a self-hoster running unattended still has nothing that stops a Lua state
growing until the node dies. [asobi, lua, state] and the lua_state_large
warning make it visible, which is a real improvement, but visibility is not a
backstop.
Proposal
asobi_lua.max_state_words, default infinity, honoured on zones only.
Past the ceiling the zone does {stop, {lua_state_ceiling, Words}, State}
rather than killing anything. That works because the machinery already exists:
asobi_zone_snapshotter persists the zone's game_state
(asobi_lua_world:dump_zone_state/1 decodes it to a plain map).
asobi_zone:maybe_restore_from_snapshot/1 (src/world/asobi_zone.erl:316)
already rebuilds a zone from that snapshot on boot, and
asobi_lua_world:restore_game_state/2 re-encodes it into a fresh VM.
So the supervisor restart reconstructs a clean Luerl state from the last
snapshot and restores service, visibly, with a crash report. A zone whose
state has passed such a ceiling is already dead as a gameplay entity anyway -
at ~7 ms per MB of copying, a 690 MB state is a 4.8 s tick.
Constraints
- Default must be
infinity. A library does not get to decide for an
operator that a 500 MB state is fatal.
- Zones only. A match has no snapshot:
asobi_match_server holds
game_state in the FSM and never persists it, so a restart there loses the
match. Document it as zone-only rather than wiring it to matches by symmetry.
- What a script keeps outside
game_state (module locals, upvalues, globals)
does not survive the restart. asobi has never promised to persist those, but
it has also never destroyed them mid-session, so this needs saying in the
guide.
Why it was not in #538
The security review of that PR asked for the opposite shape - an absolute
ceiling on the eval worker, min(2 * Base + Budget, ceiling). That was
declined for the reason above: killing the eval worker reclaims nothing, so a
ceiling there converts "the node runs out of memory" into "the zone stops
ticking forever and the node still holds the state". That is #536 by
construction, at a different number. The reporter reached 4.5 GB in one zone
with the old absolute cap in place, which is the evidence that the eval cap
is not where this belongs.
Not included
The architecture review also suggested replacing the collector's cost-based
back-off with a yield-based one: a cost-only loop cannot tell "cheap
because it reclaimed a lot" from "cheap because there was nothing to reclaim",
so a script holding a genuinely live large table pays for a collection every
interval and reclaims nothing. #536's fix capped the budget instead, which
keeps the back-off reachable but does not close that. Worth doing separately.
Follow-up to #536, raised by the architecture review of its fix (merged as
#538, v0.94.1).
#536 made the per-eval heap budget relative to the state the callback was
handed, which is correct: the old absolute cap bounded the persistent Luerl
state instead of the callback, and killed handlers that had allocated nothing.
But it is worth being precise about what that did and did not remove. There
was never a hard ceiling on the persistent state. The old cap killed the
eval worker; the state lived in the parent gen_server and was untouched, so
a 690 MB state was 690 MB before the kill and 690 MB after. What #536 removed
was a mechanism that produced a denial of service on the callback without
reclaiming a byte.
So a self-hoster running unattended still has nothing that stops a Lua state
growing until the node dies.
[asobi, lua, state]and thelua_state_largewarning make it visible, which is a real improvement, but visibility is not a
backstop.
Proposal
asobi_lua.max_state_words, defaultinfinity, honoured on zones only.Past the ceiling the zone does
{stop, {lua_state_ceiling, Words}, State}rather than killing anything. That works because the machinery already exists:
asobi_zone_snapshotterpersists the zone'sgame_state(
asobi_lua_world:dump_zone_state/1decodes it to a plain map).asobi_zone:maybe_restore_from_snapshot/1(src/world/asobi_zone.erl:316)already rebuilds a zone from that snapshot on boot, and
asobi_lua_world:restore_game_state/2re-encodes it into a fresh VM.So the supervisor restart reconstructs a clean Luerl state from the last
snapshot and restores service, visibly, with a crash report. A zone whose
state has passed such a ceiling is already dead as a gameplay entity anyway -
at ~7 ms per MB of copying, a 690 MB state is a 4.8 s tick.
Constraints
infinity. A library does not get to decide for anoperator that a 500 MB state is fatal.
asobi_match_serverholdsgame_statein the FSM and never persists it, so a restart there loses thematch. Document it as zone-only rather than wiring it to matches by symmetry.
game_state(module locals, upvalues, globals)does not survive the restart. asobi has never promised to persist those, but
it has also never destroyed them mid-session, so this needs saying in the
guide.
Why it was not in #538
The security review of that PR asked for the opposite shape - an absolute
ceiling on the eval worker,
min(2 * Base + Budget, ceiling). That wasdeclined for the reason above: killing the eval worker reclaims nothing, so a
ceiling there converts "the node runs out of memory" into "the zone stops
ticking forever and the node still holds the state". That is #536 by
construction, at a different number. The reporter reached 4.5 GB in one zone
with the old absolute cap in place, which is the evidence that the eval cap
is not where this belongs.
Not included
The architecture review also suggested replacing the collector's cost-based
back-off with a yield-based one: a cost-only loop cannot tell "cheap
because it reclaimed a lot" from "cheap because there was nothing to reclaim",
so a script holding a genuinely live large table pays for a collection every
interval and reclaims nothing. #536's fix capped the budget instead, which
keeps the back-off reachable but does not close that. Worth doing separately.