Skip to content

feat(guardrails): multi-lane inference pool for the local-model guardrail #1001

Description

@membphis

Context

PR #999 ships the local-model guardrail MVP with ONE inference lane: a single ONNX session behind Semaphore(1) + Mutex, measured at ~50 inferences/s (p50 ≈ 19 ms per ~35-token window on a 12-core avx2+vnni host). The acceptance shape spends two inferences per request (input + output pass), so the single lane saturates around ~25 guardrail-active requests/s and all cores' requests queue on it.

Design (decided in the MVP review; recording it here so it doesn't rot in the PR thread)

Scale throughput by lanes, not by intra-op threads:

  • GUARDRAIL_LOCAL_MODEL_LANES=N (env, default 1 — same experimental env surface as the existing knobs).
  • N sessions, ONE process-wide admission semaphore with N permits, and a centralized free-list of sessions: a granted task takes whichever session is idle and returns it afterwards.
  • Explicitly NO worker↔session binding. Sessions are stateless loaded-model instances and fully interchangeable, so binding buys no correctness or locality; TPC workers receive connections unevenly (kernel accept distribution), so a static per-core split would idle some lanes while others queue. The centralized queue load-balances by construction. The only 1:1 binding that may exist is in the hard-isolation variant below, and it is inference-thread↔session — never request-worker↔session.
  • Keep intra_op = 1 per session: short windows (~35 tokens) parallelize poorly across intra-op threads; the throughput axis is lane count, and one thread per lane is the best per-core efficiency.
  • Optional hard-isolation variant (only if soft partitioning proves insufficient): replace spawn_blocking with a dedicated pinned inference thread pool, each thread OWNING one session (drops the mutex), consuming one bounded task channel; pin business workers and inference threads to disjoint core sets.

Cost / constraint to record

Each lane pays its own ~190 MiB int8 weight copy: ort 2.0.0-rc.13 types Session::run(&mut self), which forbids the shared-weights concurrent-Run form the ONNX Runtime C API documents as thread-safe, and aisix-guardrails is #![forbid(unsafe_code)]. Paths back to a single weight copy, any one of which closes this: an upstream &self run signature, a thin unsafe shim crate outside the forbid boundary, or the sidecar deployment form. Until then, N lanes ≈ N × weight memory (4 lanes ≈ +600–700 MiB) — size against the target host.

Acceptance

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Long-tail integrations — backlogenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions