Skip to content

🤖 perf: instrument chat-switch latency (server replay phases + switch-to-first-message) and add a perf scenario #4504

Description

@ThomasK33

Problem

On xum server, switching from one running chat to another sometimes shows a skeleton for several seconds before any messages appear. A research pass (2026-09-25) found the mechanism:

  • The renderer hides cached rows while the switch is hydrating (cachedTranscriptStale, WorkspaceStore.ts ~1735-1745 and ~2465-2471; ChatPane.tsx ~603-618) until caught-up arrives or a 10 s deadline passes.
  • The server sends nothing until replayHistory finishes (src/node/orpc/streamBridge.ts ~107-167).

There is no measurement of the real per-phase costs, so we can't rank the causes on a real server or prove the planned fixes (A: show cached rows on incremental returns; B: progressive replay and keepalives). This issue adds that measurement. It is step F of F → A → B.

Scope

  1. Server replay timing. For each onChat replay, record the phase durations: file-lock wait, history read and parse, prior-history fingerprint, emitting history rows, in-flight stream replay, and total time to caught-up. Also record the rows and bytes sent, the replay mode (full or since) and any downgrade reason. Emit one structured log line per replay; use log.debug, or a higher level only above a threshold, e.g. >1 s. Reuse existing timing helpers if suitable. No new persisted state.
  2. Renderer switch timing. Using performance.mark/measure or an equivalent that e2e tests can read, time from setActiveWorkspaceId to (a) the first transcript row painted, (b) caught-up, and (c) the skeleton shown and hidden. Keep the overhead negligible, with no new UI.
  3. Perf e2e scenario. Add tests/e2e/scenarios/perf.chatSwitch.spec.ts (or similar): two workspaces, B with a large history, both streaming via mock AI. Switch A → B → A → B and record in-page milestones: reuse tests/e2e/utils/pageMilestones.ts from 🤖 tests: record in-page perf milestones for workspace open #4456, and extend it if the switch needs different selectors. Write them to perf-summary.json so the nightly Perf Profiles workflow collects them. The scenario must show the skeleton window on the switch back to a mid-stream chat. Assert only a generous upper bound, so it isn't flaky.

Acceptance

  • One server log line per replay with the phase breakdown. Verified by a test on the timing wrapper's behavior, not on log text.
  • Renderer marks exist, and the new perf scenario records them.
  • make test-e2e-perf passes locally, including the new scenario.
  • A before-numbers table in the PR body (median of ≥5 runs) that A and B will compare against.

Constraints

Instrumentation and measurement only: no behavior change to gating or replay. Follow AGENTS.md (no manual memoization, log helper). Refs the research report in the perf-owner workspace; see the follow-ups for A and B.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions