Skip to content

πŸ€– perf: measure Perf Profiles milestones in-page; accept the 09-21 workspace-open-large stepΒ #4441

Description

@ThomasK33

Summary

The nightly Perf Profiles workspace-open-large numbers stepped up on 09-21: script ~354 β†’ ~500 ms, task minus DevTools ~469 β†’ ~650 ms, wall ~720 β†’ ~1320 ms. The step is mostly a change in what the spec measures, not extra work. Accept it as the new baseline, and make the harness robust to this kind of change.

What happened

#4293 (4e258e6) changed data-loaded to wait until the tail-first reveal has mounted every row (ChatPane.tsx, the data-loaded attribute). perf.workspaceOpen.spec.ts stops the wall clock and reads Performance.getMetrics when Playwright's toHaveAttribute("data-loaded", "true") passes. Three effects follow:

  1. Poll quantization. Playwright polls once, then at +100/+250/+500/+1000 ms (playwright-core server/frames.js). The gap between the in-page data-loaded flip and the passing poll went from a median of 27 ms to 423 ms (local, n=6 each).
  2. Window shift. About 100 ms of post-load work used to run after the metrics were read. It now runs inside the window. With a fixed 2 s window after load, the script difference shrinks to +24/+30 ms (two rounds, n=6 each). CPU-profile non-idle time is ~800 ms at both endpoints.
  3. Real, intended trade. Full reveal takes +~150 ms (MutationObserver flip time: ~1020 β†’ ~1190 ms). In exchange, the longest main-thread task drops from ~395 to ~161 ms.

The small real leftover (redundant MessageRenderer re-renders on each reveal step, from #4292's inline handleEditUserMessage) is fixed by #4427.

Proposed follow-up (harness only)

  • Record milestone timestamps in-page (a MutationObserver on the settled marker, measured from the click). Do not rely on the Playwright poll that happens to observe it.
  • Report "useful content ready" (first transcript rows) separately from "fully revealed". Keep longest-task as its own metric.
  • Version the measurement contract. When a completion milestone changes meaning, reset the trend baseline instead of flagging a regression.

Do not use a fixed sleep or a vague "quiescence" wait. It was useful as a diagnostic here, but it depends on the workload and adds suite time.


Generated with xum β€’ Model: anthropic:claude-opus-5-5 β€’ Thinking: high

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions