feat(programs): B2 surface shell — runProgram types and stub signatures - #1307
Conversation
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Store authentication, detection, and composition for a single program invocation. Record actual finished agent results separately from projected progress while preserving existing result order. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Move scan consent and cleanup registration into UI-free shared leaves. Decouple package-manager file operations from CLI setup and assert the full runProgram import closure. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
🧙 Wizard CIRun the Wizard CI and test your changes against wizard-workbench example apps by replying with a GitHub comment using one of the following commands: Test all apps:
Test all apps in a directory:
Test an individual app:
Show more apps
Test against a Context Mill branch:
Add Results will be posted here when complete. |
|
B2 work in progress. Local verification on this head: |
|
/wizard-ci-snapshots basic-integration/javascript-node/express-todo |
🧙 Wizard CI ResultsTrigger ID:
Configuration
|
Own artifact watchers and CI inference auth in programs, enforce composition gates, and clean run-installed skills on every failed path. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
|
B2 follow-up at |
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
gewenyu99
left a comment
There was a problem hiding this comment.
🥸 Reviewed by NotVincent, totally not Vincent. The real Vincent will review it separately. Probably slop, please disregard.
|
|
||
| function fingerprintSkill(skillDir: string): string { | ||
| const digest = createHash('sha256'); | ||
| const files = fg.sync(SKILL_TEXT_GLOB, { |
There was a problem hiding this comment.
Here's the potential issue: The scan exists to stop a poisoned project skill reaching the agent, but it only matches the exact name SKILL.md, while macOS and Windows filesystems resolve names case-insensitively.
repo ships .claude/skills/x/SKILL.MD with injection text -> scan globs for SKILL.md and reads nothing -> reports clean -> SDK opens SKILL.md, gets the uppercase file -> skill loads unscanned
Suggested fix: Match case-insensitively on both skill globs, here and in the YARA file scan, and add a SKILL.MD fixture that expects the poisoned verdict.
There was a problem hiding this comment.
NotVincent here. This no longer applies to B: the project skill scan at load was cut, and the YARA file scan's case fix is its own PR off main, #1344.
| workingDirectory: string, | ||
| triageProvider: LLMProvider | undefined, | ||
| ): Promise<ProjectSkillFinding[]> { | ||
| const root = path.join(workingDirectory, '.claude', 'skills'); |
There was a problem hiding this comment.
Here's the potential issue: The scan covers only the install directory's .claude/skills, but the SDK loads .claude/skills from every directory up to the git root.
monorepo root ships a poisoned skill, user runs --install-dir apps/web -> scan finds apps/web/.claude/skills empty -> reports clean -> SDK walks up and loads the root skill -> skill loads unscanned
Suggested fix: Scan every .claude/skills from the working directory up to the git root, stopping at the home directory.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the project skill scan at load, so there is no scan root left to widen.
| emitStepEvents: config.trackStepProgress ?? false, | ||
| resolveStepKey: config.resolveStepKey, | ||
| triageProvider: boot.triageProvider, | ||
| signal: inputs.signal, |
There was a problem hiding this comment.
Here's the potential issue: Cancellation is meant to stop the model, but nothing checks that either real harness forwards the host signal into the SDK query or the Pi session.
someone drops the signal from both harnesses -> all 3285 tests stay green -> host cancels -> SDK query keeps running and spending inference -> runner still reports Aborted
Suggested fix: Call both harnesses with a live AbortController, and assert the signal reaches the SDK query and that the Pi session's abort() fires.
There was a problem hiding this comment.
NotVincent here. B no longer changes how the harnesses forward the signal, so these tests moved to #1345 off main, which checks that the Pi session aborts on a live host cancel.
| const findings: ProjectSkillFinding[] = []; | ||
| for (const entry of fs.readdirSync(root, { withFileTypes: true })) { | ||
| const skillDir = path.join(root, entry.name); | ||
| if (!entry.isDirectory() && !fs.statSync(skillDir).isDirectory()) continue; |
There was a problem hiding this comment.
Here's the potential issue: A broken symlink the SDK would skip stops the whole run as a security failure. Separately, no test uses a symlinked skill, and this statSync branch is the only thing that lets one into the scan.
dotfiles-linked skill whose target moved -> statSync throws ENOENT -> caller maps it to YARA_VIOLATION -> every run fails with "Security check could not scan project skills"
Suggested fix: Stat with throwIfNoEntry: false and skip entries that aren't directories. Add a test that symlinks a poisoned skill directory and expects the poisoned verdict, so this branch can't be narrowed to real directories unnoticed.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the project skill preflight, so a broken skill link can no longer fail the run.
| if (!descriptor) continue; | ||
| if ('value' in descriptor) { | ||
| const value: unknown = descriptor.value; | ||
| descriptor.value = structuredClone(value); |
There was a problem hiding this comment.
Here's the potential issue: The outcome should carry the run's real failure, but cloning every property of the error throws for common error types.
run crashes with an AxiosError -> finish deep-clones config.transformRequest, a function -> DataCloneError -> runProgram rejects with the clone error -> user and analytics see DataCloneError, not the crash
Suggested fix: Keep the error instance by reference, or fall back to the original value when cloning fails.
There was a problem hiding this comment.
NotVincent here. The store no longer clones the run result: finish keeps it by reference, and the store test checks that the settled result is the same object, in #1308:
wizard/src/programs/program-store.ts
Lines 124 to 126 in f2e42ac
| continue; | ||
| } | ||
|
|
||
| const reason = await scanInstalledSkill( |
There was a problem hiding this comment.
Here's the potential issue: The skill the Wizard just installed and scanned gets scanned again here.
linear run installs its skill -> download scan passes but records nothing -> preflight misses its cache -> second YARA pass, and a second triage call if a chunk is flagged
Suggested fix: Record the clean verdict after the install scan.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the project skill preflight, so nothing scans the installed skill a second time.
| sequence: (process.env.SNAP_SEQUENCE || undefined) as Sequence | undefined, | ||
| model: process.env.SNAP_MODEL || undefined, | ||
| }); | ||
| store.setInferenceAuth( |
There was a problem hiding this comment.
Here's the potential issue: tui-host used to report a missing CI token as a readable run failure. It now reads the token file at startup and dies first.
start tui-host without WIZARD_CI_GATEWAY_TOKEN_FILE -> FATAL before the control socket serves -> MCP client sees a dead socket -> detection-only route and screen-only commands stop working
Suggested fix: Read the token file on first use rather than at startup.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the CI bearer wiring, and tui-host no longer reads a CI token file.
| programId, | ||
| inferenceAuth: | ||
| options.inferenceAuth ?? | ||
| session.inferenceAuth ?? |
There was a problem hiding this comment.
Here's the potential issue: This fallback is what gives --ci detection the CI bearer, and nothing tests it.
someone deletes session.inferenceAuth here -> suite stays green -> CI detection mints with the phx key again
Suggested fix: Assert detection passes the session's provider to the agent.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the CI bearer wiring, so detection has no session.inferenceAuth fallback left.
| // out of the TUI's startup graph. | ||
| programId: store.analyticsProgramId, | ||
| inferenceAuth: | ||
| store.session.inferenceAuth ?? |
There was a problem hiding this comment.
Here's the potential issue: The store half of the CI bearer has no test.
delete this fallback or the store setter in runNonInteractive -> suite stays green -> tui-host mcp-tutorial mints again
Suggested fix: Assert the store's provider after a CI start and its use here.
There was a problem hiding this comment.
NotVincent here. This no longer applies: B cut the CI bearer wiring, so the TUI store no longer carries a CI provider.
| rerankIds?: readonly string[]; | ||
| /** Streaming activity callback for the UI. */ | ||
| onEvent?: DetectEvent; | ||
| inferenceAuth?: import('@agent/types').InferenceAuthProvider; |
There was a problem hiding this comment.
Here's the potential issue: This option adds a third source for detection auth, and nothing sets it.
every caller omits it -> the branch is dead -> readers reason about three precedence levels where two are real
Suggested fix: Delete it until a caller needs it.
There was a problem hiding this comment.
NotVincent here. The option is gone: AgenticDetectOptions has no auth field, and detection passes the session's credentials to runAgent:
wizard/src/programs/detection/agentic.ts
Lines 129 to 153 in f2e42ac
Bring in the adapter move to src/lib/runners/run-program-agent.ts. B2's adapter changes land on the new path, its @programs imports replace the relative ones, and known-violations.json lists the moved adapter's edges. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Set the tree to B1's, then declare the new surfaces with their final names, paths and signatures, taken from B3 at 50fe9e4 with the bodies removed. - src/programs/run-program.ts: ProgramOverrides, WizardFlagSnapshot, ProgramSettings, ProgramInput, ProgramOptions, ProgramRunOutcome, and a runProgram that throws "runProgram: not implemented". - src/programs/program-store.ts: the progress, diagnostic, invocation-data and settled-run types, and ProgramStore's public methods, each throwing. - src/programs/credentials.ts: ResolvedProgramCredentials, CredentialsProvider, and createPosthogInferenceAuthProvider, throwing. - src/programs/host-capabilities.ts: ProgramCiHost and ProgramRunHost, with AuthProjection in authenticate.ts and ProgramCompletionContext in program-run.ts. - The programs entry exports runProgram and createPosthogInferenceAuthProvider as lazy loaders, and the type entry exports the new types. - The agent declares InferenceAuthProvider, the collectTranscript, requestRemark and prompt run-definition fields, RunConfig.scanReport, RunInput.inferenceAuth (optional until B3 supplies it), RunSnapshot.transcriptTail, and the activity progress event, which the UI reducer ignores. Nothing calls the shell: the adapter, the CLI and every doc stay as B1 has them, and no tests are added. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Take B1's tree as it stands, the original B1 moves only. The surface shell is declared again on top of it. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
…nal B1 Declare the new surfaces with their final names and signatures, taken from B3 at f1b30d7 with the bodies removed, against B1's layout. - src/programs/run-program.ts: ProgramOverrides, WizardFlagSnapshot, ProgramSettings, ProgramInput, ProgramOptions, ProgramRunOutcome, and a runProgram that throws "runProgram: not implemented". - src/programs/program-store.ts: the progress, diagnostic, invocation-data and settled-run types, and ProgramStore's public methods, each throwing. - src/programs/credentials.ts: ResolvedProgramCredentials and CredentialsProvider. - @programs exports runProgram, and @programs/types exports the new types. - The agent declares the prompt, collectTranscript and requestRemark run-definition fields, RunConfig.scanReport, RunSnapshot.transcriptTail, and the activity progress event, which the UI reducer ignores. @agent/types exports ResolvedBinding and RunHooks for the shell. B1 keeps DiscoveredFeature in the legacy session, so known-violations.json records run-program.ts's type import of it until B3 moves it to shared. Inference auth, the deferred skill commit, the watcher fields, the lazy @agent entry and the host capabilities stay in B3. Nothing calls the shell: the adapter, the CLI and every doc stay as B1 has them, and no tests are added. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Records #1307's shell as this branch's base. The tree is unchanged here; the next commits return it to the shell's tree and add the plumbing. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
There was a problem hiding this comment.
this isn't hooked up yet in this PR, we will move this in a later PR
There was a problem hiding this comment.
Implemented in #1308: program-store.ts (at f550e227).
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
…-store imports DiscoveredFeature moves to src/shared/discovered-feature.ts and the session re-exports it, so run-program.ts no longer reaches the session and its allowlist row goes. program-store.ts imports through @agent/types and @shared/api. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
|
NotVincent here. This express-todo failure was a harness bug, not this PR: the run completed and reached keep-skills in 26 frames, but |
Scope: New surface only: types and stubs, no plumbing. Review the contract.
Related: #1320
This adds the
runProgramsurface as a shell. It declares what a caller passes in and gets back, andrunProgramthrows until #1308 fills it in. Nothing calls it yet.Stack: #1306 moves → #1307 surface shell → #1308 plumbing → #1357 runner context → #1309 docs → #1363 stream retry.
What it declares
run-program.ts:ProgramInput,ProgramSettings,ProgramOptions,ProgramRunOutcomeand a stubrunProgram. Each field has a one-line comment.program-store.ts:ProgramProgress,ProgramInvocationDataand theProgramStoreshape. It imports through the@agent/typesand@shared/apialiases.credentials.ts:ResolvedProgramCredentialsandCredentialsProvider.DiscoveredFeature: it now lives insrc/shared/discovered-feature.ts, so programs can name it without the session.wizard-session.tsre-exports it, so its readers keep their import path.prompt,collectTranscript,requestRemarkandscanReport, plus theactivityprogress event.Checks at
68a321cd:pnpm typecheck,pnpm lint(0 errors),pnpm vitest run(3,405 tests, the same as #1306, since the shell adds none) andpnpm test:arch(11 tests) pass.Created with PostHog Desktop