Local ai branch - #10
Merged
Merged
Conversation
Replaces the dead tag-triggered publish.yml with a release.yml workflow that runs on every push to main: tests gate the release, then semantic-release derives the next version from Conventional Commits, updates CHANGELOG.md/package.json, tags, publishes to the VS Code Marketplace via vsce, and creates a GitHub Release. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Selecting an AI provider/model stalled for up to a minute on both macOS and Windows. The setup flow was a chain of serial, mostly-unbounded awaits with no UI rendered during any of it, so the wait read as a hang. Three causes: 1. listCopilotModels awaited vscode.lm.selectChatModels with no timeout. That call does not read a cached list — it waits for Copilot Chat to activate, sign in and fetch its catalog, which is unbounded and cold for exactly the users who never touch Copilot: warming it is now skipped unless Copilot is the selected backend, so the probe pays the full activation itself, outside the 5s guard in tryEnsureCopilotActivated. Now capped via opts.timeoutMs (default 3000) and throws on timeout so the output channel gets a real diagnostic instead of a silent empty list. The abandoned promise gets a no-op catch so a late rejection cannot surface as an unhandled rejection. selectModel() is left uncapped on purpose — that is the request path, where waiting for the model is the work. 2. speak() awaited speakMessage, which resolves from say.speak's completion callback — i.e. only once the sentence has finished being read aloud. Measured at ~5.2s for a short line, paid before every quick pick. Now synchronous fire-and-forget. Since speakMessage stops whatever is currently being spoken, back-to-back speech truncates rather than queues; the one affected pair (the post-install confirmation) is folded into a single utterance via saveOllamaChoice's new spokenPrefix argument. 3. detectProviders probed Copilot, Ollama and the hosted API sequentially, costing the sum of three unrelated backends. Now Promise.all, so it costs the slowest. Verified against a mock where Copilot hangs for 60s, Ollama is down and no API is configured: detectProviders returns in 3001ms (was unbounded), 202ms when Copilot is warm, and no unhandled rejection follows the abandoned probe. Note: `npm test` does not run on this branch and did not before this change — test/aiProviderSetup.test.ts fails to load with "require is not defined in ES module scope". Confirmed pre-existing by stashing. Verified with a standalone harness instead; `npm run compile` passes clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Voice setup could re-enter the venv create/install path indefinitely, showing repeated "Creating isolated Python environment..." notifications and re-running expensive work on every activation and every feature reload. Several distinct causes, each able to produce that on its own: - Only `pip install` was guarded against repeated failures; creating the venv was not. A venv that kept failing to build re-prompted and retried forever. The attempt counter now gates the whole slow path. - `python -m venv` can exit 0 without producing the expected interpreter (a layout difference, or a half-written directory from an earlier crash). That state recorded no failure, so the path stayed missing and every activation tried again. Creation is now verified, and a mismatch records a failure and logs what the venv directory actually contains. - extension.js constructs a new DependencyManager on activation *and* on every feature reload (a UserImplementation file change or a featureImplementation setting change each trigger one). Instance state could not stop those from re-entering setup, so the "already attempted" memo is module-level, scoped to the extension host. Only the slow path is memoised; the cheap "is it already installed" check still runs every call. - No cross-window lock: every VS Code window ran setup against the same venv directory, so two windows could run `python -m venv` into it concurrently, or two pip installs into one site-packages. Now an O_EXCL lock file, with a 20-minute staleness window so a crashed window cannot wedge voice forever. - No timeout on any shell-out. `exec` has no default, and on Windows `python --version` can hang indefinitely on the Microsoft Store App Execution Alias — indistinguishable from the extension never finishing. Every command is now bounded (15s probes, 3min venv, 20min pip). - `exec`'s default maxBuffer is 1MB. pip can overrun it, which kills the child with ENOBUFS and surfaces as a failed install that then retries on every launch. Raised to 32MB. Also: - `pip install --upgrade pip` ran unconditionally, costing a measured 468ms of network round-trip even when already satisfied. It now runs only right after venv creation, and its failure no longer aborts setup. - Install success is confirmed against site-packages rather than pip's exit code, since a partially-installed wheel can still exit 0. - Every command logs its duration, so a slow activation is diagnosable from the output channel instead of guessed at. - Python candidates are newest-first (3.12 -> 3.11 -> 3.10 -> python3); macOS's bundled Xcode python3 is 3.9 and should only ever be the last resort. - The first-run prompt gained "Don't ask again", remembered in globalState, so declining no longer re-prompts on every launch. The warm path is now two stats with no process launches: 1ms against an existing venv, where it previously could spawn Python to `import faster_whisper` (measured at 1519ms cold, worse on Windows where ctranslate2 and onnxruntime DLLs are scanned). Verified against harnesses covering each path: fast path, lock contention, stale-lock reclaim, timeout enforcement, the attempt limit, and six consecutive activations with a venv that never materialises. Not yet exercised against a live extension host. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
28 diagrams across 14 features, one .drawio file each with a page per level. Layer vocabulary and styling follow the "Our Proposed Approach" architecture diagram in the research slides, so these read as zoom-ins on it rather than a separate notation: Input Layer, VS Code Environment, EchoCode Core, AI Services Layer, External Services, plus the standalone Speech Output / TTS and Local Storage & State sinks. Level 0 shows the feature as one numbered process (N.0) between its sources and sinks, every flow labelled. Level 1 opens it into numbered sub-processes (N.1 ... N.5) inside the EchoCode Core band, with the data stores each one touches. Generated rather than hand-drawn: features.js holds each feature's processes, stores and flows; generate.js holds style and layout. A style change reapplies to all 28 at once, which is what makes a convention correction cheap. Two layout rules that are load-bearing, both documented in the README because dropping either brings back a specific annoyance: - Every edge pins explicit exit and entry points. Without them draw.io picks perimeter points itself, so edges detour around shapes and converging flows stack on a single spot. Flows now fan across an edge. - The EchoCode Core band is a real container and the step boxes are its children. A sibling band renders on top of the steps, since later cells draw over earlier ones, which meant sending it to the back by hand on every open. As a container it stays behind permanently and drags the pipeline with it. Data stores sit along the bottom rather than in a fourth column, which removes the crossings over the services column structurally instead of routing around them. Verified: XML parses, no dangling endpoints, all 241 edges pin both ends, no overlapping shapes, nothing outside the page bounds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds render_svg.py, which reads geometry straight out of the .drawio files and emits SVG. No draw.io CLI or headless browser is available here, and reading the committed XML rather than recomputing layout means the images cannot drift from the editable source. Each feature page now carries both levels, respecting the page's users-above-developers ordering: the Level 0 context diagram closes the "For users" half, and the Level 1 decomposition opens "For developers". Both link back to the .drawio for editing. Two routing fixes in the generator, found by rasterising the output and looking at it: - Data store flows left through the bottom of their step, which meant dropping straight through the Core band and every step box below. They now exit sideways and fall through a clear corridor: the gap between the inputs column and the band on the left, or between the band and the services column on the right. - Input flows all entered at the same height, so two feeding nearby steps ran their horizontal legs, and therefore their labels, on top of each other. Entry heights are now staggered. The SVG renderer prefers horizontal segments for edge labels, since text reads along them and two flows sharing a vertical corridor would otherwise stack their labels. Verified: all 56 image references resolve, 28 SVGs present, no overlapping shapes, nothing outside the page bounds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Removes wiki/diagrams/ from version control and gitignores it: the .drawio sources, generate.js and features.js, render_svg.py, and the rendered SVGs. Files stay on disk; only the index changes. Everything under it is reproducible from the two generators, so it is regenerable rather than source of record: node wiki/diagrams/generate.js && python3 wiki/diagrams/render_svg.py Also includes the .obsidian config entries that were already in the working tree. Note that the 14 feature wiki pages still reference diagrams/svg/*.svg, so those 56 images will not resolve in a fresh clone. Left in place deliberately rather than stripped, since the pages render correctly for anyone who has generated the diagrams locally. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The wiki pages embed 28 SVGs, so those have to resolve in a fresh clone
or every feature page shows broken images. The .drawio sources and the
two generators stay untracked — they are the inputs, reproducible on any
machine that wants to regenerate.
The ignore rules are shaped deliberately. Excluding "wiki/diagrams/"
outright makes the negations impossible, because git cannot re-include a
file whose parent directory is excluded; the contents are excluded and
svg/ re-included instead:
wiki/diagrams/*
!wiki/diagrams/svg/
wiki/diagrams/svg/*
!wiki/diagrams/svg/*.svg
Also delinks the 28 "Editable source" references on the feature pages.
They pointed at diagrams/*.drawio, which is no longer tracked, so as
links they would have 404'd for every reader. They now name the path as
plain text and point at the local README, keeping the information
without the dead link.
Verified: 145 markdown references across the wiki resolve, none missing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The per-feature Level 0/1 set is reference material: 28 diagrams, dense, portrait, one per feature. Wrong shape for a projector. These are 16:9, sparse, and sized to be read from the back of a room. They also take the obvious compression. Fourteen features reduce to five repeated shapes, so rather than showing each feature, S2 names each shape once and lists its instances underneath: A Direct read-out 6 features, no AI, works offline B AI code analysis 4 features, local rules pass first C External validation 2 features, pylint and g++ D Voice intake 3 modes over one pipeline E Workspace mutation 3 features, never overwrites silently The set: S1 Component Architecture the five layers and what sits in each S2 Interaction Patterns fourteen features, five shapes S3 Provider Abstraction one funnel, three backends S4 Voice Pipeline transcribed on-device S5 Output Path everything reaches the user as speech S1 is the one the deck's "Our Proposed Approach" slide expands into: same layer vocabulary and palette, but each layer opened up to name its components. Only the rendered SVGs are tracked, matching the existing split — the .drawio sources and slides.js stay local. Regenerate with: node wiki/diagrams/slides.js && python3 wiki/diagrams/render_svg.py Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Blind and Low Vision" is the currently preferred term over "Blind and Visually Impaired". Changes the actor label in both diagram generators and regenerates all 33 affected files. The literal acronym appeared in exactly two source lines; everything else is derived from them. Note that a case-insensitive search for "bvi" reports many false hits, because "webview" contains it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.