Skip to content

Local ai branch - #10

Merged
Lunar2074 merged 26 commits into
mainfrom
Local-AI-branch
Oct 1, 2026
Merged

Lunar2074 merged 26 commits into
mainfrom
Local-AI-branch

Conversation

@Lunar2074

Copy link
Copy Markdown
Collaborator

No description provided.

Lunar2074 and others added 26 commits July 7, 2026 17:35
Replaces the dead tag-triggered publish.yml with a release.yml workflow
that runs on every push to main: tests gate the release, then
semantic-release derives the next version from Conventional Commits,
updates CHANGELOG.md/package.json, tags, publishes to the VS Code
Marketplace via vsce, and creates a GitHub Release.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Selecting an AI provider/model stalled for up to a minute on both macOS and
Windows. The setup flow was a chain of serial, mostly-unbounded awaits with no
UI rendered during any of it, so the wait read as a hang. Three causes:

1. listCopilotModels awaited vscode.lm.selectChatModels with no timeout.
   That call does not read a cached list — it waits for Copilot Chat to
   activate, sign in and fetch its catalog, which is unbounded and cold for
   exactly the users who never touch Copilot: warming it is now skipped unless
   Copilot is the selected backend, so the probe pays the full activation
   itself, outside the 5s guard in tryEnsureCopilotActivated.

   Now capped via opts.timeoutMs (default 3000) and throws on timeout so the
   output channel gets a real diagnostic instead of a silent empty list. The
   abandoned promise gets a no-op catch so a late rejection cannot surface as
   an unhandled rejection. selectModel() is left uncapped on purpose — that is
   the request path, where waiting for the model is the work.

2. speak() awaited speakMessage, which resolves from say.speak's completion
   callback — i.e. only once the sentence has finished being read aloud.
   Measured at ~5.2s for a short line, paid before every quick pick.

   Now synchronous fire-and-forget. Since speakMessage stops whatever is
   currently being spoken, back-to-back speech truncates rather than queues;
   the one affected pair (the post-install confirmation) is folded into a
   single utterance via saveOllamaChoice's new spokenPrefix argument.

3. detectProviders probed Copilot, Ollama and the hosted API sequentially,
   costing the sum of three unrelated backends. Now Promise.all, so it costs
   the slowest.

Verified against a mock where Copilot hangs for 60s, Ollama is down and no API
is configured: detectProviders returns in 3001ms (was unbounded), 202ms when
Copilot is warm, and no unhandled rejection follows the abandoned probe.

Note: `npm test` does not run on this branch and did not before this change —
test/aiProviderSetup.test.ts fails to load with "require is not defined in ES
module scope". Confirmed pre-existing by stashing. Verified with a standalone
harness instead; `npm run compile` passes clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Voice setup could re-enter the venv create/install path indefinitely, showing
repeated "Creating isolated Python environment..." notifications and re-running
expensive work on every activation and every feature reload. Several distinct
causes, each able to produce that on its own:

- Only `pip install` was guarded against repeated failures; creating the venv
  was not. A venv that kept failing to build re-prompted and retried forever.
  The attempt counter now gates the whole slow path.

- `python -m venv` can exit 0 without producing the expected interpreter (a
  layout difference, or a half-written directory from an earlier crash). That
  state recorded no failure, so the path stayed missing and every activation
  tried again. Creation is now verified, and a mismatch records a failure and
  logs what the venv directory actually contains.

- extension.js constructs a new DependencyManager on activation *and* on every
  feature reload (a UserImplementation file change or a featureImplementation
  setting change each trigger one). Instance state could not stop those from
  re-entering setup, so the "already attempted" memo is module-level, scoped to
  the extension host. Only the slow path is memoised; the cheap "is it already
  installed" check still runs every call.

- No cross-window lock: every VS Code window ran setup against the same venv
  directory, so two windows could run `python -m venv` into it concurrently, or
  two pip installs into one site-packages. Now an O_EXCL lock file, with a
  20-minute staleness window so a crashed window cannot wedge voice forever.

- No timeout on any shell-out. `exec` has no default, and on Windows
  `python --version` can hang indefinitely on the Microsoft Store App Execution
  Alias — indistinguishable from the extension never finishing. Every command
  is now bounded (15s probes, 3min venv, 20min pip).

- `exec`'s default maxBuffer is 1MB. pip can overrun it, which kills the child
  with ENOBUFS and surfaces as a failed install that then retries on every
  launch. Raised to 32MB.

Also:
- `pip install --upgrade pip` ran unconditionally, costing a measured 468ms of
  network round-trip even when already satisfied. It now runs only right after
  venv creation, and its failure no longer aborts setup.
- Install success is confirmed against site-packages rather than pip's exit
  code, since a partially-installed wheel can still exit 0.
- Every command logs its duration, so a slow activation is diagnosable from the
  output channel instead of guessed at.
- Python candidates are newest-first (3.12 -> 3.11 -> 3.10 -> python3); macOS's
  bundled Xcode python3 is 3.9 and should only ever be the last resort.
- The first-run prompt gained "Don't ask again", remembered in globalState, so
  declining no longer re-prompts on every launch.

The warm path is now two stats with no process launches: 1ms against an
existing venv, where it previously could spawn Python to `import
faster_whisper` (measured at 1519ms cold, worse on Windows where ctranslate2
and onnxruntime DLLs are scanned).

Verified against harnesses covering each path: fast path, lock contention,
stale-lock reclaim, timeout enforcement, the attempt limit, and six
consecutive activations with a venv that never materialises. Not yet exercised
against a live extension host.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
28 diagrams across 14 features, one .drawio file each with a page per
level. Layer vocabulary and styling follow the "Our Proposed Approach"
architecture diagram in the research slides, so these read as zoom-ins
on it rather than a separate notation: Input Layer, VS Code Environment,
EchoCode Core, AI Services Layer, External Services, plus the standalone
Speech Output / TTS and Local Storage & State sinks.

Level 0 shows the feature as one numbered process (N.0) between its
sources and sinks, every flow labelled. Level 1 opens it into numbered
sub-processes (N.1 ... N.5) inside the EchoCode Core band, with the data
stores each one touches.

Generated rather than hand-drawn: features.js holds each feature's
processes, stores and flows; generate.js holds style and layout. A style
change reapplies to all 28 at once, which is what makes a convention
correction cheap.

Two layout rules that are load-bearing, both documented in the README
because dropping either brings back a specific annoyance:

- Every edge pins explicit exit and entry points. Without them draw.io
  picks perimeter points itself, so edges detour around shapes and
  converging flows stack on a single spot. Flows now fan across an edge.

- The EchoCode Core band is a real container and the step boxes are its
  children. A sibling band renders on top of the steps, since later cells
  draw over earlier ones, which meant sending it to the back by hand on
  every open. As a container it stays behind permanently and drags the
  pipeline with it.

Data stores sit along the bottom rather than in a fourth column, which
removes the crossings over the services column structurally instead of
routing around them.

Verified: XML parses, no dangling endpoints, all 241 edges pin both ends,
no overlapping shapes, nothing outside the page bounds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds render_svg.py, which reads geometry straight out of the .drawio
files and emits SVG. No draw.io CLI or headless browser is available
here, and reading the committed XML rather than recomputing layout means
the images cannot drift from the editable source.

Each feature page now carries both levels, respecting the page's
users-above-developers ordering: the Level 0 context diagram closes the
"For users" half, and the Level 1 decomposition opens "For developers".
Both link back to the .drawio for editing.

Two routing fixes in the generator, found by rasterising the output and
looking at it:

- Data store flows left through the bottom of their step, which meant
  dropping straight through the Core band and every step box below.
  They now exit sideways and fall through a clear corridor: the gap
  between the inputs column and the band on the left, or between the
  band and the services column on the right.

- Input flows all entered at the same height, so two feeding nearby
  steps ran their horizontal legs, and therefore their labels, on top of
  each other. Entry heights are now staggered.

The SVG renderer prefers horizontal segments for edge labels, since text
reads along them and two flows sharing a vertical corridor would
otherwise stack their labels.

Verified: all 56 image references resolve, 28 SVGs present, no
overlapping shapes, nothing outside the page bounds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Removes wiki/diagrams/ from version control and gitignores it: the
.drawio sources, generate.js and features.js, render_svg.py, and the
rendered SVGs. Files stay on disk; only the index changes.

Everything under it is reproducible from the two generators, so it is
regenerable rather than source of record:

  node wiki/diagrams/generate.js && python3 wiki/diagrams/render_svg.py

Also includes the .obsidian config entries that were already in the
working tree.

Note that the 14 feature wiki pages still reference diagrams/svg/*.svg,
so those 56 images will not resolve in a fresh clone. Left in place
deliberately rather than stripped, since the pages render correctly for
anyone who has generated the diagrams locally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The wiki pages embed 28 SVGs, so those have to resolve in a fresh clone
or every feature page shows broken images. The .drawio sources and the
two generators stay untracked — they are the inputs, reproducible on any
machine that wants to regenerate.

The ignore rules are shaped deliberately. Excluding "wiki/diagrams/"
outright makes the negations impossible, because git cannot re-include a
file whose parent directory is excluded; the contents are excluded and
svg/ re-included instead:

    wiki/diagrams/*
    !wiki/diagrams/svg/
    wiki/diagrams/svg/*
    !wiki/diagrams/svg/*.svg

Also delinks the 28 "Editable source" references on the feature pages.
They pointed at diagrams/*.drawio, which is no longer tracked, so as
links they would have 404'd for every reader. They now name the path as
plain text and point at the local README, keeping the information
without the dead link.

Verified: 145 markdown references across the wiki resolve, none missing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The per-feature Level 0/1 set is reference material: 28 diagrams, dense,
portrait, one per feature. Wrong shape for a projector. These are 16:9,
sparse, and sized to be read from the back of a room.

They also take the obvious compression. Fourteen features reduce to five
repeated shapes, so rather than showing each feature, S2 names each shape
once and lists its instances underneath:

  A  Direct read-out      6 features, no AI, works offline
  B  AI code analysis     4 features, local rules pass first
  C  External validation  2 features, pylint and g++
  D  Voice intake         3 modes over one pipeline
  E  Workspace mutation   3 features, never overwrites silently

The set:

  S1  Component Architecture  the five layers and what sits in each
  S2  Interaction Patterns    fourteen features, five shapes
  S3  Provider Abstraction    one funnel, three backends
  S4  Voice Pipeline          transcribed on-device
  S5  Output Path             everything reaches the user as speech

S1 is the one the deck's "Our Proposed Approach" slide expands into:
same layer vocabulary and palette, but each layer opened up to name its
components.

Only the rendered SVGs are tracked, matching the existing split — the
.drawio sources and slides.js stay local.

Regenerate with:
  node wiki/diagrams/slides.js && python3 wiki/diagrams/render_svg.py

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Blind and Low Vision" is the currently preferred term over "Blind and
Visually Impaired". Changes the actor label in both diagram generators
and regenerates all 33 affected files.

The literal acronym appeared in exactly two source lines; everything else
is derived from them. Note that a case-insensitive search for "bvi"
reports many false hits, because "webview" contains it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Lunar2074
Lunar2074 merged commit f3d1b83 into main Oct 1, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant