Skip to content

Repository files navigation

cripterAI

100% local AI meeting transcription and notes — not a single cloud call.

A desktop app (Electron) that captures your system audio, transcribes it with faster-whisper, tells the speakers apart, and drafts meeting notes with a local LLM via Ollama. Everything — audio, transcripts, models, summaries — stays on your machine.

What it does

  • System-audio capture — on Linux (PulseAudio/PipeWire) it enumerates microphones, sink monitors and per-application outputs and records any combination of them (parec + ffmpeg, multi-source amix mixing). On Windows a "System audio" source captures the default output in-app via WASAPI loopback. macOS system capture is not implemented yet; local audio files work everywhere.
  • Transcription with faster-whisper — a Python/FastAPI sidecar that probes CUDA properly (loads libcublas/libcudnn via ctypes before committing to GPU) and falls back cleanly to CPU/int8. One resident Whisper model with explicit eviction, so it can share a GPU with Ollama. Per-segment confidence and optional word timestamps.
  • Speaker diarization with no gated models — a custom pipeline: sliding windows → ECAPA-TDNN embeddings (SpeechBrain, Apache-2.0) → agglomerative clustering. No HuggingFace token, no license click-through, fully offline. A voiceprint registry accumulates speaker centroids and recognizes the same person across recordings (pyannote is available as an optional backend if you bring your own HF_TOKEN).
  • Meeting notes with a local LLM — summaries, action items and custom prompt templates via Ollama. The app auto-detects installed models (and other local OpenAI-compatible servers such as LM Studio or vLLM), recommends a model for your VRAM, and can install models from the UI with streamed download progress. Long transcripts get num_ctx: 8192 so Ollama doesn't silently truncate them.
  • Translation — batch transcript translation with a local LLM, with per-segment fallback.
  • Export — Markdown, TXT, HTML, JSON, DOCX and PDF, plus bulk ZIP export.
  • Library — SQLite + plain Markdown files under ~/.cripter-ai/ (your data outlives the app), with a persisted job queue (cancel/retry, AbortController) and live SSE progress streams end-to-end into the UI.
  • UI — Angular 21 standalone + NgRx SignalStore (zero .subscribe() in components), bilingual EN/ES, light/dark theming on a custom Tailwind v4 design system ("Aurora", 31 components — no Angular Material).

Workspace layout

apps/
  renderer/      Angular 21 standalone SPA (UI)
  desktop/       Electron main process (launches the bridge API; custom cripter:// scheme)
  bridge-api/    Express 5 local API (SQLite, jobs, channel/capture/speaker services)
  whisper-svc/   Python FastAPI (faster-whisper + local diarization)

packages/
  shared-types/  Cross-package TS interfaces & enums
  data-access/   Angular services, signal stores
  ui/            Aurora design-system standalone components (Tailwind v4)
  ui-styles/     Design tokens & theming (light/dark)
  utils/         Isomorphic utility functions

Nx monorepo, TypeScript strict, pnpm.

Getting started

Requirements: Node 22 LTS, pnpm, Python 3.11+, ffmpeg (plus pulseaudio-utils on Linux), optionally an NVIDIA GPU + Ollama for the AI features. Linux and Windows are supported for development; see Windows for the Windows specifics.

pnpm install
cp .env.example .env

# First time only — set up the Python venv for the transcription service:
pnpm nx install whisper-svc

# Run everything in dev (renderer + bridge-api + desktop + whisper-svc):
pnpm dev

Desktop packaging

pnpm build
pnpm nx pack desktop   # produces .deb and AppImage (electron-builder)

The packaged app bundles the renderer and the bridge API; whisper-svc runs as a separate service (e.g. systemd --user) so a single GPU-resident instance can serve the app.

Stop every dev service (and unload Ollama models) with pnpm stop, or pnpm stop:keep-ollama to leave Ollama running. pnpm ollama:start starts Ollama and pulls the models your saved agents use.

Windows

All dev scripts are plain Node, so everything runs from cmd or PowerShell — no bash or WSL needed.

Prerequisites

  • Node 22 LTS, with pnpm enabled via corepack enable.
  • Python 3.11+ from python.org with the py launcher (the default installer option).
  • ffmpeg: winget install Gyan.FFmpeg, or drop ffmpeg.exe and ffprobe.exe into tools/bin/ (nx stage desktop copies them into the packaged app's resources/bin).
  • Ollama for Windows for the LLM features.
  • Visual Studio Build Tools (C++ workload) only if better-sqlite3 has no prebuilt binary for your Node version.

Run

pnpm install
pnpm exec nx install whisper-svc
pnpm dev
pnpm stop

GPU transcription — faster-whisper needs the CUDA 12 cuBLAS and cuDNN 9 DLLs, either from a CUDA Toolkit 12 + cuDNN 9 install with their bin folders on PATH, or from the NVIDIA wheels inside the service's venv:

cd apps\whisper-svc
.venv\Scripts\pip install nvidia-cublas-cu12 nvidia-cudnn-cu12

Without them the service falls back to CPU/int8.

System-audio capture — add a System audio source to the channel. The app captures whatever plays on the default output device in-app via WASAPI loopback; no virtual audio cable is needed. The PulseAudio OS sources (sink monitors, per-application outputs) are Linux-only.

Packaging — run pnpm build then pnpm exec nx pack desktop on a Windows machine; the NSIS installer is written to dist/desktop-pack.

Privacy model

All processing is local by default: audio never leaves the machine, models run on your hardware, storage is local SQLite + Markdown. Cloud LLM providers (OpenAI/Anthropic) exist only as an explicit opt-in — you create such an agent yourself and bring your own key; nothing is configured out of the box.

License

MIT

About

100% local AI meeting transcription & notes: Electron + Angular 21 + Express 5 + FastAPI/faster-whisper, custom speaker diarization with no gated models, summaries via Ollama.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages