diff --git a/README.md b/README.md index 2638fbd..e5364b1 100644 --- a/README.md +++ b/README.md @@ -15,9 +15,9 @@ ![Trinity Assistant Banner](assets/banner.png) > [!NOTE] > **Aktuelle Highlights:** -> - **v0.17.3:** Eve stellt standardmaessig zwei parallele Realtime-Pipelines bereit. Dadurch kann die lokale Mac-/Windows-Stimme aktiv bleiben, waehrend ein iPhone oder iPad eine eigene Sprachsession nutzt. Der geschuetzte Voice-Port akzeptiert neben einem bewusst getrennten Voice-Token auch den vorhandenen Companion-Bridge-Token; bestehende Konfigurationen bleiben gueltig. -> - **v0.17.2:** `Soul.md` und `User.md` sind gegen leere Einstellungen abgesichert. Native UI, WebUI und Companion-Bridge lehnen leere Prompttexte ab, schreiben Änderungen atomar und legen vor echten Änderungen eine private Recovery-Kopie an. Mac- und Windows-Updates übernehmen nur nichtleere Promptdateien. Eve bestätigt eine Unterbrechung nun kurz und behält den akustischen Wiedergabe-Fingerabdruck, damit sie nach einer Unterbrechung spätere Antworten wieder vollständig spricht. -> - **v0.17.1:** Der lokale Eve-Desktoppfad öffnet Mikrofon und Lautsprecher jetzt als getrennte Audiostreams. Das verhindert insbesondere auf Macs mit unterschiedlichen Ein-/Ausgabe-Sampleraten einen blockierenden CoreAudio-Duplexstart. Erkannte echte Sprache leert Eves Ausgabe sofort und sendet zusätzlich ein lokales `response.cancel`; die serverseitige VAD bleibt als zweite Sicherung aktiv. +> - **v0.17.9:** Eine Trinity-Instanz in einer Windows-VM kann Eve-STT und Eve-TTS jetzt von einem privaten Ubuntu-NVIDIA-Host beziehen. Windows bleibt Control Plane mit Sessions, Memory, Agenten und UI; Ubuntu uebernimmt nur GPU-Inferenz. Dafuer gibt es die Profile `eve-windows-remote` und `eve-linux-gpu-server`, getrennte Tokens, Diagnosepruefungen und Installationsskripte. +> - **v0.17.8:** ClassicUI, iPhone und iPad zeigen nur noch einen zentralen Lautsprecherknopf. Aktiviert ein Client seine Ausgabe, wird er Bridge-weit zum einzigen Sprecher und alle anderen verbundenen Clients zeigen den durchgestrichenen Lautsprecher. Ein erneuter Klick auf dem aktiven Geraet schaltet Trinity ueberall stumm; auch Eves direkter Desktop- und Companion-Audiopfad folgt dieser Auswahl. +> - **v0.17.7:** Die Even G2 erhaelt einen kompakten, zwischengespeicherten Umgebungskontext: die neueste n-tv-Schlagzeile, Wetter und Temperatur der naechsten Stunde sowie optional den naechsten Kalendertermin. iPhone/iPad koennen Standort und Kalender einzeln und widerrufbar an die private Bridge melden; ohne Freigabe gilt der in der G2 gewaehlte Ersatzort. > - Die vollstaendige Historie steht in **[RELEASES.md](RELEASES.md)** und in den detaillierten **[Release Notes](docs/release_notes/)**. > [!IMPORTANT] diff --git a/RELEASES.md b/RELEASES.md index 81bb1ec..f7e30d8 100644 --- a/RELEASES.md +++ b/RELEASES.md @@ -6,6 +6,12 @@ Einzelnotizen liegen unter [docs/release_notes](docs/release_notes/). ## Aktuelle Highlights +- **v0.17.9:** Windows kann als Trinity-Control-Plane in einer VM laufen, waehrend ein privater Ubuntu-Host mit NVIDIA-GPU Parakeet-STT, Qwen3-TTS/Eve und das konfigurierte OpenAI-kompatible LLM bereitstellt. Die neuen Profile `eve-windows-remote` und `eve-linux-gpu-server`, getrennte Voice-/Core-Tokens, Remote-Diagnosen und Installationsskripte vermeiden unnoetiges GPU-Passthrough. Native Mac- und Windows-Eve-Profile sowie Legacy STT/TTS bleiben unveraendert waehlbar. +- **v0.17.8:** ClassicUI, iPhone und iPad verwenden einen einzigen Bridge-weiten Lautsprecherknopf. Das ausgewaehlte Geraet ist exklusiv; alle anderen Clients werden samt direkter Eve-Wiedergabe sofort stumm und zeigen den durchgestrichenen Lautsprecher. Ein zweiter Klick auf dem aktiven Geraet setzt die Ausgabe explizit auf `Stumm`, statt sie unbemerkt an einen anderen Client weiterzureichen. +- **v0.17.7:** Even G2 zeigt eine neue n-tv-Schlagzeile kurz als Laufhinweis sowie Wetter und Temperatur der naechsten Stunde im kompakten Statusfeld. Optional meldet der Companion den aktuellen Standort und den naechsten Kalendertermin an die private Bridge; beides bleibt nur im Arbeitsspeicher und wird beim Abschalten entfernt. Ohne Freigabe verwendet Trinity den in G2 konfigurierten Ersatzort. +- **v0.17.6:** Eine neue Bridge-weite Ein-Geraet-Auswahl legt eindeutig fest, wo Trinity spricht. ClassicUI, iPhone und iPad koennen mit `Ich spreche hier` die Ausgabe uebernehmen; die Auswahl wird in der Konfiguration persistiert und ueber `/speaker` sowie den normalen Ereignisstatus synchronisiert. Desktop-TTS bleibt still, wenn ein Companion-Client ausgewaehlt ist. Even G2 fragt das konkrete Ziel vor einem Auftrag ab und zeigt dessen Geraetenamen im HUD. +- **v0.17.5:** Even G2 verwirft im Zurufmodus lokal alle Transkripte ohne Wakeword, bevor Befehle, Checklisten oder Antworten angestossen werden. Der Buero-/Konversationsmodus bleibt bewusst direkt. G2 nutzt nach der Konfigurationsmigration das praezise STT-Profil; Checklistenbegriffe sind nicht mehr als Whisper-Hotwords hinterlegt. Companion-Ausgabe meldet deaktiviertes `Hoeren` sichtbar und versucht eine kurzfristig noch nicht verfuegbare Eve-Verbindung einmal erneut. +- **v0.17.4:** Even G2 kann in Vorlesungen als reines Mikrofon arbeiten. Das G2-Profil routet die Antwort wahlweise an den iPhone-/iPad-Companion mit aktivem `Hoeren`, an Trinity Desktop oder nur als Text an HUD und Chat. Die Quellkennung bleibt durch Bridge, STT-Feed und Antwort erhalten; Desktop-Ausgabe wird an die bereits laufende Eve-Realtime-Engine uebergeben und faellt nur ohne Eve kontrolliert auf die bisherige Systemstimme zurueck. `Vorlesung` reagiert ausschliesslich nach dem Wakeword, `Buero` direkt; der fruehere Chat-Betriebsmodus wird kompatibel nach Buero migriert. - **v0.17.3:** Eve Voice reserviert standardmaessig zwei parallele Realtime-Pipelines: eine fuer die lokale Desktop-Unterhaltung und eine fuer einen iPhone-/iPad-Companion. Der authentifizierte Voice-Proxy akzeptiert sowohl einen bewusst getrennten Voice-Token als auch den bereits konfigurierten Companion-Bridge-Token. Damit bleibt die bestehende Token-Trennung moeglich, waehrend normale Companion-Profile ohne doppelte Tokenpflege funktionieren. Die Companion-App v0.17.7 wartet zudem auf die echte Serverbestaetigung und beendet Wiederholungsversuche bei Authentifizierungs- oder Kapazitaetsfehlern kontrolliert. - **v0.17.2:** Private Persona-Dateien werden beim Speichern nicht mehr durch leere UI-Inhalte ersetzt. Native Einstellungen, WebUI und Companion-Bridge nutzen denselben validierten, atomaren Schreibpfad und sichern die vorherige nichtleere Fassung privat unter `TrinityRuntime/recovery/prompts`. Mac- und Windows-Installer sichern und restaurieren ausschließlich nichtleere `Soul.md` und `User.md`; neutrale Vorlagen greifen nur, wenn keine verwertbare private Datei existiert. Eves lokale Unterbrechung wartet auf vier aufeinanderfolgende Sprachblöcke, vergleicht zusätzlich das Frequenzspektrum und behält den jüngsten Wiedergabe-Fingerabdruck nach `response.cancel`. Dadurch bleibt echtes Barge-in schnell, ohne nach einer ersten Unterbrechung jede weitere Antwort selbst abzuschneiden. - **v0.17.1:** Lokale Eve-Unterhaltung auf Mac und Windows nutzt getrennte Mikrofon- und Lautsprecherstreams statt eines gekoppelten PortAudio-Duplexstreams. Das vermeidet blockierende Starts bei unterschiedlichen nativen Geräteraten und hält Ein- und Ausgabefehler voneinander isoliert. Sobald die Echo-Prüfung echte neue Sprache erkennt, leert Trinity die lokale Wiedergabe und sendet `response.cancel`, bevor die serverseitige VAD-Unterbrechung eintrifft. Legacy und Serverprofile bleiben unverändert verfügbar. diff --git a/core/ambient_context.py b/core/ambient_context.py new file mode 100644 index 0000000..24b44e2 --- /dev/null +++ b/core/ambient_context.py @@ -0,0 +1,219 @@ +"""Short, cached glance information for low-distraction companion displays.""" + +from __future__ import annotations + +import json +import threading +import time +import xml.etree.ElementTree as ET +from concurrent.futures import ThreadPoolExecutor +from email.utils import parsedate_to_datetime +from urllib.parse import urlencode +from urllib.request import Request, urlopen + + +NTV_RSS_URL = "https://www.n-tv.de/rss" +OPEN_METEO_GEOCODING_URL = "https://geocoding-api.open-meteo.com/v1/search" +OPEN_METEO_FORECAST_URL = "https://api.open-meteo.com/v1/forecast" + + +class AmbientContextService: + """Combine public glance data with ephemeral device context. + + Device coordinates and calendar titles deliberately remain in memory. The + server only caches public news/weather responses and never writes this + context into Trinity's synchronized vault. + """ + + def __init__(self, fetch=None, clock=None): + self._fetch = fetch or self._fetch_url + self._clock = clock or time.time + self._lock = threading.Lock() + self._cache = {} + self._device = {} + + def report_device(self, payload): + if not isinstance(payload, dict): + raise ValueError("Geraetekontext muss ein Objekt sein.") + now = self._clock() + latitude = _coordinate(payload.get("latitude"), -90, 90) + longitude = _coordinate(payload.get("longitude"), -180, 180) + calendar_title = _clean(payload.get("calendar_title"), 40) + try: + calendar_start = float(payload.get("calendar_start") or 0) + except (TypeError, ValueError): + calendar_start = 0.0 + with self._lock: + self._device = { + "latitude": latitude, + "longitude": longitude, + "calendar_title": calendar_title, + "calendar_start": calendar_start, + "updated_at": now, + } + return {"ok": True, "updated_at": now} + + def snapshot(self, fallback_place="Filderstadt"): + place = _clean(fallback_place, 80) or "Filderstadt" + now = self._clock() + with self._lock: + device = dict(self._device) + + coordinate = None + location_source = "fallback" + if ( + device.get("latitude") is not None + and device.get("longitude") is not None + and now - float(device.get("updated_at") or 0) <= 2 * 60 * 60 + ): + coordinate = (device["latitude"], device["longitude"]) + location_source = "device" + + with ThreadPoolExecutor(max_workers=2) as pool: + news_future = pool.submit(self._cached, "ntv", 300, self._latest_ntv) + if coordinate: + weather_key = f"weather:{coordinate[0]:.3f}:{coordinate[1]:.3f}" + weather_future = pool.submit( + self._cached, + weather_key, + 600, + lambda: self._weather_at(*coordinate, place="Aktueller Standort"), + ) + else: + weather_future = pool.submit( + self._cached, + f"weather-place:{place.casefold()}", + 600, + lambda: self._weather_for_place(place), + ) + try: + headline = news_future.result() + except Exception: + headline = {} + try: + weather = weather_future.result() + except Exception: + weather = {} + + calendar = {} + start = float(device.get("calendar_start") or 0) + title = str(device.get("calendar_title") or "") + if title and now - 300 <= start <= now + 24 * 60 * 60: + calendar = {"title": title, "start": start} + + return { + "ok": True, + "headline": headline, + "weather": {**weather, "source": location_source} if weather else {}, + "calendar": calendar, + "updated_at": now, + } + + def _cached(self, key, ttl, loader): + now = self._clock() + with self._lock: + cached = self._cache.get(key) + if cached and cached[0] > now: + return dict(cached[1]) + value = loader() + with self._lock: + self._cache[key] = (now + ttl, dict(value)) + return value + + def _latest_ntv(self): + root = ET.fromstring(self._fetch(NTV_RSS_URL)) + candidates = [] + for item in root.findall("./channel/item"): + title = _clean(item.findtext("title"), 220) + if not title: + continue + published = item.findtext("pubDate") or "" + try: + timestamp = parsedate_to_datetime(published).timestamp() + except (TypeError, ValueError, OverflowError): + timestamp = 0.0 + candidates.append((timestamp, title)) + if not candidates: + return {} + published_at, title = max(candidates, key=lambda item: item[0]) + return {"source": "n-tv", "title": title, "published_at": published_at} + + def _weather_for_place(self, place): + query = urlencode({"name": place, "count": 1, "language": "de", "format": "json"}) + payload = json.loads(self._fetch(f"{OPEN_METEO_GEOCODING_URL}?{query}")) + results = payload.get("results") or [] + if not results: + return {} + result = results[0] + return self._weather_at( + float(result["latitude"]), + float(result["longitude"]), + place=_clean(result.get("name"), 40) or place, + ) + + def _weather_at(self, latitude, longitude, place): + query = urlencode( + { + "latitude": latitude, + "longitude": longitude, + "hourly": "temperature_2m,weather_code", + "forecast_hours": 2, + "timezone": "auto", + } + ) + payload = json.loads(self._fetch(f"{OPEN_METEO_FORECAST_URL}?{query}")) + hourly = payload.get("hourly") or {} + temperatures = hourly.get("temperature_2m") or [] + codes = hourly.get("weather_code") or [] + if not temperatures: + return {} + index = 1 if len(temperatures) > 1 else 0 + code = int(codes[index] if index < len(codes) else 0) + return { + "place": place, + "temperature": round(float(temperatures[index])), + "code": code, + "symbol": _weather_symbol(code), + } + + @staticmethod + def _fetch_url(url): + request = Request(url, headers={"User-Agent": "Trinity-Assistant/0.17"}) + with urlopen(request, timeout=3.5) as response: + return response.read() + + +def _coordinate(value, minimum, maximum): + if value in (None, ""): + return None + try: + number = float(value) + except (TypeError, ValueError) as exc: + raise ValueError("Ungueltige Standortkoordinate.") from exc + if not minimum <= number <= maximum: + raise ValueError("Standortkoordinate ausserhalb des gueltigen Bereichs.") + return number + + +def _clean(value, limit): + return " ".join(str(value or "").split())[:limit] + + +def _weather_symbol(code): + if code == 0: + return "SUN" + if code in {1, 2}: + return "PART" + if code == 3: + return "CLOUD" + if code in {45, 48}: + return "FOG" + if code in {51, 53, 55, 56, 57}: + return "DRIZ" + if code in {61, 63, 65, 66, 67, 80, 81, 82}: + return "RAIN" + if code in {71, 73, 75, 77, 85, 86}: + return "SNOW" + if code in {95, 96, 99}: + return "STORM" + return "WX" diff --git a/core/bridge_audio.py b/core/bridge_audio.py index 627481a..0bcf263 100644 --- a/core/bridge_audio.py +++ b/core/bridge_audio.py @@ -11,13 +11,13 @@ MAX_AUDIO_SECONDS = 20 TRINITY_VOCABULARY = ( "Trinity. Deutsche Vorlesung und natuerliche Konversation. " - "Wichtige Begriffe: Trinity, Hilfe, Modus Zuruf, Zurufmodus, " - "Konversationsmodus, Wakeword, Checkliste, Stichwoerter, Arbeitsraum, " + "Kommandos: Trinity, Hilfe, Modus Zuruf, Zurufmodus, " + "Konversationsmodus, Wakeword, Arbeitsraum, " "Schnellsession, Nash-Gleichgewicht, Gefangenendilemma und Kooperation." ) TRINITY_HOTWORDS = ( - "Trinity Hilfe Zuruf Zurufmodus Konversationsmodus Wakeword Checkliste " - "Stichwoerter Arbeitsraum Schnellsession Nash-Gleichgewicht " + "Trinity Hilfe Zuruf Zurufmodus Konversationsmodus Wakeword " + "Arbeitsraum Schnellsession Nash-Gleichgewicht " "Gefangenendilemma Kooperation" ) diff --git a/core/config.json.example b/core/config.json.example index 4057e28..b6035b8 100644 --- a/core/config.json.example +++ b/core/config.json.example @@ -47,6 +47,13 @@ "fallback_to_legacy": true, "language_policy": "de_only", "access_token": "", + "backend_host": "127.0.0.1", + "backend_port": 18767, + "backend_token": "", + "remote_voice_url": "", + "remote_voice_token": "", + "remote_core_base_url": "", + "remote_core_api_key": "", "reference_audio": "", "reference_text": "Herzlich willkommen lieber Schülerinnen und Schüler des Max Planck Gymnasiums. Schön, dass Ihr heute hier seid. Unser Schülervortrag dreht sich um das Thema KI und Wir. Lasst uns gemeinsam entdecken wie künstliche Intelligenz unseren Alltag verändert.", "stt_model": "mlx-community/parakeet-tdt-0.6b-v3", @@ -104,6 +111,34 @@ "stt_model": "nvidia/parakeet-tdt-0.6b-v3", "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", "tts_backend": "torch" + }, + "eve-linux-gpu-server": { + "mode": "realtime", + "device": "cuda", + "runtime_role": "server", + "conversation_backend": "remote", + "bind_host": "0.0.0.0", + "public_port": 8766, + "internal_port": 18766, + "local_audio": false, + "num_pipelines": 2, + "stt_model": "nvidia/parakeet-tdt-0.6b-v3", + "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", + "tts_backend": "torch" + }, + "eve-windows-remote": { + "mode": "realtime", + "device": "cpu", + "runtime_role": "client", + "conversation_backend": "trinity", + "bind_host": "127.0.0.1", + "public_port": 8766, + "internal_port": 18766, + "local_audio": true, + "num_pipelines": 1, + "stt_model": "nvidia/parakeet-tdt-0.6b-v3", + "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", + "tts_backend": "torch" } } }, diff --git a/core/remote_client.py b/core/remote_client.py index 08fce49..316e739 100644 --- a/core/remote_client.py +++ b/core/remote_client.py @@ -56,6 +56,18 @@ def latest_payload(self): def get_runtime(self): return self._request("/runtime", method="GET") + def get_speaker(self): + return self._request("/speaker", method="GET") + + def set_speaker(self, device_id, label, kind="desktop"): + return self._request( + "/speaker", + {"device_id": device_id, "label": label, "kind": kind}, + ) + + def release_speaker(self): + return self.set_speaker("none", "Stumm", kind="none") + def current_session(self): return self._request("/session/current", method="GET") diff --git a/core/settings_ui.py b/core/settings_ui.py index e9eb914..cbd10e8 100644 --- a/core/settings_ui.py +++ b/core/settings_ui.py @@ -130,6 +130,13 @@ def save_config(self, _checked=False, *, show_confirmation=True): voice["fallback_to_legacy"] = self.voice_fallback_cb.isChecked() voice["reference_audio"] = self.voice_reference_edit.text().strip() voice["access_token"] = self.voice_token_edit.text().strip() + voice["remote_voice_url"] = self.voice_remote_url_edit.text().strip() + voice["remote_voice_token"] = self.voice_token_edit.text().strip() + voice["backend_host"] = self.voice_backend_host_edit.text().strip() or "127.0.0.1" + voice["backend_port"] = self.voice_backend_port_spin.value() + voice["backend_token"] = self.voice_backend_token_edit.text().strip() + voice["remote_core_base_url"] = self.voice_remote_core_url_edit.text().strip() + voice["remote_core_api_key"] = self.voice_remote_core_token_edit.text().strip() voice["streaming_chunk_size"] = self.voice_chunk_size_spin.value() voice["audio_prebuffer_ms"] = self.voice_prebuffer_spin.value() voice["barge_in_enabled"] = self.voice_barge_in_cb.isChecked() @@ -2073,11 +2080,12 @@ def _create_system_tab(self): ui_modes = resolve_ui_modes(system_conf) self.mode_combo = QComboBox() - self.mode_combo.addItems(["office", "lecture", "chat"]) - self.mode_combo.setCurrentText(system_conf.get("mode", "office")) + self.mode_combo.addItems(["office", "lecture"]) + saved_mode = system_conf.get("mode", "office") + self.mode_combo.setCurrentText("office" if saved_mode == "chat" else saved_mode) form.addRow("Trinity Modus:", self.mode_combo) - mode_hint = QLabel("Office: Standard (STT+TTS an).
Lecture: Vorlesung optimiert.
Chat: STT+TTS aus (nur Flüstern/Telegram).") + mode_hint = QLabel("Büro: direkte Konversation.
Vorlesung: Antworten nur nach dem Wakeword.") mode_hint.setStyleSheet("color: #888; font-size: 11px;") mode_hint.setWordWrap(True) form.addRow("", mode_hint) @@ -2635,6 +2643,19 @@ def _test_comfyui_connection(self): "Wie das Mac-Serverprofil, aber für eine Windows-Workstation mit CUDA. " "Dieses Profil auf einem Mac nicht für den normalen Betrieb auswählen.", ), + ( + "Windows-VM mit Eve auf einem Ubuntu-GPU-Host", + "eve-windows-remote", + "Windows bleibt Trinity-Kern und Desktop-Oberfläche. Mikrofon und Lautsprecher " + "verbinden sich mit Eve auf dem Ubuntu-Host; Antworten laufen zur Wahrung von " + "Sessions, Memory, RAG und Agenten zurück durch diese Windows-Instanz.", + ), + ( + "Ubuntu mit NVIDIA als Eve-Sprachserver", + "eve-linux-gpu-server", + "Ubuntu führt Parakeet-STT und Qwen3-TTS auf CUDA aus. Der erkannte Text wird " + "an den geschützten Trinity-Core-Endpunkt der Windows-VM weitergereicht.", + ), ( "Diagnose: Ornith direkt, ohne Trinity", "eve-direct-ornith", @@ -2661,17 +2682,32 @@ def _load_voice_profile_form(self, _selection=None): "Benutzerdefiniertes Eve-Profil.", ) self.voice_profile_description.setText(description) - realtime = profile_name in {"eve-mac-server", "eve-windows-server"} + realtime = profile_name in {"eve-mac-server", "eve-windows-server", "eve-linux-gpu-server"} + remote_client = profile_name == "eve-windows-remote" + remote_server = profile_name == "eve-linux-gpu-server" for field in ( self.voice_bind_host_edit, self.voice_public_port_spin, - self.voice_token_edit, self.voice_realtime_hint, ): field.setVisible(realtime) label = self.voice_form.labelForField(field) if label is not None: label.setVisible(realtime) + for field, visible in ( + (self.voice_token_edit, realtime or remote_client), + (self.voice_remote_url_edit, remote_client), + (self.voice_backend_host_edit, remote_client), + (self.voice_backend_port_spin, remote_client), + (self.voice_backend_token_edit, remote_client), + (self.voice_remote_core_url_edit, remote_server), + (self.voice_remote_core_token_edit, remote_server), + (self.voice_remote_hint, remote_client or remote_server), + ): + field.setVisible(visible) + label = self.voice_form.labelForField(field) + if label is not None: + label.setVisible(visible) def _create_stt_tts_tab(self): widget = QWidget() @@ -2764,6 +2800,42 @@ def _create_stt_tts_tab(self): self.voice_token_edit.setPlaceholderText("Erforderlich bei 0.0.0.0 / Tailscale") voice_form.addRow("Realtime Token:", self.voice_token_edit) + self.voice_remote_url_edit = QLineEdit(str(voice_conf.get("remote_voice_url") or "")) + self.voice_remote_url_edit.setPlaceholderText("ws://UBUNTU-TAILSCALE-IP:8766/v1/realtime") + voice_form.addRow("Ubuntu Voice URL:", self.voice_remote_url_edit) + + self.voice_backend_host_edit = QLineEdit(str(voice_conf.get("backend_host") or "127.0.0.1")) + self.voice_backend_host_edit.setPlaceholderText("0.0.0.0") + voice_form.addRow("Windows Core Bind-Adresse:", self.voice_backend_host_edit) + + self.voice_backend_port_spin = QSpinBox() + self.voice_backend_port_spin.setRange(1024, 65535) + self.voice_backend_port_spin.setValue(int(voice_conf.get("backend_port") or 18767)) + voice_form.addRow("Windows Core Port:", self.voice_backend_port_spin) + + self.voice_backend_token_edit = QLineEdit(str(voice_conf.get("backend_token") or "")) + self.voice_backend_token_edit.setEchoMode(QLineEdit.Password) + self.voice_backend_token_edit.setPlaceholderText("Separater langer Rückkanal-Token") + voice_form.addRow("Windows Core Token:", self.voice_backend_token_edit) + + self.voice_remote_core_url_edit = QLineEdit(str(voice_conf.get("remote_core_base_url") or "")) + self.voice_remote_core_url_edit.setPlaceholderText("http://WINDOWS-TAILSCALE-IP:18767/v1") + voice_form.addRow("Windows Trinity Core URL:", self.voice_remote_core_url_edit) + + self.voice_remote_core_token_edit = QLineEdit(str(voice_conf.get("remote_core_api_key") or "")) + self.voice_remote_core_token_edit.setEchoMode(QLineEdit.Password) + self.voice_remote_core_token_edit.setPlaceholderText("Gleich wie Windows Core Token") + voice_form.addRow("Windows Trinity Core Token:", self.voice_remote_core_token_edit) + + self.voice_remote_hint = QLabel( + "Ubuntu behält die NVIDIA-GPU. Windows bleibt die kanonische Trinity mit " + "Sessions, Memory und Agenten. Beide Rechner dürfen diese Ports nur im " + "privaten LAN oder Tailnet freigeben, niemals am öffentlichen Router." + ) + self.voice_remote_hint.setWordWrap(True) + self.voice_remote_hint.setStyleSheet("color: #d29922; font-size: 11px;") + voice_form.addRow("", self.voice_remote_hint) + self.voice_realtime_hint = QLabel( "Diese drei Felder werden nur benötigt, wenn iPhone oder iPad Eve über " "den Mac beziehungsweise Windows-PC nutzen. Für Tailscale: 0.0.0.0, " diff --git a/core/transcriber.py b/core/transcriber.py index 39c2b44..0997dfe 100644 --- a/core/transcriber.py +++ b/core/transcriber.py @@ -2,6 +2,7 @@ os.environ["OMP_NUM_THREADS"] = "8" os.environ["KMP_BLOCKTIME"] = "1" import html as html_lib +import json import time import queue import sys @@ -315,8 +316,11 @@ def load_config(self): # System Config self.system_cfg = config.get("system", {}) self.mode = self.system_cfg.get("mode", "office") + if self.mode == "chat": + self.mode = "office" self.microphone_enabled = self.system_cfg.get("microphone_enabled", True) self.tts_enabled = self.system_cfg.get("tts_enabled", True) + self.speech_output = self.system_cfg.get("speech_output", {}) self.speech_input_enabled = ( sys.platform != "win32" or self.system_cfg.get("windows_speech_enabled", False) @@ -341,6 +345,7 @@ def load_config(self): self.mode = "office" self.microphone_enabled = True self.tts_enabled = True + self.speech_output = {} self.speech_input_enabled = ( sys.platform != "win32" and os.environ.get("TRINITY_SERVER") != "1" @@ -1090,7 +1095,7 @@ def _process_external_stt_feed(self): speak = bool(event.get("speak", False)) source = str(event.get("source") or "ios-stt").strip().lower() marker = "final" if is_final else "live" - source_label = "G2-STT" if source == "g2-stt" else "iPhone-STT" + source_label = "G2-STT" if source.startswith("g2-stt") else "iPhone-STT" print(f"📱 {source_label} ({marker}): {text}") if not is_final: continue @@ -1100,7 +1105,7 @@ def _process_external_stt_feed(self): self.recent_chunks.clear() continue - if getattr(self, "mode", "office") == "chat" and not has_trigger( + if getattr(self, "mode", "office") in {"office", "chat"} and not has_trigger( text, self.trigger_variants, ): @@ -1120,7 +1125,7 @@ def _process_external_stt_feed(self): continue request = None - if source in {"ios-stt", "g2-stt"}: + if source == "ios-stt" or source.startswith("g2-stt"): request = { "request_id": event.get("event_id"), "source": source, @@ -1192,7 +1197,21 @@ def _speak_thread(self, text): if not getattr(self, "tts_enabled", True): set_state("idle") return + speech_output = getattr(self, "speech_output", {}) + if isinstance(speech_output, dict) and speech_output.get("device_id"): + if str(speech_output.get("kind") or "desktop") != "desktop": + print( + "🔇 Desktop-TTS bleibt stumm; aktive Sprechstelle: " + f"{speech_output.get('label') or speech_output.get('device_id')}" + ) + set_state("idle") + return set_state("speaking") + + if self._queue_eve_desktop_speech(text): + print(f"🔊 Eve spricht auf diesem Desktop: {text[:60]}...") + set_state("idle") + return target_device = self.audio_routing.get("private_device", "Standard") if "[SPEAKER]" in text: @@ -1217,6 +1236,30 @@ def _speak_thread(self, text): if self.speak_process and self.speak_process.returncode == 0: set_state("idle") + def _queue_eve_desktop_speech(self, text): + """Hand external TTS to the active local Eve realtime audio client.""" + ready_path = os.path.join( + PROJECT_DIR, "TrinityRuntime", "voice", "desktop_eve_audio.ready" + ) + if not os.path.isfile(ready_path): + return False + queue_path = os.path.join( + PROJECT_DIR, "TrinityRuntime", "voice", "desktop_speech_queue.jsonl" + ) + try: + os.makedirs(os.path.dirname(queue_path), exist_ok=True) + payload = {"text": str(text or "").strip(), "timestamp": time.time()} + if not payload["text"]: + return False + with open(queue_path, "a", encoding="utf-8") as handle: + handle.write(json.dumps(payload, ensure_ascii=False) + "\n") + handle.flush() + os.fsync(handle.fileno()) + return True + except OSError as exc: + print(f"⚠️ Eve-Ausgabequeue nicht verfügbar: {exc}") + return False + def switch_mode(self, new_mode): old_mode = getattr(self, 'mode', 'office') if old_mode == new_mode: return diff --git a/core/trinity_bridge.py b/core/trinity_bridge.py index 7d8d50a..486a881 100644 --- a/core/trinity_bridge.py +++ b/core/trinity_bridge.py @@ -8,6 +8,7 @@ import json import mimetypes import os +import platform import re import shutil import sqlite3 @@ -22,6 +23,7 @@ from urllib.request import url2pathname from agent_catalog import build_agent_catalog, normalize_catalog_overrides +from ambient_context import AmbientContextService from bridge_audio import BridgeAudioTranscriber from chat_attachments import attachment_kind from chat_protocol import ( @@ -73,6 +75,8 @@ "/workspaces", "/dashboard", "/memory/graph", + "/speaker", + "/ambient", } @@ -166,6 +170,7 @@ def __init__(self, home, token="", auth_enabled=False): self._summary_jobs = set() self._audio_transcriber = None self._canvas_status_cache = (0.0, {}) + self.ambient = AmbientContextService() self.sessions = UnifiedSessionStore(self.home, load_config(self.config_path)) self.workbench = WorkbenchManager(self.home) @@ -193,6 +198,7 @@ def instance_state(self): return { "profile": paths.profile, "session": self.current_session(), + "speaker": self.get_speaker(), "knowledge": { "vault_root": str(paths.vault_root), "vault_available": paths.vault_root.is_dir(), @@ -203,6 +209,59 @@ def instance_state(self): }, } + def _default_speaker(self): + hostname = platform.node().strip() or "Desktop" + return { + "device_id": f"desktop:{self.profile.lower()}:{hostname}", + "label": f"Trinity Desktop · {hostname}", + "kind": "desktop", + "updated_at": 0.0, + } + + def get_speaker(self): + """Return the one device currently allowed to render Trinity speech.""" + config = load_config(self.config_path) + value = config.get("system", {}).get("speech_output") + if not isinstance(value, dict): + value = self._default_speaker() + kind = str(value.get("kind") or "").strip().lower() + device_id = str(value.get("device_id") or "").strip() + if kind not in {"desktop", "companion", "none"} or not device_id: + value = self._default_speaker() + return { + "device_id": str(value.get("device_id") or "")[:160], + "label": str(value.get("label") or "Trinity-Ausgabe")[:100], + "kind": str(value.get("kind") or "desktop"), + "updated_at": float(value.get("updated_at") or 0.0), + } + + def set_speaker(self, payload): + if not isinstance(payload, dict): + raise ValueError("Sprechstelle muss ein Objekt sein.") + kind = str(payload.get("kind") or "").strip().lower() + device_id = str(payload.get("device_id") or "").strip()[:160] + label = str(payload.get("label") or "").strip()[:100] + if kind not in {"desktop", "companion", "none"}: + raise ValueError("Sprechstelle muss desktop, companion oder none sein.") + if kind == "none": + device_id = "none" + label = "Stumm" + if not device_id: + raise ValueError("Eine Geräte-ID wird benötigt.") + if not label: + label = "Trinity Desktop" if kind == "desktop" else "Trinity Companion" + speaker = { + "device_id": device_id, + "label": label, + "kind": kind, + "updated_at": time.time(), + } + with self._lock: + config = load_config(self.config_path) + config.setdefault("system", {})["speech_output"] = speaker + save_config(self.config_path, config) + return {"ok": True, **speaker} + def validate_client_profile(self, value): expected = str(value or "").strip().upper() if expected and expected != self.profile: @@ -311,7 +370,7 @@ def send_stt(self, payload, user=None): raise ValueError("STT-Text darf nicht leer sein.") is_final = bool(payload.get("is_final", False)) source = str(payload.get("source") or "ios-stt").strip().lower() - if source not in {"ios-stt", "g2-stt"}: + if source != "ios-stt" and not source.startswith("g2-stt"): source = "ios-stt" canonical = self.sessions.canonicalize(payload, source=source) event = append_external_stt_event( @@ -696,14 +755,18 @@ def _run_session_summary_job( def get_mode(self): config = self._read_config() mode = str(config.get("system", {}).get("mode", "office") or "office") - if mode not in {"office", "lecture", "chat"}: + if mode == "chat": + mode = "office" + if mode not in {"office", "lecture"}: mode = "office" return {"ok": True, "mode": mode} def set_mode(self, payload): mode = str(payload.get("mode", "")).strip().lower() - if mode not in {"office", "lecture", "chat"}: - raise ValueError("Modus muss office, lecture oder chat sein.") + if mode == "chat": + mode = "office" + if mode not in {"office", "lecture"}: + raise ValueError("Modus muss office oder lecture sein.") with self._lock: config = self._read_config() config.setdefault("system", {})["mode"] = mode @@ -717,9 +780,14 @@ def set_mode(self, payload): def get_runtime(self): config = load_config(self.config_path) system = config.get("system", {}) + runtime_mode = str(system.get("mode", "lecture") or "lecture").strip().lower() + if runtime_mode == "chat": + runtime_mode = "office" + if runtime_mode not in {"office", "lecture"}: + runtime_mode = "lecture" return { "ok": True, - "mode": str(system.get("mode", "lecture") or "lecture"), + "mode": runtime_mode, "microphone_enabled": bool(system.get("microphone_enabled", True)), "audio_capture_mode": str(system.get("audio_capture_mode", "mic_only") or "mic_only"), "tts_enabled": bool(system.get("tts_enabled", True)), @@ -1089,8 +1157,10 @@ def set_runtime(self, payload): system = config.setdefault("system", {}) if "mode" in payload: mode = str(payload["mode"] or "").strip().lower() - if mode not in {"office", "lecture", "chat"}: - raise ValueError("Modus muss office, lecture oder chat sein.") + if mode == "chat": + mode = "office" + if mode not in {"office", "lecture"}: + raise ValueError("Modus muss office oder lecture sein.") system["mode"] = mode for key in ("microphone_enabled", "tts_enabled"): if key in payload: @@ -1770,6 +1840,14 @@ def do_GET(self): # noqa: N802 _json_response(self, 200, {"ok": True, **bridge.latest_bubble()}) elif parsed.path == "/mode": _json_response(self, 200, bridge.get_mode()) + elif parsed.path == "/speaker": + _json_response(self, 200, {"ok": True, **bridge.get_speaker()}) + elif parsed.path == "/ambient": + _json_response( + self, + 200, + bridge.ambient.snapshot(query.get("place", ["Filderstadt"])[0]), + ) elif parsed.path == "/runtime": _json_response(self, 200, bridge.get_runtime()) elif parsed.path == "/dashboard": @@ -2005,6 +2083,10 @@ def do_POST(self): # noqa: N802 _json_response(self, 200, bridge.import_offline_events(_read_json(self), user=user)) elif parsed.path == "/mode": _json_response(self, 200, bridge.set_mode(_read_json(self))) + elif parsed.path == "/speaker": + _json_response(self, 200, bridge.set_speaker(_read_json(self))) + elif parsed.path == "/ambient/device": + _json_response(self, 200, bridge.ambient.report_device(_read_json(self))) elif parsed.path == "/runtime": _json_response(self, 200, bridge.set_runtime(_read_json(self))) elif parsed.path == "/agent/update": diff --git a/core/voice/command_builder.py b/core/voice/command_builder.py index 419be6f..af56b11 100644 --- a/core/voice/command_builder.py +++ b/core/voice/command_builder.py @@ -21,6 +21,15 @@ def build_speech_to_speech_command(config: VoiceConfig) -> list[str]: profile = config.profile command = _entrypoint(config) + if profile.conversation_backend == "trinity": + conversation_base_url = f"http://{config.backend_host}:{config.backend_port}/v1" + conversation_api_key = config.backend_token + elif profile.conversation_backend == "remote": + conversation_base_url = config.remote_core_base_url.rstrip("/") + conversation_api_key = config.remote_core_api_key + else: + conversation_base_url = config.direct_llm_base_url.rstrip("/") + conversation_api_key = config.direct_llm_api_key or "local" command.extend([ "--mode", profile.mode, "--device", profile.device, @@ -35,15 +44,9 @@ def build_speech_to_speech_command(config: VoiceConfig) -> list[str]: "--min_silence_ms", "96", "--speech_pad_ms", "320", "--llm_backend", "chat-completions", - "--model_name", "trinity-core" if profile.conversation_backend == "trinity" else config.direct_llm_model, - "--responses_api_base_url", - ( - f"http://{config.backend_host}:{config.backend_port}/v1" - if profile.conversation_backend == "trinity" - else config.direct_llm_base_url.rstrip("/") - ), - "--responses_api_api_key", - config.backend_token if profile.conversation_backend == "trinity" else (config.direct_llm_api_key or "local"), + "--model_name", "trinity-core" if profile.conversation_backend in {"trinity", "remote"} else config.direct_llm_model, + "--responses_api_base_url", conversation_base_url, + "--responses_api_api_key", conversation_api_key, "--responses_api_stream", "true", "--responses_api_disable_thinking", "true", "--init_chat_prompt", diff --git a/core/voice/config.py b/core/voice/config.py index 852d801..7214619 100644 --- a/core/voice/config.py +++ b/core/voice/config.py @@ -71,6 +71,34 @@ "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", "tts_backend": "torch", }, + "eve-linux-gpu-server": { + "mode": "realtime", + "device": "cuda", + "runtime_role": "server", + "conversation_backend": "remote", + "bind_host": "0.0.0.0", + "public_port": 8766, + "internal_port": 18766, + "local_audio": False, + "num_pipelines": 2, + "stt_model": "nvidia/parakeet-tdt-0.6b-v3", + "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", + "tts_backend": "torch", + }, + "eve-windows-remote": { + "mode": "realtime", + "device": "cpu", + "runtime_role": "client", + "conversation_backend": "trinity", + "bind_host": "127.0.0.1", + "public_port": 8766, + "internal_port": 18766, + "local_audio": True, + "num_pipelines": 1, + "stt_model": "nvidia/parakeet-tdt-0.6b-v3", + "tts_model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base", + "tts_backend": "torch", + }, "eve-direct-ornith": { "mode": "local", "device": "mps", @@ -118,6 +146,7 @@ class VoiceProfile: name: str mode: str device: str + runtime_role: str conversation_backend: str bind_host: str public_port: int @@ -141,6 +170,10 @@ class VoiceConfig: backend_host: str = "127.0.0.1" backend_port: int = 18767 backend_token: str = field(default_factory=lambda: secrets.token_urlsafe(24)) + remote_core_base_url: str = "" + remote_core_api_key: str = "" + remote_voice_url: str = "" + remote_voice_token: str = "" stt_model: str = "mlx-community/parakeet-tdt-0.6b-v3" tts_model: str = "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-6bit" reference_audio: Path = field(default_factory=Path) @@ -176,6 +209,7 @@ def profile(self) -> VoiceProfile: name=self.profile_name, mode=str(raw.get("mode") or "local"), device=device, + runtime_role=str(raw.get("runtime_role") or "integrated"), conversation_backend=str(raw.get("conversation_backend") or "trinity"), bind_host=str(raw.get("bind_host") or "127.0.0.1"), public_port=int(raw.get("public_port") or 8766), @@ -196,15 +230,32 @@ def validate(self) -> list[str]: errors.append("Aktuell wird nur voice.language_policy=de_only unterstützt.") if profile.mode not in {"local", "realtime"}: errors.append(f"Unbekannter Voice-Modus: {profile.mode}") - if profile.conversation_backend not in {"trinity", "direct"}: + if profile.runtime_role not in {"integrated", "server", "client"}: + errors.append(f"Unbekannte Voice-Rolle: {profile.runtime_role}") + if profile.conversation_backend not in {"trinity", "direct", "remote"}: errors.append(f"Unbekanntes Conversation-Backend: {profile.conversation_backend}") if ( - profile.mode == "realtime" + profile.runtime_role != "client" + and profile.mode == "realtime" and not _loopback(profile.bind_host) and not (self.access_token or self.companion_access_token) ): errors.append("Ein extern gebundener Voice-Server braucht voice.access_token.") - if self.enabled and not self.reference_audio.is_file(): + if profile.runtime_role == "client" and not self.remote_voice_url.strip(): + errors.append("Das Windows-Remoteprofil braucht voice.remote_voice_url.") + if profile.runtime_role == "client" and not (self.remote_voice_token or self.access_token): + errors.append("Das Windows-Remoteprofil braucht einen Voice-Token.") + if profile.conversation_backend == "remote" and not self.remote_core_base_url.strip(): + errors.append("Der Ubuntu-Voice-Server braucht voice.remote_core_base_url.") + if profile.conversation_backend == "remote" and not self.remote_core_api_key.strip(): + errors.append("Der Ubuntu-Voice-Server braucht voice.remote_core_api_key.") + if ( + profile.conversation_backend == "trinity" + and not _loopback(self.backend_host) + and not self.backend_token + ): + errors.append("Ein extern gebundener Trinity-Core-Backend braucht voice.backend_token.") + if self.enabled and profile.runtime_role != "client" and not self.reference_audio.is_file(): errors.append(f"Eve-Referenzaudio fehlt: {self.reference_audio or '(nicht konfiguriert)'}") if self.enabled and not self.reference_text.strip(): errors.append("Eve-Referenztranskript fehlt.") @@ -228,6 +279,11 @@ def default_voice_config() -> dict[str, Any]: "access_token": "", "backend_host": "127.0.0.1", "backend_port": 18767, + "backend_token": "", + "remote_core_base_url": "", + "remote_core_api_key": "", + "remote_voice_url": "", + "remote_voice_token": "", "stt_model": "mlx-community/parakeet-tdt-0.6b-v3", "tts_model": "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-6bit", "reference_audio": "", @@ -311,6 +367,31 @@ def load_voice_config(home: str | Path, config: dict[str, Any], profile_name: st ), backend_host=str(raw.get("backend_host") or "127.0.0.1"), backend_port=int(raw.get("backend_port") or 18767), + backend_token=str( + os.environ.get("TRINITY_VOICE_BACKEND_TOKEN") + or raw.get("backend_token") + or "" + ), + remote_core_base_url=str( + os.environ.get("TRINITY_REMOTE_CORE_BASE_URL") + or raw.get("remote_core_base_url") + or "" + ), + remote_core_api_key=str( + os.environ.get("TRINITY_REMOTE_CORE_API_KEY") + or raw.get("remote_core_api_key") + or "" + ), + remote_voice_url=str( + os.environ.get("TRINITY_REMOTE_VOICE_URL") + or raw.get("remote_voice_url") + or "" + ), + remote_voice_token=str( + os.environ.get("TRINITY_REMOTE_VOICE_TOKEN") + or raw.get("remote_voice_token") + or "" + ), stt_model=str(raw.get("stt_model") or "mlx-community/parakeet-tdt-0.6b-v3"), tts_model=str(raw.get("tts_model") or "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-6bit"), reference_audio=_expand_path(configured_audio, root), diff --git a/core/voice/conversation/trinity_backend.py b/core/voice/conversation/trinity_backend.py index 4a07a6e..e25576e 100644 --- a/core/voice/conversation/trinity_backend.py +++ b/core/voice/conversation/trinity_backend.py @@ -6,6 +6,7 @@ import re import threading import time +import unicodedata import uuid from collections.abc import Iterable from http import HTTPStatus @@ -17,6 +18,78 @@ from ..language_policy import enforce_input_language, segment_for_speech +DEFAULT_WAKEWORD_VARIANTS = ( + "trinity", + "triniti", + "trindy", + "trinnity", + "trinitiy", + "trinitys", + "trinitie", + "drinity", + "trinidi", + "trenty", + "trendy", +) + + +def _normalize_wakeword_text(value: Any) -> str: + decomposed = unicodedata.normalize("NFKD", str(value or "").casefold()) + asciiish = "".join(char for char in decomposed if not unicodedata.combining(char)) + return re.sub(r"[^a-z0-9]+", " ", asciiish.replace("ß", "ss")).strip() + + +def _bounded_levenshtein(left: str, right: str, max_distance: int) -> int: + if left == right: + return 0 + if abs(len(left) - len(right)) > max_distance: + return max_distance + 1 + previous = list(range(len(right) + 1)) + for row, left_char in enumerate(left, start=1): + current = [row] + row_min = row + for column, right_char in enumerate(right, start=1): + current.append(min( + current[column - 1] + 1, + previous[column] + 1, + previous[column - 1] + (left_char != right_char), + )) + row_min = min(row_min, current[-1]) + if row_min > max_distance: + return max_distance + 1 + previous = current + return previous[-1] + + +def _has_wakeword(text: str, variants: Iterable[str]) -> bool: + normalized = _normalize_wakeword_text(text) + if not normalized: + return False + compact = normalized.replace(" ", "") + tokens = normalized.split() + for raw_candidate in variants: + candidate = _normalize_wakeword_text(raw_candidate).replace(" ", "") + if len(candidate) < 5: + continue + forms = {candidate} + if candidate.endswith("y"): + forms.update({f"{candidate[:-1]}i", f"{candidate[:-1]}ie"}) + if candidate.endswith("i"): + forms.add(f"{candidate}e") + if any(form in compact for form in forms): + return True + if not (candidate.startswith("trini") or candidate.startswith("drini")): + continue + for token in tokens: + if len(token) < 5: + continue + for form in forms: + distance = 1 if len(form) < 8 else 2 + if _bounded_levenshtein(token, form, distance) <= distance: + return True + return False + + def _message_text(content: Any) -> str: if isinstance(content, str): return content.strip() @@ -35,6 +108,7 @@ class TrinityConversationBackend(ConversationBackend): def __init__(self, home: str | Path): self.home = Path(home).expanduser().resolve() self.core_dir = self.home / "core" + self.config_path = self.core_dir / "config.json" self.transcript_path = self.home / "TrinityRuntime" / "voice" / "voice_session.md" self.transcript_path.parent.mkdir(parents=True, exist_ok=True) if not self.transcript_path.exists(): @@ -42,6 +116,20 @@ def __init__(self, home: str | Path): self._brain = None self._brain_lock = threading.RLock() + def _runtime_voice_policy(self) -> tuple[str, tuple[str, ...]]: + try: + config = json.loads(self.config_path.read_text(encoding="utf-8")) + except (OSError, ValueError, TypeError): + config = {} + mode = str(config.get("system", {}).get("mode", "office") or "office").strip().lower() + if mode == "chat": + mode = "office" + if mode not in {"lecture", "office"}: + mode = "office" + configured = config.get("persona", {}).get("trigger_variants") or DEFAULT_WAKEWORD_VARIANTS + variants = tuple(str(item) for item in configured if str(item).strip()) + return mode, variants or DEFAULT_WAKEWORD_VARIANTS + def _ensure_brain(self): if self._brain is None: import sys @@ -98,12 +186,16 @@ def _append_chat_events(self, user_text: str, answer: str, request_id: str) -> N ) def respond(self, text: str, *, session_id: str = "", turn_id: str = "") -> Iterable[str]: - rejection = enforce_input_language(text) - if rejection: - return [rejection] query = str(text or "").strip() if not query: return [] + mode, wakeword_variants = self._runtime_voice_policy() + if mode == "lecture" and not _has_wakeword(query, wakeword_variants): + self._append_transcript("Lecture (ohne Wakeword)", query) + return [] + rejection = enforce_input_language(query) + if rejection: + return [rejection] request_id = turn_id or uuid.uuid4().hex with self._brain_lock: self._append_transcript("User", query) diff --git a/core/voice/doctor.py b/core/voice/doctor.py index 6d88f72..a8f8ac6 100644 --- a/core/voice/doctor.py +++ b/core/voice/doctor.py @@ -53,27 +53,34 @@ def _llm_health(base_url: str, api_key: str = "") -> tuple[bool, str]: def run_checks(config: VoiceConfig) -> list[Check]: profile = config.profile + inference_required = profile.runtime_role != "client" s2s_version = _package_version("speech-to-speech") mlx_audio_version = _package_version("mlx-audio") checks = [ Check("Python", sys.version_info >= (3, 10), platform.python_version()), Check("Engine", config.engine in {"legacy", "eve"}, config.engine), Check("Profil", not config.validate(), profile.name if not config.validate() else "; ".join(config.validate())), - Check("speech-to-speech", s2s_version == "0.2.11", s2s_version or "nicht installiert"), - Check("mlx-audio", bool(mlx_audio_version), mlx_audio_version or "nicht installiert", required=platform.system() == "Darwin"), - Check("Parakeet-Modul", importlib.util.find_spec("speech_to_speech") is not None, config.stt_model), - Check("Eve-Referenzaudio", config.reference_audio.is_file(), str(config.reference_audio)), - Check("Backend-Port", _port_available(config.backend_host, config.backend_port), f"{config.backend_host}:{config.backend_port}"), + Check("speech-to-speech", s2s_version == "0.2.11", s2s_version or "nur auf dem GPU-Host benötigt", required=inference_required), + Check("mlx-audio", bool(mlx_audio_version), mlx_audio_version or "nicht installiert", required=inference_required and platform.system() == "Darwin"), + Check("Parakeet-Modul", importlib.util.find_spec("speech_to_speech") is not None, config.stt_model, required=inference_required), + Check("Eve-Referenzaudio", config.reference_audio.is_file(), str(config.reference_audio), required=inference_required), ] - if profile.mode == "realtime": + if profile.conversation_backend == "trinity": + checks.append(Check("Backend-Port", _port_available(config.backend_host, config.backend_port), f"{config.backend_host}:{config.backend_port}")) + if profile.mode == "realtime" and profile.runtime_role != "client": checks.extend([ Check("Interner Voice-Port", _port_available("127.0.0.1", profile.internal_port), str(profile.internal_port)), Check("Öffentlicher Voice-Port", _port_available(profile.bind_host, profile.public_port), f"{profile.bind_host}:{profile.public_port}"), Check("Realtime-Token", bool(config.access_token) or profile.bind_host in {"127.0.0.1", "localhost", "::1"}, "gesetzt" if config.access_token else "nur Loopback"), ]) + if profile.runtime_role == "client": + checks.append(Check("Ubuntu Voice URL", bool(config.remote_voice_url), config.remote_voice_url or "nicht gesetzt")) if profile.conversation_backend == "direct": ok, detail = _llm_health(config.direct_llm_base_url, config.direct_llm_api_key) checks.append(Check("Direktes Diagnose-LLM", ok, detail)) + if profile.conversation_backend == "remote": + ok, detail = _llm_health(config.remote_core_base_url, config.remote_core_api_key) + checks.append(Check("Windows Trinity Core", ok, detail)) checks.append(Check("Tailscale", shutil.which("tailscale") is not None, shutil.which("tailscale") or "optional", required=False)) return checks diff --git a/core/voice/local_realtime_client.py b/core/voice/local_realtime_client.py index f380524..9e03678 100644 --- a/core/voice/local_realtime_client.py +++ b/core/voice/local_realtime_client.py @@ -6,11 +6,13 @@ import json import logging import math +import os import threading import time from collections import deque from queue import Empty, Full, Queue from typing import Any +from urllib.parse import parse_qsl, urlencode, urlsplit, urlunsplit import numpy as np @@ -33,10 +35,18 @@ class LocalRealtimeAudioClient: loudspeaker echo; headphones remain the recommended route for barge-in. """ - def __init__(self, config: VoiceConfig, host: str = "127.0.0.1"): + def __init__( + self, + config: VoiceConfig, + host: str = "127.0.0.1", + endpoint: str = "", + access_token: str = "", + ): self.config = config self.host = host self.port = config.profile.internal_port + self.endpoint = endpoint.strip() + self.access_token = access_token.strip() self._stop = threading.Event() self._ready = threading.Event() self._thread: threading.Thread | None = None @@ -49,6 +59,12 @@ def __init__(self, config: VoiceConfig, host: str = "127.0.0.1"): self._barge_in_candidate: deque[bytes] = deque(maxlen=BARGE_IN_CONFIRM_BLOCKS) self._last_output_at = 0.0 self._last_cancel_at = 0.0 + self._speech_queue_path = config.home / "TrinityRuntime" / "voice" / "desktop_speech_queue.jsonl" + self._ready_path = config.home / "TrinityRuntime" / "voice" / "desktop_eve_audio.ready" + self._trinity_config_path = config.home / "core" / "config.json" + self._speaker_check_at = 0.0 + self._desktop_output_enabled = True + self._speech_queue_offset = 0 def start(self, timeout: float = 20.0) -> None: self._thread = threading.Thread( @@ -64,6 +80,7 @@ def start(self, timeout: float = 20.0) -> None: def stop(self) -> None: self._stop.set() + self._remove_ready_marker() connection = self._connection if connection is not None: try: @@ -127,9 +144,12 @@ def _run(self) -> None: import sounddevice as sd from websockets.sync.client import connect - uri = f"ws://{self.host}:{self.port}/v1/realtime" + uri = self._connection_uri() with connect(uri, open_timeout=12, max_size=None, proxy=None) as connection: self._connection = connection + self._speech_queue_path.parent.mkdir(parents=True, exist_ok=True) + self._speech_queue_path.touch(exist_ok=True) + self._speech_queue_offset = self._speech_queue_path.stat().st_size connection.send(json.dumps(self._session_update(), ensure_ascii=False)) sender = threading.Thread( target=self._send_loop, @@ -156,8 +176,10 @@ def _run(self) -> None: callback=self._input_callback, ): self._ready.set() + self._write_ready_marker() print("Eve Desktop-Audio bereit: Unterbrechen durch Sprechen ist aktiv.") while not self._stop.is_set(): + self._consume_speech_queue() try: raw = connection.recv(timeout=0.1) except TimeoutError: @@ -174,9 +196,67 @@ def _run(self) -> None: LOGGER.exception("Lokaler Eve-Audioclient beendet") finally: self._stop.set() + self._remove_ready_marker() if sender: sender.join(timeout=2) + def _connection_uri(self) -> str: + raw = self.endpoint or f"ws://{self.host}:{self.port}/v1/realtime" + parts = urlsplit(raw) + path = parts.path or "/v1/realtime" + query = dict(parse_qsl(parts.query, keep_blank_values=True)) + if self.access_token: + query["access_token"] = self.access_token + return urlunsplit((parts.scheme, parts.netloc, path, urlencode(query), parts.fragment)) + + def _write_ready_marker(self) -> None: + try: + self._ready_path.parent.mkdir(parents=True, exist_ok=True) + self._ready_path.write_text(str(os.getpid()), encoding="utf-8") + except OSError: + LOGGER.debug("Eve-Bereitschaftsmarker konnte nicht geschrieben werden", exc_info=True) + + def _remove_ready_marker(self) -> None: + try: + self._ready_path.unlink(missing_ok=True) + except OSError: + LOGGER.debug("Eve-Bereitschaftsmarker konnte nicht entfernt werden", exc_info=True) + + def _consume_speech_queue(self) -> None: + try: + with self._speech_queue_path.open("r", encoding="utf-8") as handle: + handle.seek(self._speech_queue_offset) + lines = handle.readlines() + self._speech_queue_offset = handle.tell() + except OSError: + return + for line in lines: + try: + payload = json.loads(line) + except (TypeError, ValueError): + continue + text = str(payload.get("text") or "").strip() + if not text: + continue + self._queue_event({ + "type": "response.create", + "response": { + "conversation": "none", + "input": [{ + "type": "message", + "role": "user", + "content": [{"type": "input_text", "text": text}], + }], + "instructions": ( + "Lies den bereitgestellten deutschen Text wortgetreu vor. " + "Gib ausschliesslich diesen Text aus." + ), + "output_modalities": ["audio"], + "max_output_tokens": 4096, + "metadata": {"trinity_action": "desktop_read_aloud"}, + }, + }) + def _send_loop(self, connection) -> None: while not self._stop.is_set(): try: @@ -193,6 +273,10 @@ def _output_callback(self, outdata, frames, _time_info, status) -> None: if status: LOGGER.debug("Desktop-Audioausgabe: %s", status) wanted = frames * SAMPLE_BYTES + if not self._desktop_speaker_selected(): + self._clear_output() + outdata[:] = b"\x00" * wanted + return with self._output_lock: take = min(wanted, len(self._output)) outgoing = bytes(self._output[:take]) @@ -206,6 +290,22 @@ def _output_callback(self, outdata, frames, _time_info, status) -> None: self._played_output.append(output_samples) self._last_output_at = time.monotonic() + def _desktop_speaker_selected(self) -> bool: + now = time.monotonic() + if now < self._speaker_check_at: + return self._desktop_output_enabled + self._speaker_check_at = now + 0.35 + try: + config = json.loads(self._trinity_config_path.read_text(encoding="utf-8")) + speaker = config.get("system", {}).get("speech_output", {}) + self._desktop_output_enabled = not speaker or str( + speaker.get("kind") or "desktop" + ).strip().lower() == "desktop" + except (OSError, ValueError, TypeError): + # A transient write/read race must not unexpectedly mute the active desktop. + pass + return self._desktop_output_enabled + def _input_callback(self, indata, _frames, _time_info, status) -> None: if status: LOGGER.debug("Desktop-Audioeingabe: %s", status) @@ -309,6 +409,9 @@ def _handle_event(self, raw: str | bytes) -> None: if event_type == "input_audio_buffer.speech_started": self._clear_output() elif event_type == "response.output_audio.delta": + if not self._desktop_speaker_selected(): + self._clear_output() + return encoded = str(event.get("delta") or "") try: audio = base64.b64decode(encoded, validate=True) diff --git a/core/voice/runtime.py b/core/voice/runtime.py index c4774a4..c739f37 100644 --- a/core/voice/runtime.py +++ b/core/voice/runtime.py @@ -58,13 +58,16 @@ def start(self) -> None: profile = self.config.profile if profile.conversation_backend == "trinity": backend = TrinityConversationBackend(self.config.home) - else: + elif profile.conversation_backend == "direct": backend = DirectLLMConversationBackend( self.config.direct_llm_base_url, self.config.direct_llm_model, self.config.direct_llm_api_key, ) + else: + backend = None if profile.conversation_backend == "trinity": + assert backend is not None self.backend_server = TrinityConversationHTTPServer( backend, self.config.backend_host, @@ -73,6 +76,15 @@ def start(self) -> None: ) self.backend_server.start() + if profile.runtime_role == "client": + self.local_audio_client = LocalRealtimeAudioClient( + self.config, + endpoint=self.config.remote_voice_url, + access_token=self.config.remote_voice_token or self.config.access_token, + ) + self.local_audio_client.start() + return + command = build_speech_to_speech_command(self.config) env = os.environ.copy() env["TOKENIZERS_PARALLELISM"] = "false" @@ -91,6 +103,10 @@ def start(self) -> None: self.local_audio_client.start() def wait(self) -> int: + if not self.process and self.local_audio_client: + while self.local_audio_client.is_alive and self.local_audio_client.failure is None: + time.sleep(0.25) + return 1 if self.local_audio_client.failure else 0 if not self.process: return 0 while self.process is not None and self.process.poll() is None: diff --git a/core/web_ui.py b/core/web_ui.py index bfa50c4..f37029a 100644 --- a/core/web_ui.py +++ b/core/web_ui.py @@ -121,7 +121,7 @@ def render_web_ui(auth_enabled=False): {title:'Persona',fields:[['persona.agent_name','Name von Trinity','text'],['persona.trigger_variants','Wakeword-Varianten (kommagetrennt)','text','list']]}, {title:'LLM',fields:[['llm.active_slot','Aktiver Slot','select',['local','remote_1','remote_2']],['llm.local.url','Lokale URL','text'],['llm.local.model','Lokales Modell','text'],['llm.local.api_key','Lokaler API-Schlüssel','password'],['llm.remote_1.url','Remote 1 URL','text'],['llm.remote_1.model','Remote 1 Modell','text'],['llm.remote_1.api_key','Remote 1 API-Schlüssel','password'],['llm.remote_2.url','Remote 2 URL','text'],['llm.remote_2.model','Remote 2 Modell','text'],['llm.remote_2.api_key','Remote 2 API-Schlüssel','password']]}, {title:'Sprache und Audio',fields:[['stt.model','STT-Modell','select',['tiny','base','small','medium','large']],['stt.silence_threshold','Stille-Schwelle','number'],['stt.chunk_duration','Chunk-Dauer (Sekunden)','number'],['stt.show_volume_meter','Pegelanzeige','checkbox'],['tts.voice','Desktop-TTS-Stimme','text'],['audio_routing.private_device','Privates Audiogeraet','text'],['audio_routing.public_device','Oeffentliches Audiogeraet','text']]}, - {title:'Proaktiv und Oberflaechen',fields:[['proactive.heartbeat_enabled','Heartbeat aktiv','checkbox'],['proactive.bubbles_enabled','Bubbles aktiv','checkbox'],['proactive.visuals_enabled','Visuelle Hinweise aktiv','checkbox'],['proactive.interval_minutes','Heartbeat-Intervall (Minuten)','number'],['system.mode','Modus','select',['office','lecture','chat']],['system.eyes_ui_enabled','Augen-UI','checkbox'],['system.classic_ui_enabled','ClassicUI','checkbox'],['system.web_ui_enabled','WebUI','checkbox'],['system.terminal_cli_enabled','Terminal-CLI','checkbox'],['system.show_terminal','Terminalfenster anzeigen','checkbox'],['system.windows_speech_enabled','Windows-Spracheingabe','checkbox']]}, + {title:'Proaktiv und Oberflaechen',fields:[['proactive.heartbeat_enabled','Heartbeat aktiv','checkbox'],['proactive.bubbles_enabled','Bubbles aktiv','checkbox'],['proactive.visuals_enabled','Visuelle Hinweise aktiv','checkbox'],['proactive.interval_minutes','Heartbeat-Intervall (Minuten)','number'],['system.mode','Modus','select',['office','lecture']],['system.eyes_ui_enabled','Augen-UI','checkbox'],['system.classic_ui_enabled','ClassicUI','checkbox'],['system.web_ui_enabled','WebUI','checkbox'],['system.terminal_cli_enabled','Terminal-CLI','checkbox'],['system.show_terminal','Terminalfenster anzeigen','checkbox'],['system.windows_speech_enabled','Windows-Spracheingabe','checkbox']]}, {title:'APIs und Bild',fields:[['apis.tavily','Tavily API-Schlüssel','password'],['apis.fal_ai','fal.ai API-Schlüssel','password'],['image.primary_model','Primäres Bildmodell','text'],['image.fallback_model','Fallback-Bildmodell','text'],['comfyui.enabled','ComfyUI aktiv','checkbox'],['comfyui.server_url','ComfyUI-Server','text'],['comfyui.default_workflow','ComfyUI-Workflow','text']]}, {title:'Agenten-Frameworks',fields:[['codex.enabled','Codex-Auftraege erlauben','checkbox'],['codex.executable','Codex-Programm','text'],['codex.projects','Codex-Projekte: Alias = Pfad','textarea','projects'],['codex.default_project','Codex-Standardprojekt','text'],['codex.sandbox','Codex-Sandbox','select',['read-only','workspace-write']],['codex.timeout_seconds','Codex-Zeitlimit (Sekunden)','number'],['codex.max_output_chars','Codex-Antwortlaenge','number'],['codex.ephemeral','Keine dauerhafte Codex-Sitzung','checkbox'],['codex.network_access','Netzwerkzugriff fuer Codex','checkbox'],['pi.enabled','Pi-Auftraege erlauben','checkbox'],['pi.executable','Pi-Programm oder Wrapper','text'],['pi.projects','Pi-Projekte: Alias = Pfad','textarea','projects'],['pi.default_project','Pi-Standardprojekt','text'],['pi.arguments','Pi-Argumente (optional, {prompt} als Platzhalter)','textarea','list'],['pi.timeout_seconds','Pi-Zeitlimit (Sekunden)','number'],['pi.max_output_chars','Pi-Antwortlaenge','number'],['opencode.enabled','OpenCode-Auftraege erlauben','checkbox'],['opencode.executable','OpenCode-Programm','text'],['opencode.server_url','Laufender OpenCode-Dienst','text'],['opencode.projects','OpenCode-Projekte: Alias = Pfad','textarea','projects'],['opencode.default_project','OpenCode-Standardprojekt','text'],['opencode.agent','OpenCode-Agent','text'],['opencode.model','OpenCode-Modell (optional)','text'],['opencode.timeout_seconds','OpenCode-Zeitlimit (Sekunden)','number'],['opencode.max_output_chars','OpenCode-Antwortlaenge','number'],['harness_routing.frameworks','Rollen je Framework inkl. Trinity (JSON)','textarea','json'],['harness_routing.agent_assignments','Agenten-Ausfuehrung je Framework (JSON)','textarea','json'],['agent_catalog.agents','Agentenkatalog: Reifegrad, Rechte, Freigaben, Limits (JSON)','textarea','json']]}, {title:'Trinity-Ablagen',fields:[['control_plane.enabled','Control Plane aktiv','checkbox'],['control_plane.runtime_root','Lokale Laufzeitdaten','text'],['control_plane.vault_root','Cloud-Vault für dauerhafte Inhalte','text'],['control_plane.external_agents_root','Lokale Wurzel mit .agents','text'],['control_plane.default_brainvault_harness','Standard-Harness für externe Agenten','select',['codex','pi','opencode','trinity']]]}, diff --git a/docs/EVEN_G2.md b/docs/EVEN_G2.md index 4ee131f..bf2b7e4 100644 --- a/docs/EVEN_G2.md +++ b/docs/EVEN_G2.md @@ -17,10 +17,12 @@ Die Brille verbindet sich nicht direkt mit Mac oder Windows. Das Telefon mit der ## G2-Oberflaeche -- links oben: Bubble-Hinweis und optional bis zu fuenf Vorbereitungspunkte -- rechts oben: Anzahl neuer Medien sowie aktives Profil und Modus +- links oben: Bubble-Hinweis und optional bis zu sechs Vorbereitungspunkte +- rechts oben: Uhrzeit, Modus, Ausgabegeraet, Wetter der naechsten Stunde, + optionaler Kurztermin und Trinity-Status - links unten: letzte Transkription oder Trinity-Antwort -- rechts unten: Uhrzeit +- rechts unten: Anzahl laufender oder fertiger Medien +- ganz unten: eine neue n-tv-Schlagzeile fuer etwa acht Sekunden - Mitte: bewusst frei Das G2-Display ist monochrom. Die Trinity-Ampel wird daher mit `!`, `!!` und `!!!` dargestellt. Die Brille hat keinen Lautsprecher; Antworten erscheinen als Text. @@ -31,6 +33,20 @@ geoeffnet werden. Ein weiterer Tipp wechselt zum naechsten Ergebnis, Doppeltippen schliesst die Vorschau. Fuer Audio, Video und HTML zeigt das HUD eine kompakte Ergebniskarte. +### Wetter, Termin und Schlagzeile + +Trinity Desktop liest die offizielle n-tv-RSS-Schlagzeile und die +stundengenaue Open-Meteo-Prognose serverseitig ein und stellt beides +zwischengespeichert unter `GET /ambient` bereit. Der in der G2-Einstellung +`Wetter-Ersatzort` gesetzte Ort gilt, solange kein Companion-Client einen +aktuellen Standort meldet; Standard ist `Filderstadt`. + +Unter iPhone/iPad **Settings -> G2 HUD-Kontext** koennen Standort und naechster +Kalendertermin getrennt freigegeben werden. Die Meldung ist optional, +widerrufbar und bleibt nur im Arbeitsspeicher der privaten Trinity-Bridge. Beim +Abschalten werden die zuvor gemeldeten Werte dort entfernt. Ohne die +Freigaben funktionieren Schlagzeile und Ersatzort-Wetter weiterhin. + ## Modi - **Zuruf:** Gesprochenes wird transkribiert und durch Trinitys normalen Wakeword-Pfad verarbeitet. diff --git a/docs/VOICE_ARCHITECTURE.md b/docs/VOICE_ARCHITECTURE.md index 83f7ff5..1de3742 100644 --- a/docs/VOICE_ARCHITECTURE.md +++ b/docs/VOICE_ARCHITECTURE.md @@ -27,6 +27,8 @@ proxy adds access-token enforcement before a remote client can reach it. | `eve-mac-server` | iPhone/iPad thin clients over Tailscale | MLX/MPS | Remote Apple device | | `eve-windows-local` | Windows microphone and speaker with barge-in | CUDA | Local PC | | `eve-windows-server` | iPhone/iPad thin clients over Tailscale | CUDA | Remote Apple device | +| `eve-windows-remote` | Windows control plane using an Ubuntu GPU host | Remote CUDA | Windows or Companion | +| `eve-linux-gpu-server` | Parakeet and Eve inference for a remote Windows core | CUDA | Remote desktop or Companion | | `eve-trinity` | Auto-selected local Trinity path | Auto | Local desktop | | `eve-direct-ornith` | Isolated half-duplex LLM diagnostics | MLX/MPS | Local desktop | @@ -54,3 +56,5 @@ sync through the normal Trinity Bridge independently of that audio choice. Configuration details are in `core/config.json.example`; security boundaries are documented in [VOICE_SECURITY.md](VOICE_SECURITY.md). +The split Ubuntu/Windows deployment is documented in +[VOICE_UBUNTU_HOST.md](VOICE_UBUNTU_HOST.md). diff --git a/docs/VOICE_SECURITY.md b/docs/VOICE_SECURITY.md index 25d0b80..04a2e75 100644 --- a/docs/VOICE_SECURITY.md +++ b/docs/VOICE_SECURITY.md @@ -2,7 +2,11 @@ - Voice binds to loopback by default. Remote bind is an explicit setting. - Remote Realtime access requires a separate Voice token; use Tailscale or WSS. -- Do not expose ports 8765/8766 through a public router. +- Do not expose ports 8765/8766/18767 through a public router. +- In the Ubuntu/Windows split layout, port `18767` is the authenticated Trinity + Core endpoint. Bind it only to a private LAN/Tailscale interface and use a + separate Core token. The Ubuntu Voice token and Windows Core token must not + be reused as public API credentials. - Tokens belong in local configuration/Keychain, never source control or logs. - Standard logs contain session/turn IDs and timings, not reference audio, complete emails or full private prompts. diff --git a/docs/VOICE_UBUNTU_HOST.md b/docs/VOICE_UBUNTU_HOST.md new file mode 100644 index 0000000..40afcb2 --- /dev/null +++ b/docs/VOICE_UBUNTU_HOST.md @@ -0,0 +1,80 @@ +# Ubuntu NVIDIA host for Trinity Eve Voice + +This layout keeps the GPU on Ubuntu and runs Trinity's control plane inside a +Windows VM. Ubuntu provides compute services only; Windows remains responsible +for sessions, memory, agents, tools, approvals and the user interface. + +```mermaid +flowchart LR + C["Windows, iPhone, iPad or G2 audio client"] <-->|"PCM and realtime events :8766"| U["Ubuntu Eve Voice"] + U -->|"transcribed text :18767"| W["Windows Trinity Core"] + W -->|"OpenAI-compatible API"| L["Ubuntu LLM :1234"] + W -->|"answer text"| U + U -->|"Eve audio"| C +``` + +## 1. Prepare Ubuntu + +Install the NVIDIA driver and verify CUDA visibility: + +```bash +nvidia-smi +``` + +Clone Trinity into a local, non-synchronized directory, create its normal +Python environment, then install the optional voice runtime with an authorized +Eve reference sample: + +```bash +cd ~/Trinity_Assistant +./scripts/install_voice_ubuntu.sh /secure/path/Eve_Schule.mp3 +``` + +The sample is biometric/personal data. Keep it outside Git and cloud-synced +project folders. + +## 2. Configure the Ubuntu voice server + +In Ubuntu's local `core/config.json`, use: + +```json +{ + "voice": { + "engine": "eve", + "profile": "eve-linux-gpu-server", + "access_token": "VOICE_TOKEN", + "remote_core_base_url": "http://WINDOWS_TAILSCALE_IP:18767/v1", + "remote_core_api_key": "CORE_TOKEN" + } +} +``` + +Validate and start it: + +```bash +./venv/bin/trinity voice doctor --profile eve-linux-gpu-server +./venv/bin/trinity voice serve --profile eve-linux-gpu-server +``` + +The Voice Gateway listens on port `8766`. Allow access only from the private +LAN or Tailscale interface. + +## 3. Configure Windows + +Follow [VOICE_WINDOWS.md](VOICE_WINDOWS.md) and use profile +`eve-windows-remote`. Its Core token must match `CORE_TOKEN`; its Voice token +must match `VOICE_TOKEN`. + +## 4. Verify the complete path + +1. `trinity voice doctor --profile eve-linux-gpu-server` succeeds on Ubuntu. +2. `trinity voice doctor --profile eve-windows-remote` succeeds on Windows. +3. A typed Windows chat request returns normally through Trinity Core. +4. A Windows/iPhone/iPad microphone request reaches Ubuntu STT, appears in the + same Trinity session, and returns as Eve audio on the selected speaker. +5. Disabling Ubuntu leaves the Windows UI usable; selecting Legacy restores the + previous Windows STT/TTS path. + +GPU passthrough is intentionally not used. It would usually remove the GPU from +the Ubuntu host and adds VM/driver fragility without improving this networked +speech pipeline. diff --git a/docs/VOICE_WINDOWS.md b/docs/VOICE_WINDOWS.md index 931c954..803e68e 100644 --- a/docs/VOICE_WINDOWS.md +++ b/docs/VOICE_WINDOWS.md @@ -1,10 +1,49 @@ # Eve Voice on Windows 11 -Windows supports both a local desktop conversation and a headless Voice Gateway -for iPhone/iPad. Both require a compatible NVIDIA GPU visible inside Windows and -a CUDA/PyTorch Qwen3-TTS Base checkpoint. A VM therefore needs working GPU -passthrough; CPU-only synthesis is not a productive Eve configuration. The -checkpoint is configurable because the Apple MLX identifier is not portable. +Windows supports two production layouts: + +1. **Recommended for a VM:** Windows runs Trinity's UI, sessions, memory, + agents and policy layer. A private Ubuntu host with an NVIDIA GPU runs + Parakeet STT and Qwen3-TTS/Eve. No PCI passthrough is required. +2. **Native Windows GPU:** Windows runs Trinity and the CUDA voice models on a + GPU that is directly visible inside Windows. + +## Windows VM with an Ubuntu GPU host + +Run in an elevated PowerShell: + +```powershell +cd $env:LOCALAPPDATA\Trinity +.\scripts\install_voice_windows.ps1 -RemoteGPUClient -OpenFirewall +``` + +In **Settings -> Voice**, select +`Windows VM with Eve on an Ubuntu GPU host` and configure: + +- Ubuntu Voice URL: `ws://UBUNTU_TAILSCALE_IP:8766/v1/realtime` +- Voice token: a long, random token shared only with Ubuntu +- Windows Core bind: `0.0.0.0` +- Windows Core port: `18767` +- Windows Core token: a second long, random token shared only with Ubuntu + +The normal Trinity LLM provider can point to an OpenAI-compatible endpoint on +Ubuntu, for example `http://UBUNTU_TAILSCALE_IP:1234/v1`. Restart Trinity and +run: + +```powershell +trinity voice doctor --profile eve-windows-remote +``` + +Ubuntu setup is documented in +[VOICE_UBUNTU_HOST.md](VOICE_UBUNTU_HOST.md). Restrict ports `8766`, `18767` +and the LLM port to the private LAN/Tailnet. Do not expose them on a public +router. + +## Native Windows GPU + +Native Windows voice requires a compatible NVIDIA GPU visible inside Windows +and a CUDA/PyTorch Qwen3-TTS Base checkpoint. The checkpoint is configurable +because the Apple MLX identifier is not portable. ```powershell cd $env:LOCALAPPDATA\Trinity diff --git a/docs/release_notes/v0.17.4.md b/docs/release_notes/v0.17.4.md new file mode 100644 index 0000000..aab11da --- /dev/null +++ b/docs/release_notes/v0.17.4.md @@ -0,0 +1,24 @@ +# Trinity v0.17.4 + +## Even G2 microphone routing + +- Even G2 can act as the lecture microphone without becoming an audio output. +- Each G2 profile routes spoken answers to the iPhone/iPad Companion, Trinity + Desktop, or HUD/chat only. +- Companion output is spoken only on a mobile device with `Listen` enabled. +- Desktop output is handed to the active Eve realtime voice and falls back to + the legacy system voice only when Eve is unavailable. +- G2 source and output-target metadata survive the complete bridge, STT and + response path. + +## Runtime modes + +- `Lecture` requires the configured Trinity wake word. +- `Office` responds directly and replaces the redundant former `Chat` mode. +- Existing `chat` settings migrate to `office` without manual changes. + +## Verification + +- Desktop voice, wake-word, bridge and runtime regression tests pass. +- Even G2 unit tests verify output-target routing. +- Companion simulator and physical-device builds pass. diff --git a/docs/release_notes/v0.17.5.md b/docs/release_notes/v0.17.5.md new file mode 100644 index 0000000..91b20b9 --- /dev/null +++ b/docs/release_notes/v0.17.5.md @@ -0,0 +1,23 @@ +# Trinity v0.17.5 + +## Reliable G2 wake mode + +- Enforces the wake word on Even G2 before any transcript, HUD command, or + checklist command is routed to Trinity. +- Keeps Office mode as an explicit continuous-conversation mode. +- Uses precise G2 speech recognition by default and removes checklist phrases + from the Whisper prompt bias. +- Prevents ordinary lecture speech and assistant responses from creating + unsolicited checklists. + +## Companion audio feedback + +- Retries a temporarily unavailable Eve playback session once. +- Shows whether mobile playback was skipped because `Listen` is disabled or + because Eve could not be reached. + +## Verification + +- Desktop bridge-audio regression tests pass. +- Even G2 command, checklist, routing, and HUD tests pass. +- The Trinity Companion simulator build passes. diff --git a/docs/release_notes/v0.17.6.md b/docs/release_notes/v0.17.6.md new file mode 100644 index 0000000..ef7b48e --- /dev/null +++ b/docs/release_notes/v0.17.6.md @@ -0,0 +1,18 @@ +# Trinity v0.17.6 + +## Synchronized single-speaker output + +- Adds a Bridge-wide speaker selection with a stable identity for Trinity + Desktop, each iPhone, and each iPad. +- Adds `Ich spreche hier` to ClassicUI and persists the selected output device. +- Stops Desktop TTS when a Companion client owns the output, preventing two + devices from speaking the same response. +- Exposes the current target through `/speaker` and the regular instance/event + state so G2 and Companion clients remain synchronized. + +## Compatibility + +- Existing installations default to their Trinity Desktop until another client + explicitly takes over. +- Voice profiles, Bridge tokens, Eve/Legacy selection, and G2 microphone input + remain unchanged. diff --git a/docs/release_notes/v0.17.7.md b/docs/release_notes/v0.17.7.md new file mode 100644 index 0000000..f66cbaa --- /dev/null +++ b/docs/release_notes/v0.17.7.md @@ -0,0 +1,18 @@ +# Trinity v0.17.7 + +## G2 glance context + +- Adds an authenticated ambient-context endpoint for the Even G2 HUD. +- Fetches the latest n-tv RSS headline and the next-hour Open-Meteo forecast + with short server-side caches. +- Accepts optional current-location and next-calendar-event context from an + authorized iPhone or iPad Companion. +- Keeps device location and calendar data in memory only and clears them when + sharing is disabled. +- Falls back to the place configured in the G2 profile, defaulting to + Filderstadt. + +## Verification + +- Ambient service and authenticated HTTP endpoint tests pass together with + the existing Bridge and G2-audio suites. diff --git a/docs/release_notes/v0.17.8.md b/docs/release_notes/v0.17.8.md new file mode 100644 index 0000000..1fa121f --- /dev/null +++ b/docs/release_notes/v0.17.8.md @@ -0,0 +1,16 @@ +# Trinity v0.17.8 + +## One Speaker, One Output + +- ClassicUI, iPhone and iPad expose one speaker control in their main toolbar. +- Claiming output on one client makes that device the exclusive Trinity speaker. +- Every other connected client follows the shared Bridge state and shows a muted speaker. +- Pressing the active speaker again mutes Trinity on every device. +- Direct Eve playback on desktop and Companion now follows the same selection instead of bypassing it. +- Separate toolbar controls for TTS and the local iOS audio route were removed; detailed audio routing remains an operating-system or settings concern. + +## Verification + +- Full desktop Python test suite +- Focused Bridge and Eve playback tests +- iPhone/iPad simulator build without code signing diff --git a/docs/release_notes/v0.17.9.md b/docs/release_notes/v0.17.9.md new file mode 100644 index 0000000..8f3a263 --- /dev/null +++ b/docs/release_notes/v0.17.9.md @@ -0,0 +1,8 @@ +# Trinity v0.17.9 - Remote GPU Voice Runtime + +- Adds an Ubuntu NVIDIA voice-server profile for Parakeet STT and Qwen3-TTS/Eve. +- Adds a Windows VM client profile that keeps Trinity Core, sessions, memory, + agents and UI on Windows while using Ubuntu for GPU inference. +- Adds authenticated remote Core and Voice links with diagnostics and private + firewall guidance. +- Keeps native macOS, native Windows CUDA and Legacy STT/TTS paths intact. diff --git a/pyproject.toml b/pyproject.toml index 6834a09..1f1a32b 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "trinity-assistant" -version = "0.17.3" +version = "0.17.8" description = "Local-first academic personal concierge for macOS, Windows 11, and Linux servers" readme = "README.md" requires-python = ">=3.10,<3.15" diff --git a/scripts/install_voice_ubuntu.sh b/scripts/install_voice_ubuntu.sh new file mode 100755 index 0000000..a3d4aae --- /dev/null +++ b/scripts/install_voice_ubuntu.sh @@ -0,0 +1,51 @@ +#!/usr/bin/env bash +set -euo pipefail + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +PYTHON="$ROOT/venv/bin/python" +VOICE_DIR="$ROOT/TrinityRuntime/voices/eve" +VOICE_SOURCE="${1:-}" + +if [[ "$(uname -s)" != "Linux" ]]; then + echo "This installer is intended for the Ubuntu GPU host." >&2 + exit 1 +fi +if [[ ! -x "$PYTHON" ]]; then + echo "Trinity Python environment not found: $PYTHON" >&2 + exit 1 +fi +if ! command -v nvidia-smi >/dev/null 2>&1; then + echo "NVIDIA driver not found. Install and verify it on Ubuntu first." >&2 + exit 1 +fi + +nvidia-smi >/dev/null +"$PYTHON" -m pip install "speech-to-speech==0.2.11" "websockets>=14,<18" + +mkdir -p "$VOICE_DIR" +cp "$ROOT/assets/voices/eve/ref_text.txt" "$VOICE_DIR/ref_text.txt" +if [[ -n "$VOICE_SOURCE" ]]; then + if [[ ! -f "$VOICE_SOURCE" ]]; then + echo "Voice sample not found: $VOICE_SOURCE" >&2 + exit 1 + fi + cp "$VOICE_SOURCE" "$VOICE_DIR/Eve_Schule.mp3" +else + echo "No authorized Eve sample copied. Pass its path as the first argument." >&2 +fi + +cat <<'EOF' +Ubuntu Eve dependencies are installed. + +Next configure Trinity's voice section with: + profile: eve-linux-gpu-server + access_token: a long Voice token + remote_core_base_url: http://WINDOWS-TAILSCALE-IP:18767/v1 + remote_core_api_key: the separate Windows Core token + +Then run: + venv/bin/trinity voice doctor --profile eve-linux-gpu-server + venv/bin/trinity voice serve --profile eve-linux-gpu-server + +Do not expose ports 8766 or 18767 through a public router. +EOF diff --git a/scripts/install_voice_windows.ps1 b/scripts/install_voice_windows.ps1 index 9b6ffd1..76bea22 100644 --- a/scripts/install_voice_windows.ps1 +++ b/scripts/install_voice_windows.ps1 @@ -2,6 +2,7 @@ param( [string]$TrinityRoot = "$env:LOCALAPPDATA\Trinity", [string]$VoiceSource = "", + [switch]$RemoteGPUClient, [switch]$OpenFirewall ) @@ -17,6 +18,20 @@ if ([Environment]::OSVersion.Version.Major -lt 10) { throw "The Eve server profile requires Windows 11." } +if ($RemoteGPUClient) { + if ($OpenFirewall) { + $ruleName = "Trinity Voice Core 18767" + if (-not (Get-NetFirewallRule -DisplayName $ruleName -ErrorAction SilentlyContinue)) { + New-NetFirewallRule -DisplayName $ruleName -Direction Inbound -Action Allow -Protocol TCP -LocalPort 18767 -Profile Private -ErrorAction Stop | Out-Null + } + Write-Host "Private-network firewall rule created for TCP 18767." + } + Write-Host "Windows remote-GPU mode prepared; no CUDA voice packages were installed in the VM." + Write-Host "Select 'Windows-VM mit Eve auf einem Ubuntu-GPU-Host' in Trinity Settings." + Write-Host "Use ws://UBUNTU-TAILSCALE-IP:8766/v1/realtime and a separate Core token." + exit 0 +} + & $python -m pip install "speech-to-speech==0.2.11" "websockets>=14,<18" if ($LASTEXITCODE -ne 0) { throw "Voice dependencies could not be installed." } @@ -41,7 +56,9 @@ Write-Host $(if ($tailscale) { "Tailscale detected." } else { "Tailscale not det if ($OpenFirewall) { $ruleName = "Trinity Eve Voice 8766" - New-NetFirewallRule -DisplayName $ruleName -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8766 -Profile Private -ErrorAction Stop | Out-Null + if (-not (Get-NetFirewallRule -DisplayName $ruleName -ErrorAction SilentlyContinue)) { + New-NetFirewallRule -DisplayName $ruleName -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8766 -Profile Private -ErrorAction Stop | Out-Null + } Write-Host "Private-network firewall rule created for TCP 8766." } diff --git a/tests/test_ambient_context.py b/tests/test_ambient_context.py new file mode 100644 index 0000000..a476aee --- /dev/null +++ b/tests/test_ambient_context.py @@ -0,0 +1,113 @@ +import json +import threading +from datetime import datetime, timezone +from http.server import ThreadingHTTPServer +from urllib.parse import quote +from urllib.request import Request, urlopen + +from ambient_context import AmbientContextService +from trinity_bridge import TrinityBridge, make_handler + + +def test_ambient_context_uses_latest_rss_item_and_fallback_weather(): + rss = b""" + AltFri, 31 Jul 2026 10:00:00 +0000 + Neueste MeldungFri, 31 Jul 2026 12:00:00 +0000 + """ + + def fetch(url): + if "n-tv.de" in url: + return rss + if "geocoding" in url: + return json.dumps({"results": [{"latitude": 48.68, "longitude": 9.22, "name": "Filderstadt"}]}).encode() + return json.dumps({"hourly": {"temperature_2m": [20.1, 21.6], "weather_code": [1, 3]}}).encode() + + now = datetime(2026, 7, 31, 13, tzinfo=timezone.utc).timestamp() + result = AmbientContextService(fetch=fetch, clock=lambda: now).snapshot("Filderstadt") + + assert result["headline"]["title"] == "Neueste Meldung" + assert result["weather"] == { + "place": "Filderstadt", "temperature": 22, "code": 3, "symbol": "CLOUD", "source": "fallback" + } + + +def test_ambient_context_prefers_fresh_device_location_and_calendar_without_persisting_it(): + calls = [] + + def fetch(url): + calls.append(url) + if "n-tv.de" in url: + return b"Headline" + return json.dumps({"hourly": {"temperature_2m": [17, 18], "weather_code": [0, 1]}}).encode() + + now = 1_800_000_000.0 + service = AmbientContextService(fetch=fetch, clock=lambda: now) + service.report_device( + { + "latitude": 52.52, + "longitude": 13.405, + "calendar_title": "Vorlesung Wirtschaftsinformatik", + "calendar_start": now + 3600, + } + ) + result = service.snapshot("Filderstadt") + + assert result["weather"]["source"] == "device" + assert result["weather"]["temperature"] == 18 + assert result["calendar"]["title"] == "Vorlesung Wirtschaftsinformatik" + assert not any("geocoding" in url for url in calls) + + +def test_ambient_context_returns_partial_data_when_public_services_fail(): + def fetch(_url): + raise OSError("offline") + + result = AmbientContextService(fetch=fetch).snapshot("Filderstadt") + + assert result["ok"] is True + assert result["headline"] == {} + assert result["weather"] == {} + + +def test_ambient_http_endpoints_accept_authenticated_device_context(tmp_path): + (tmp_path / "core").mkdir() + (tmp_path / "memory").mkdir() + bridge = TrinityBridge(tmp_path, token="secret") + + class FakeAmbient: + def __init__(self): + self.device = None + + def report_device(self, payload): + self.device = payload + return {"ok": True} + + def snapshot(self, fallback_place): + return {"ok": True, "place": fallback_place, "device": self.device} + + bridge.ambient = FakeAmbient() + server = ThreadingHTTPServer(("127.0.0.1", 0), make_handler(bridge)) + thread = threading.Thread(target=server.serve_forever, daemon=True) + thread.start() + try: + base_url = f"http://127.0.0.1:{server.server_port}" + report = Request( + f"{base_url}/ambient/device", + data=json.dumps({"latitude": 48.68, "longitude": 9.22}).encode("utf-8"), + headers={"Authorization": "Bearer secret", "Content-Type": "application/json"}, + method="POST", + ) + with urlopen(report, timeout=3) as response: + assert json.load(response)["ok"] is True + + snapshot = Request( + f"{base_url}/ambient?place={quote('Filderstadt Plattenhardt')}", + headers={"Authorization": "Bearer secret"}, + ) + with urlopen(snapshot, timeout=3) as response: + payload = json.load(response) + assert payload["place"] == "Filderstadt Plattenhardt" + assert payload["device"] == {"latitude": 48.68, "longitude": 9.22} + finally: + server.shutdown() + server.server_close() diff --git a/tests/test_bridge_audio.py b/tests/test_bridge_audio.py index 6385c33..24e95b9 100644 --- a/tests/test_bridge_audio.py +++ b/tests/test_bridge_audio.py @@ -7,7 +7,12 @@ import numpy as np import pytest -from bridge_audio import BridgeAudioTranscriber, G2_SAMPLE_RATE +from bridge_audio import ( + BridgeAudioTranscriber, + G2_SAMPLE_RATE, + TRINITY_HOTWORDS, + TRINITY_VOCABULARY, +) from trinity_bridge import TrinityBridge, make_handler @@ -66,6 +71,14 @@ def transcribe(self, _audio, **kwargs): transcriber.transcribe(encoded, quality="maximum") +def test_bridge_audio_prompt_does_not_bias_spontaneous_checklist_hallucinations(): + prompt = f"{TRINITY_VOCABULARY} {TRINITY_HOTWORDS}".lower() + + assert "checkliste" not in prompt + assert "wichtige begriffe" not in prompt + assert "stichwoerter" not in prompt + + def test_audio_transcription_http_endpoint_accepts_authenticated_g2_request(tmp_path): (tmp_path / "core").mkdir() (tmp_path / "memory").mkdir() diff --git a/tests/test_runtime_resilience.py b/tests/test_runtime_resilience.py index bb21bbd..13a908e 100644 --- a/tests/test_runtime_resilience.py +++ b/tests/test_runtime_resilience.py @@ -100,6 +100,20 @@ def test_windows_speech_can_be_enabled_explicitly(tmp_path, monkeypatch): assert ear.speech_input_enabled is True +def test_desktop_response_is_queued_for_running_eve_client(tmp_path, monkeypatch): + runtime_voice = tmp_path / "TrinityRuntime" / "voice" + runtime_voice.mkdir(parents=True) + (runtime_voice / "desktop_eve_audio.ready").write_text("123", encoding="utf-8") + monkeypatch.setattr(transcriber, "PROJECT_DIR", str(tmp_path)) + ear = object.__new__(transcriber.TrinityEar) + + assert ear._queue_eve_desktop_speech("Hallo von Eve") is True + + queue = runtime_voice / "desktop_speech_queue.jsonl" + payload = json.loads(queue.read_text(encoding="utf-8").strip()) + assert payload["text"] == "Hallo von Eve" + + def test_runtime_reload_applies_saved_settings(tmp_path, monkeypatch): config_path = tmp_path / "config.json" config_path.write_text( @@ -152,7 +166,7 @@ def test_runtime_reload_applies_saved_settings(tmp_path, monkeypatch): assert ear.reload_config_if_changed() is True assert ear.voice == "New Voice" - assert ear.mode == "chat" + assert ear.mode == "office" assert ear.telegram_cfg["enabled"] is True diff --git a/tests/test_trinity_bridge.py b/tests/test_trinity_bridge.py index b9fc7af..904e5b0 100644 --- a/tests/test_trinity_bridge.py +++ b/tests/test_trinity_bridge.py @@ -79,6 +79,30 @@ def test_bridge_runtime_updates_saved_config(tmp_path): assert bridge.get_runtime()["mode"] == "lecture" +def test_bridge_speaker_claim_is_persisted_and_exposed_in_instance_state(tmp_path): + bridge = TrinityBridge(tmp_path) + + claimed = bridge.set_speaker( + { + "device_id": "companion:ipad-lecture", + "label": "iPad Vorlesung", + "kind": "companion", + } + ) + + assert claimed["ok"] is True + assert bridge.get_speaker()["device_id"] == "companion:ipad-lecture" + assert bridge.instance_state()["speaker"]["label"] == "iPad Vorlesung" + + released = bridge.set_speaker( + {"device_id": "ignored", "label": "ignored", "kind": "none"} + ) + + assert released["device_id"] == "none" + assert released["label"] == "Stumm" + assert bridge.get_speaker()["kind"] == "none" + + def test_bridge_web_settings_round_trip_and_keeps_unknown_values(tmp_path): home = tmp_path (home / "core").mkdir() @@ -479,6 +503,23 @@ def transcribe(self, audio_base64, **kwargs): assert events[0]["text"] == "Trinity erklaere Spieltheorie" +def test_bridge_preserves_g2_output_target_in_stt_source(tmp_path): + home = tmp_path + (home / "core").mkdir() + (home / "memory").mkdir() + bridge = TrinityBridge(home) + + bridge.send_stt({ + "text": "Trinity, bitte erklaeren.", + "source": "g2-stt-companion", + "is_final": True, + "speak": False, + }) + + events = pop_external_stt_events(home / "core" / "ios_stt_feed.jsonl") + assert events[0]["source"] == "g2-stt-companion" + + def test_bridge_transcribes_g2_audio_and_routes_continuous_conversation(tmp_path): home = tmp_path (home / "core").mkdir() diff --git a/tests/voice/test_local_realtime_client.py b/tests/voice/test_local_realtime_client.py index 6c67cee..38adbc7 100644 --- a/tests/voice/test_local_realtime_client.py +++ b/tests/voice/test_local_realtime_client.py @@ -1,143 +1,24 @@ -import base64 -import json -import time +from core.voice.config import default_voice_config, load_voice_config +from core.voice.local_realtime_client import LocalRealtimeAudioClient -import numpy as np -from voice.config import default_voice_config, load_voice_config -from voice.local_realtime_client import LocalRealtimeAudioClient - - -def client(tmp_path): - reference = tmp_path / "eve.mp3" - reference.write_bytes(b"voice") +def test_remote_connection_uri_adds_token_and_preserves_query(tmp_path): raw = default_voice_config() - raw.update({ - "engine": "eve", - "profile": "eve-mac-local", - "reference_audio": str(reference), - }) - return LocalRealtimeAudioClient(load_voice_config(tmp_path, {"voice": raw})) - - -def test_session_enables_server_side_interruption(tmp_path): - update = client(tmp_path)._session_update() - - turn_detection = update["session"]["audio"]["input"]["turn_detection"] - assert turn_detection["interrupt_response"] is True - assert turn_detection["create_response"] is True - - -def test_speech_started_flushes_buffered_audio(tmp_path): - local = client(tmp_path) - local._output.extend(b"audio") - local._handle_event(json.dumps({"type": "input_audio_buffer.speech_started"})) - - assert local._output == bytearray() - - -def test_audio_delta_is_buffered_for_playback(tmp_path): - local = client(tmp_path) - audio = b"\x01\x00" * 32 - local._handle_event(json.dumps({ - "type": "response.output_audio.delta", - "delta": base64.b64encode(audio).decode("ascii"), - })) - - assert bytes(local._output) == audio - - -def test_obvious_playback_echo_is_not_forwarded(tmp_path): - local = client(tmp_path) - samples = np.full(512, 2_000, dtype=np.int16) - local._played_output.append(samples.copy()) - local._last_output_at = time.monotonic() - - assert local._should_forward_microphone(samples.tobytes()) is False - - -def test_distinct_loud_speech_can_interrupt_playback(tmp_path): - local = client(tmp_path) - local._played_output.append(np.full(512, 2_000, dtype=np.int16)) - local._last_output_at = time.monotonic() - speech = np.tile(np.array([1_400, -1_100, 800, -500], dtype=np.int16), 128) - - assert local._should_forward_microphone(speech.tobytes()) is True - - -def test_distinct_speech_cancels_current_response_before_forwarding(tmp_path): - local = client(tmp_path) - local._output.extend(b"buffered-audio") - local._played_output.append(np.full(512, 2_000, dtype=np.int16)) - local._last_output_at = time.monotonic() - speech = np.tile(np.array([1_400, -1_100, 800, -500], dtype=np.int16), 128) - - for _ in range(4): - local._input_callback(speech.tobytes(), 512, None, None) - - assert local._output == bytearray() - assert local._send_queue.get_nowait() == {"type": "response.cancel"} - forwarded = [local._send_queue.get_nowait() for _ in range(4)] - assert all(event["type"] == "input_audio_buffer.append" for event in forwarded) - - -def test_single_distinct_block_does_not_false_trigger_barge_in(tmp_path): - local = client(tmp_path) - local._output.extend(b"buffered-audio") - local._played_output.append(np.full(512, 2_000, dtype=np.int16)) - local._last_output_at = time.monotonic() - speech = np.tile(np.array([1_400, -1_100, 800, -500], dtype=np.int16), 128) - - local._input_callback(speech.tobytes(), 512, None, None) - - assert local._output == bytearray(b"buffered-audio") - assert local._send_queue.empty() - - -def test_interrupt_keeps_echo_history_for_the_speaker_tail(tmp_path): - local = client(tmp_path) - playback = np.tile(np.array([1_500, -1_000, 600, -300], dtype=np.int16), 128) - local._played_output.append(playback.copy()) - local._output.extend(playback.tobytes()) - local._last_output_at = time.monotonic() - - local._clear_output() - local._last_output_at = time.monotonic() - - assert len(local._played_output) == 1 - assert local._should_forward_microphone(playback.tobytes()) is False - - -def test_completed_barge_in_does_not_cancel_the_following_response(tmp_path): - local = client(tmp_path) - playback = np.tile(np.array([1_500, -1_000, 600, -300], dtype=np.int16), 128) - speech = np.tile(np.array([900, 1_600, -1_300, -700], dtype=np.int16), 128) - local._played_output.append(playback.copy()) - local._output.extend(playback.tobytes()) - local._last_output_at = time.monotonic() - - for _ in range(4): - local._input_callback(speech.tobytes(), 512, None, None) - while not local._send_queue.empty(): - local._send_queue.get_nowait() - - local._played_output.append(playback.copy()) - local._output.extend(playback.tobytes()) - local._last_output_at = time.monotonic() - for _ in range(10): - local._input_callback(playback.tobytes(), 512, None, None) - - assert local._send_queue.empty() - assert local._output == bytearray(playback.tobytes()) - - -def test_output_callback_consumes_audio_without_microphone_coupling(tmp_path): - local = client(tmp_path) - samples = np.full(512, 1_250, dtype=np.int16) - local._output.extend(samples.tobytes()) - output = bytearray(samples.nbytes) - - local._output_callback(output, 512, None, None) - - assert bytes(output) == samples.tobytes() - assert local._output == bytearray() + raw.update( + { + "profile": "eve-windows-remote", + "access_token": "voice secret", + "remote_voice_url": "wss://voice.example.test/v1/realtime?client=windows", + "backend_token": "core-secret", + } + ) + config = load_voice_config(tmp_path, {"voice": raw}) + client = LocalRealtimeAudioClient( + config, + endpoint=config.remote_voice_url, + access_token=config.access_token, + ) + + assert client._connection_uri() == ( + "wss://voice.example.test/v1/realtime?client=windows&access_token=voice+secret" + ) diff --git a/tests/voice/test_trinity_conversation_backend.py b/tests/voice/test_trinity_conversation_backend.py new file mode 100644 index 0000000..276c4b0 --- /dev/null +++ b/tests/voice/test_trinity_conversation_backend.py @@ -0,0 +1,55 @@ +import json + +from voice.conversation.trinity_backend import TrinityConversationBackend + + +class FakeBrain: + def __init__(self): + self.queries = [] + + def ask(self, query, *_args, **_kwargs): + self.queries.append(query) + return "Verstanden.", False + + +def backend_for(tmp_path, mode): + core = tmp_path / "core" + core.mkdir() + (core / "config.json").write_text( + json.dumps({ + "system": {"mode": mode}, + "persona": {"trigger_variants": ["trinity", "triniti"]}, + }), + encoding="utf-8", + ) + backend = TrinityConversationBackend(tmp_path) + backend._brain = FakeBrain() + backend._append_chat_events = lambda *_args, **_kwargs: None + return backend + + +def test_lecture_voice_ignores_speech_without_wakeword(tmp_path): + backend = backend_for(tmp_path, "lecture") + + assert list(backend.respond("Heute geht es um Spieltheorie")) == [] + assert backend._brain.queries == [] + + +def test_lecture_voice_answers_after_fuzzy_wakeword(tmp_path): + backend = backend_for(tmp_path, "lecture") + + assert list(backend.respond("Triniti, erkläre das Nash-Gleichgewicht")) == ["Verstanden."] + assert backend._brain.queries == ["Triniti, erkläre das Nash-Gleichgewicht"] + + +def test_office_voice_answers_without_wakeword(tmp_path): + backend = backend_for(tmp_path, "office") + + assert list(backend.respond("Erkläre das Nash-Gleichgewicht")) == ["Verstanden."] + assert backend._brain.queries == ["Erkläre das Nash-Gleichgewicht"] + + +def test_legacy_chat_mode_behaves_like_office(tmp_path): + backend = backend_for(tmp_path, "chat") + + assert list(backend.respond("Erkläre das Nash-Gleichgewicht")) == ["Verstanden."] diff --git a/tests/voice/test_voice_command.py b/tests/voice/test_voice_command.py index 7503eb7..75bd6ca 100644 --- a/tests/voice/test_voice_command.py +++ b/tests/voice/test_voice_command.py @@ -45,3 +45,24 @@ def test_windows_profile_uses_cuda_compatible_models(tmp_path): assert command[command.index("--parakeet_tdt_model_name") + 1] == "nvidia/parakeet-tdt-0.6b-v3" assert command[command.index("--qwen3_tts_model_name") + 1] == "Qwen/Qwen3-TTS-12Hz-1.7B-Base" assert command[command.index("--qwen3_tts_backend") + 1] == "torch" + + +def test_ubuntu_server_uses_remote_windows_trinity_core(tmp_path): + reference = tmp_path / "Eve.mp3" + reference.write_bytes(b"voice") + raw = default_voice_config() + raw.update({ + "engine": "eve", + "profile": "eve-linux-gpu-server", + "access_token": "voice-secret", + "reference_audio": str(reference), + "remote_core_base_url": "http://100.64.0.20:18767/v1", + "remote_core_api_key": "core-secret", + }) + config = load_voice_config(tmp_path, {"voice": raw}) + + command = build_speech_to_speech_command(config) + + assert command[command.index("--responses_api_base_url") + 1] == "http://100.64.0.20:18767/v1" + assert command[command.index("--responses_api_api_key") + 1] == "core-secret" + assert command[command.index("--model_name") + 1] == "trinity-core" diff --git a/tests/voice/test_voice_config.py b/tests/voice/test_voice_config.py index 567d58f..dd0cd0f 100644 --- a/tests/voice/test_voice_config.py +++ b/tests/voice/test_voice_config.py @@ -102,6 +102,49 @@ def test_mobile_server_profiles_bind_externally_without_desktop_audio(tmp_path): assert windows.profile.local_audio is False +def test_ubuntu_gpu_server_routes_back_to_windows_core(tmp_path): + audio = tmp_path / "Eve.mp3" + audio.write_bytes(b"voice") + config = load_voice_config( + tmp_path, + { + "voice": { + "engine": "eve", + "profile": "eve-linux-gpu-server", + "access_token": "voice-secret", + "reference_audio": str(audio), + "remote_core_base_url": "http://100.64.0.20:18767/v1", + "remote_core_api_key": "core-secret", + } + }, + ) + + assert config.profile.runtime_role == "server" + assert config.profile.conversation_backend == "remote" + assert config.profile.device == "cuda" + assert config.validate() == [] + + +def test_windows_remote_profile_needs_no_local_voice_models(tmp_path): + config = load_voice_config( + tmp_path, + { + "voice": { + "engine": "eve", + "profile": "eve-windows-remote", + "access_token": "voice-secret", + "remote_voice_url": "ws://100.64.0.10:8766/v1/realtime", + "backend_host": "0.0.0.0", + "backend_token": "core-secret", + } + }, + ) + + assert config.profile.runtime_role == "client" + assert config.profile.local_audio is True + assert config.validate() == [] + + def test_eve_requires_reference_audio(tmp_path): config = load_voice_config(tmp_path, {"voice": {"engine": "eve"}}) diff --git a/trinity_classic.py b/trinity_classic.py index 75a7408..a2f3a13 100644 --- a/trinity_classic.py +++ b/trinity_classic.py @@ -4,6 +4,7 @@ import html import json import os +import platform import re import sys import threading @@ -284,6 +285,7 @@ def __init__(self): self.remote_events = [] self.remote_after = 0.0 self._remote_next_poll = 0.0 + self._speaker_next_refresh = 0.0 self._workspace_sidebar_signature = None self._workspace_sidebar_next_refresh = 0.0 self.memory_store = MemoryStore(os.path.join(MEMORY_DIR, "trinity_memory.sqlite3")) @@ -340,7 +342,7 @@ def __init__(self): self.new_session_button.clicked.connect(self.start_new_session) self.mode_combo = QComboBox() self.mode_combo.setObjectName("toolbarCombo") - self.mode_combo.addItems(["lecture", "office", "chat"]) + self.mode_combo.addItems(["lecture", "office"]) self.mode_combo.setFixedWidth(96) self.mode_combo.setToolTip("Trinity-Betriebsmodus") self.mode_combo.currentTextChanged.connect(self.set_runtime_mode) @@ -348,10 +350,13 @@ def __init__(self): self.audio_source_button.setObjectName("subtle") self.audio_source_button.setFixedSize(46, 38) self.audio_source_button.clicked.connect(self.toggle_audio_capture_mode) - self.tts_button = QPushButton() - self.tts_button.setObjectName("subtle") - self.tts_button.setFixedSize(46, 38) - self.tts_button.clicked.connect(self.toggle_tts) + self.speaker_button = QPushButton() + self.speaker_button.setObjectName("subtle") + self.speaker_button.setFixedSize(46, 38) + self.speaker_button.setToolTip( + "Diesen Desktop als einzige Trinity-Sprachausgabe auswählen" + ) + self.speaker_button.clicked.connect(self.toggle_desktop_speaker) self.theme_button = QPushButton() self.theme_button.setObjectName("theme") self.theme_button.setFixedSize(46, 38) @@ -383,7 +388,7 @@ def __init__(self): right_cluster_layout.setSpacing(4) right_cluster_layout.addWidget(self.mode_combo) right_cluster_layout.addWidget(self.audio_source_button) - right_cluster_layout.addWidget(self.tts_button) + right_cluster_layout.addWidget(self.speaker_button) right_cluster_layout.addWidget(self.theme_button) right_cluster_layout.addWidget(settings_button) header.addWidget(right_cluster) @@ -1400,6 +1405,7 @@ def _apply_style(self): """) def refresh(self): + self._refresh_speaker_control() if self.remote_client: self._refresh_remote_chat() self._refresh_workspace_views() @@ -1856,6 +1862,68 @@ def _runtime_values(self): "tts_enabled": bool(system.get("tts_enabled", True)), } + def _desktop_speaker_identity(self): + profile = self.session_store.profile.lower() + hostname = platform.node().strip() or "Desktop" + return { + "device_id": f"desktop:{profile}:{hostname}", + "label": f"Trinity Desktop · {hostname}", + "kind": "desktop", + } + + def _speaker_values(self): + if self.remote_client: + return self.remote_client.get_speaker() + return {"ok": True, **TrinityBridge(BASE_DIR).get_speaker()} + + def _refresh_speaker_control(self, force=False): + if not hasattr(self, "speaker_button"): + return + if not force and time.monotonic() < self._speaker_next_refresh: + return + self._speaker_next_refresh = time.monotonic() + 1.2 + try: + selected = self._speaker_values() + except RuntimeError: + return + identity = self._desktop_speaker_identity() + active = selected.get("device_id") == identity["device_id"] + self.speaker_button.setText("🔊" if active else "🔇") + if active: + self.speaker_button.setToolTip( + "Dieser Desktop spricht. Klicken, um Trinity überall stummzuschalten." + ) + else: + label = str(selected.get("label") or "ein anderes Gerät") + self.speaker_button.setToolTip( + f"Aktuell spricht Trinity auf: {label}. Klicken, um die Ausgabe hierher zu holen." + ) + + def toggle_desktop_speaker(self): + identity = self._desktop_speaker_identity() + try: + selected = self._speaker_values() + active = selected.get("device_id") == identity["device_id"] + if active and self.remote_client: + result = self.remote_client.release_speaker() + elif active: + result = TrinityBridge(BASE_DIR).set_speaker( + {"device_id": "none", "label": "Stumm", "kind": "none"} + ) + elif self.remote_client: + result = self.remote_client.set_speaker(**identity) + else: + result = TrinityBridge(BASE_DIR).set_speaker(identity) + except (RuntimeError, ValueError) as exc: + self.status.setText(f"Sprechstelle konnte nicht gewählt werden: {exc}") + return + self._speaker_next_refresh = 0.0 + self._refresh_speaker_control(force=True) + self.status.setText( + "Trinity spricht nicht" if result.get("kind") == "none" + else f"Trinity spricht jetzt hier: {result.get('label')}" + ) + def _set_runtime_values(self, updates): if self.remote_client: try: @@ -1872,7 +1940,6 @@ def _set_runtime_values(self, updates): def _sync_runtime_controls(self, values=None): values = values or self._runtime_values() microphone_enabled = bool(values.get("microphone_enabled", True)) - tts_enabled = bool(values.get("tts_enabled", True)) audio_capture_mode = str(values.get("audio_capture_mode", "mic_only") or "mic_only") self.listen_button.setText("🎙" if microphone_enabled else "🔇") self.listen_button.setToolTip("Mikrofon aktiv" if microphone_enabled else "Mikrofon pausiert") @@ -1883,11 +1950,11 @@ def _sync_runtime_controls(self, values=None): else "Nur eigenes Mikro" ) self.new_session_button.setText("+") - self.tts_button.setText("🔊" if tts_enabled else "🔈") - self.tts_button.setToolTip("Desktop-TTS aktiv" if tts_enabled else "Desktop-TTS pausiert") mode = values.get("mode", "lecture") self.mode_combo.blockSignals(True) - self.mode_combo.setCurrentText(mode if mode in {"lecture", "office", "chat"} else "lecture") + if mode == "chat": + mode = "office" + self.mode_combo.setCurrentText(mode if mode in {"lecture", "office"} else "lecture") self.mode_combo.blockSignals(False) def toggle_microphone(self): diff --git a/trinity_cli.py b/trinity_cli.py index e36edfc..1b6c18a 100644 --- a/trinity_cli.py +++ b/trinity_cli.py @@ -10,7 +10,7 @@ from pathlib import Path -VERSION = "0.17.2" +VERSION = "0.17.9" def find_trinity_home(explicit=None):