Repository navigation
Replies: 2 comments
|
Thank you @ailakshya, this is one of the most useful externally-grounded reads of the spec we've had so far :) 1. Activation modelsThe accessibility argument against hold-to-talk is well taken. We'd approaching activation as a strategy rather than a gesture. The design spec already specifies silence auto-stop (~ 15s of no voice). I'll look at Escape or typing cancels transcription too. End-of-utterance is currently client signalled. Your comments are prioritising client-side VAD more in my roadmap now. 2. Injection safetyFull agreement. Our injector is current IBus over DBus, commit-only by default. We realised after drafting the spec the current short-comings in secure-field detection and so on. We will be sure to document it. Our wire protocol makes the safety disitcintions quite explicit, transcription deltas carry a committed/unstable discriminant, unstable text is only ever displayable (preedit), committed text is append-only. Another issue I noticed, on GNOME/Wayland, text-input-v3 appears to strip IBus preedit "styling" in most apps, so a cross-app consistent "unstable look" probably can't rely on in-field styling. Treating the shell chrome as fallback for now. 3. CPU-lite compute profileAgreed, and we plan to get measurements for this. We initially comitted to batch-only mode for the MVP, but based on wider discussions, it appears there will be a lot of demand for streaming mode, so we're working on supporting that. All configurable. Had discussions about potentially running an install-time benchmark to pick a good default setting for the user, but likely such polish will come in a later iteration. 4. IndicatorAlso agree, caret anchoring is ruled for the same Wayland reasons. I prototyped a GNOME Shell extension (in-compositor, so it's focus-safe where a client-side overlay can't be). It's a little VU-level ribbon plus a content-free status label, fed by a DBus signal carrying state + normalized audio levels. Transcription content never crosses that bus. 5. Warm-up / residencyYes, the inference snaps have idle-unload with a configurable timer, we measure cold-load ("time to ready") per model in our nascent benchmark matrix, and we're proposing spec language that model residency is govered by the idle timer and never by client connection close, so a burst of presses never pays a reload. The idle-vs-latency tradeoff will be documented as a tunable. Other notes
Thanks for the offer to help, once the repo is in better shape, the injection compatibility matrix is where outside exercise would help most. Happy to compare notes further here in the meantime. |
Uh oh!
There was an error while loading. Please reload this page.
Feedback on UD129 — Ubuntu Desktop STT Integration
Thanks for opening this up for feedback before the spec locks. I recently built and daily-drove a from-scratch dictation setup on Wayland (Hyprland) — whisper.cpp with CPU and Vulkan GPU backends, PipeWire capture, VAD, and text injection via both the
virtual-keyboardprotocol (wtype) anduinput(ydotool). Several of the walls I hit map directly onto decisions in this spec, so I'm sharing concrete field notes rather than opinions.What the spec gets right (so it's clear where I'm building from)
Key feedback (roughly prioritized)
1. Push-and-hold should not be the only activation model — it's both an accessibility barrier and fragile on Wayland
The entire UX is built on "press and hold, release to stop." Two problems:
Recommendation: treat activation as a configurable strategy with (at least) three first-class modes:
Auto-stop-on-silence was the single biggest UX win in my setup, and it's the accessible default. The session lifecycle already models
Finalizing; a VAD-driven end-of-utterance transition slots in cleanly.2. Keystroke-injection backends fundamentally cannot satisfy your own safety requirements — lean into IM-protocol semantics from day one
Two of your hard requirements are "never inject into the wrong target" and "block dictation in secure/password fields." It's worth stating explicitly in the spec that these are only achievable with a text-input/IM-protocol backend (IBus today,
text-input-v3/input-method-v2tomorrow):virtual-keyboard(wtype) and especiallyuinput(ydotool) have no concept of a target. They emit keystrokes to whatever the compositor currently focuses. If focus changes mid-session, the text follows focus — directly violating "don't retarget." There is no way to bind them to the surface captured at session start.content_purpose/secure-input signal, so they can't honor the password-field block.text-input-v3exposescontent_purpose = password, which is the only portable hook you'll get.wtypealso silently no-ops in several targets (some terminals/TUIs, certain Electron/XWayland cases) — it returns success but nothing lands.So: the IBus-first / IM-native-later direction is right, but I'd (a) document that keystroke injection is disqualified for this feature's safety model so nobody later adds a
wtype/ydotoolbackend assuming parity, and (b) shape the abstraction around IM semantics now (preedit/commit/surrounding_text/content_purpose), even while the MVP only usescommit. Retrofitting preedit onto a commit-only, keystroke-shaped abstraction later is painful.One concrete warning from experience: do not approximate provisional text by typing then backspacing to "revise" words. It floods the compositor with keystrokes, and a tracking bug will delete the user's existing content. Your commit-only MVP correctly avoids this — just make sure no backend is tempted to fake preedit this way.
3. Commit-only MVP is a chance to keep compute minimal — don't pay for smooth-partial decoding you aren't displaying
Since the MVP shows no partial hypotheses in the target, you don't need continuous sliding-window re-decoding (the expensive path that smooth live-preview requires). VAD-segmented, decode-per-utterance is dramatically cheaper — in my tests continuous streaming re-decode was ~10–100× the compute of transcribe-once-per-utterance for the same committed output.
This matters for your stated goals (accessibility reach, battery, low-end hardware, "runs locally on any machine"):
base/smallq5-class or equivalent) that hits real-time on CPU. Reserve continuous streaming + large models for when partials are actually shown (a later iteration) or when a GPU/NPU is present.4. A caret-anchored indicator is largely infeasible on Wayland for the MVP — plan for panel/tray/OSD
The spec floats "a small overlay near the text input target." Wayland deliberately doesn't expose caret/cursor position to external processes — only the focused app (or an IM) knows it. So a caret-anchored indicator isn't reachable from the Speech Controller for most apps; it's only possible via the IM protocol path or the app itself. I'd make the MVP indicator a compositor OSD / panel / tray element and treat caret-anchoring as an IM-backend future item. (You hedge this already — worth making explicit so nobody scopes caret-anchoring into the MVP.)
5. Warm-up dominates perceived latency — pre-warm on key-down
Your target of "first transcription within 500–800 ms" is only reachable with the model already resident. Cold-loading even a small model is ~1 s+, and that dominates the felt latency far more than inference itself. Suggest: keep the model warm/resident while the feature is enabled (or pre-warm on hotkey-down, before audio, so load overlaps the user drawing breath), and document the idle-unload-vs-latency tradeoff explicitly as a tunable.
Smaller notes
content_purposeonly helps for IM-aware widgets. Treating terminals as always protected-by-default (opt-in to enable) may be safer than trying to detect password prompts inside them.Happy to help
If useful, I can contribute a reference implementation of the CPU-lite / VAD-segmented profile and the toggle + auto-stop-on-silence activation strategy, or help exercise the injection layer against the compatibility matrix (I have working test rigs for wtype/ydotool/IM paths). Either way — great to see this being built local-first and accessibility-first.
All reactions