Skip to content

feat(voice): add configurable local voice transcription - #14882

Closed
UtkarshUsername wants to merge 231 commits into
pingdotgg:mainfrom
UtkarshUsername:feat/local-speech-to-text
Closed

UtkarshUsername wants to merge 231 commits into
pingdotgg:mainfrom
UtkarshUsername:feat/local-speech-to-text

Conversation

@UtkarshUsername

@UtkarshUsername UtkarshUsername commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

This replaces #8928 after the orchestrator V2 rewrite, as requested in #8928 (comment).

Problem

Users need local voice transcription for voice dictation.

Change

  • Add local voice transcription to desktop and web. The client captures audio, and the selected environment server transcribes it.
  • Add a searchable catalog of transcription models with download, selection, removal, and language filters.
  • Stream transcription with live previews.
  • Prepare batch models when recording starts and unload idle models to manage resources.
  • Let users choose a transcription language or use automatic detection when the model supports it.
  • Let users choose Auto, CPU, or a specific GPU for transcription.
  • Add first-time voice setup with microphone selection and setup guidance.
  • Add a Dictionary for names, technical terms, and other uncommon words.
  • Support dictionary aliases, bulk entry, and project-specific words.
  • Use project names as transcription vocabulary.
  • Add optional filler-word removal.
  • Add configurable spoken correction cues.
  • Add transcript post-processing with editable custom instructions.
  • Use composer context to improve transcript cleanup and resolve ambiguity.
  • Support translating speech to English.
  • Add dictation keyboard shortcuts.
  • Add dictation to agent comment editors, pull request comments, and reviews, and to custom agent answers with desktop annotation controls.
  • Add microphone test.
  • Add a transcription check for the selected voice model.
  • Add queue to send immediately after transcription/post-processing
  • Keep iOS on its existing on-device transcription path.
  • Package the native transcription runtime and platform libraries with desktop and server builds, and document the voice input workflow.

Scope and approval

Feature was discussed on discord with Julius.
This is the replacement implementation requested by the maintainer in the closure comment linked above. Original work and discussion: #8928.

Verification

image image image image image image image image image image image image image image image image

Focused checks on Windows:

  • Server: vp test run src/speech src/textGeneration/TranscriptionPostProcessing.test.ts src/textGeneration/TextGeneration.test.ts src/textGeneration/OpenCode2TextGeneration.test.ts src/provider/Drivers/OpenCodeDriver.test.ts src/serverSettings.test.ts: 173 passed, 2 skipped. Includes speech HTTP, lifecycle, streaming, cancellation, models, cleanup fallback, and OpenCode 2 routing/session cleanup.

  • Contracts: vp test run src/speech.test.ts src/settings.test.ts: 171 passed.

  • Client runtime: vp test run src/voice-input: 45 passed.

  • Shared: vp test run src/serverSettings.test.ts src/projectSettings.test.ts src/speech.test.ts: 61 passed. Includes ACP Registry fallback and scoped vocabulary.

  • Web: focused speech hooks, recording/resampling, Voice settings/setup tests, microphone/model checks, settings search, scoped settings, composer controls, and keybindings: 271 passed across 15 files.

  • Desktop: vp test run src/preview/Manager.test.ts: 94 passed.

  • Packaging: vp test run lib/cli-external-packages.test.ts build-desktop-artifact.test.ts: 93 passed, 2 failed. Both failures reproduce identically on untouched main cc1e634 on Windows: Linux helper executable mode (0666 versus 0755) and the cross-architecture Windows primary-native-probe assertion.

  • Targeted formatting and lint completed with no errors. Lint reports warnings in affected files.

  • Typechecks passed for server, web, contracts, client-runtime, shared, and desktop (vp exec tsc --noEmit --pretty false with each workspace tsconfig).

Model: GPT 5.6 Sol + GPT 6 Astra + GPT 6 Sol + GPT 6.1 Sol
Harness: T3 Code Codex

t3-code Bot and others added 30 commits October 3, 2026 01:44
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
- Show recording, transcription, and download status in the composer
- Replace active send controls with voice confirmation and cancellation actions
- Move attachments alongside speech status
- Place cancel control after the speech button
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
Co-authored-by: Utkarsh Patil <73941998+UtkarshUsername@users.noreply.github.com>
- Add model catalog, downloads, cancellation, selection, and removal
- Expose model management in voice settings
- Mark Whisper-family and Breeze-ASR models as supporting recognition hints
- Surface hint support in voice settings
@UtkarshUsername UtkarshUsername changed the title [WIP] feat(voice): add configurable local voice transcription feat(voice): add configurable local voice transcription Oct 5, 2026
@SaidulBadhon

Copy link
Copy Markdown

LGTM. Please marge this one. I need it 😐

@maria-rcks

Copy link
Copy Markdown
Collaborator

Note

Written by claude-opus-5-5 on behalf of Maria

Hi! We are cleaning up open PRs, and this one names a harness (Codex) but not which model version created it. If this change is really important, we recommend rebuilding the PR with a newer model and noting the model in the PR description.

@maria-rcks maria-rcks closed this Oct 11, 2026
@maria-rcks

Copy link
Copy Markdown
Collaborator

Note

Written by claude-opus-5-5 on behalf of Maria

Reopening, this was closed by mistake. Sorry for the noise!

@maria-rcks maria-rcks reopened this Oct 11, 2026
@maria-rcks

Copy link
Copy Markdown
Collaborator

Note

Written by claude-opus-5-5 on behalf of Maria

Closing again after a second look, sorry for the back and forth. This PR has merge conflicts with main and does not say which model created it. If this change is still important, please rebuild it on current main with a newer model and note the model in the PR description.

@maria-rcks maria-rcks closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants