跳到主要内容
    ↑↓ 选择↵ 打开esc 关闭
    safwan-o

    Voice

    @safwan-o/opencode-voice·v0.5.0·界面与主题

    Local voice input plugin for OpenCode (V2 TUI port)

    GitHub 星标

    0

    月装机量

    762

    近 7 天 128

    综合评分

    37.3

    生态多维模型

    最近提交

    14 天前

    2026-09-21

    快速安装与配置

    opencode.json

    写入当前项目的 opencode.json,只对这个仓库生效。

    opencode.json

    {
      "$schema": "https://opencode.ai/config.json",
      "plugin": ["@safwan-o/opencode-voice@0.5.0"]
    }

    OpenCode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。

    opencode voice logo

    Local speech-to-text for the OpenCode TUI.

    CI opencode status npm version npm downloads (monthly) npm downloads (total) license stt

    [!IMPORTANT] This is a dual-version TUI plugin: OpenCode V2 loads it via Plugin.define (./tui entry, cli.json), OpenCode V1 via {id, tui} (main entry, opencode plugin command). On 2.x it must not be installed with opencode plugin add (that registers server plugins, which this package is not). Follow the per-version install below.

    Demo

    Slash commands Voice settings Model picker
    slash commands voice settings model picker
    /voice, /voice-submit, /voice-stop, /voice-settings Record key, transcription, audio & system 87 local models with size and engine at a glance

    English | Русский | 简体中文 | Español


    Install

    OpenCode 2.x

    OpenCode 2.x loads TUI plugins from cli.json (~/.config/opencode/cli.json):

    {
      "plugins": ["@safwan-o/opencode-voice"]
    }
    

    Then fully quit the TUI and reopen it twice: the first boot downloads and installs the package in the background, the second boot runs it. (/restart only restarts the session — plugin code loads at TUI startup.)

    [!NOTE] tui.json is legacy and ignored by OpenCode 2.x — do not list the plugin there. Do not use opencode plugin add either: that registers server-side plugins and this package is CLI-only, so it would report a load failure.

    OpenCode 1.x

    OpenCode 1.x manages TUI plugins through file entries in tui.json — npm specs only feed the server loader there, so the npm package name alone will not load the voice commands (you would see must default export an object with server() in the logs instead). Register a local entry:

    mkdir -p ~/opencode-voice-v1
    cd ~/opencode-voice-v1
    ln -s ~/opencode-voice/index.js index.js
    ln -s ~/opencode-voice/lib lib
    printf '%s' '{"name":"@safwan-o/opencode-voice-v1-shim","version":"0.5.0","type":"module","main":"./index.js"}' > package.json
    

    (The minimal package.json is load-bearing: V1 silently skips file plugins whose exports map contains a ./tui subpath, and symlinks keep the entry tracking this repo. Do not add an exports map to the shim.)

    Then, per project, register the entry (project-local, invisible to V2):

    cd <project> && ocv1 plugin ~/opencode-voice-v1
    

    Then restart the V1 TUI. If you change the shim's version without re-registering, V1 keeps the old verdict — re-run the plugin command (or bump + re-register) to refresh it.

    Version differences

    Area OpenCode 1.x OpenCode 2.x
    Install local file entry in tui.json (see below) cli.json plugins array
    Entry point main → ./index.js ({id, tui}) ./tui → src/tui.ts (Plugin.define)
    Transcription lands appended into the prompt for in-place review clipboard for pasting, or submitted via /voice-submit
    Submit key extra submitHotkey binding available slash command only

    Shared on both versions: models, engines, cleanup + cutoff, ctrl+space default, and the four user-facing commands (V1 additionally exposes two hidden hold-to-talk commands). Only picker row format and diagnostics text differ cosmetically.

    On first launch, choose a model. The plugin downloads its required local runtime and model weights automatically. Audio and transcription stay on your machine.

    Optional CLI installer. It runs the same OpenCode plugin install command and pre-downloads the managed engine:

    npx @safwan-o/opencode-voice install
    

    Update an installed plugin to the latest published version, then restart OpenCode:

    npx @safwan-o/opencode-voice update
    

    Add --global if the plugin was installed in OpenCode's global configuration.

    Do not clone the repo unless you want to develop the plugin.

    [!TIP] First launch opens a model picker. Choose a local model, let it download, then use ctrl+space to dictate (change it in /voice-settings). On 2.x with auto-submit off, transcriptions land in your clipboard for pasting; on 1.x they are appended into the prompt for in-place review.

    What It Installs

    The plugin manages the STT engine and models:

    • downloads whisper.cpp or the Rust transcribe-cpp sidecar from the opencode-voice GitHub Release registry
    • stores each runtime in ~/.cache/opencode-voice/engines/<engine>/<platform>-<arch>/
    • downloads the selected model on first setup

    Manual runtime installation is optional. Existing whisper-cli or opencode-voice-transcribe binaries can also be imported through the CLI.

    Check your machine:

    npx @safwan-o/opencode-voice doctor
    

    Install or inspect a managed runtime without opening OpenCode:

    npx @safwan-o/opencode-voice engine install transcribe-cpp
    npx @safwan-o/opencode-voice engine status transcribe-cpp
    

    Use It

    Commands:

    • /voice - toggle recording; transcription is delivered for review (clipboard on 2.x, appended into the prompt on 1.x), or submits when auto-submit is on
    • /voice-submit - toggle recording and submit after transcription, even with auto-submit off
    • /voice-stop - cancel active recording or transcription
    • /voice-settings - open model, hotkey, microphone, and diagnostics settings

    Default hotkey:

    ctrl+space -> start recording
    ctrl+space -> stop, transcribe, and deliver (or submit)
    

    In /voice-settings -> Recording, choose one Record key. Toggle recording starts on the first press and transcribes on the second. (ctrl+r is intentionally not the default: OpenCode binds it to session rename.) Hold-to-talk is shown as unavailable until OpenCode exposes terminal key-release events to TUI plugins.

    Delivery: with auto-submit off (default), the transcription goes out for review and a toast confirms. On 2.x it is copied to the OS clipboard (wl-copy / pbcopy / clip / xclip / xsel, first available wins) — paste it into the main textbox with Ctrl+V, where images can also be attached. There is no composer-append API in OpenCode V2, so the clipboard is the edit path. On 1.x it is appended into the prompt for in-place review instead. With auto-submit on (or via /voice-submit), the text is sent as a user message immediately on both versions.

    Settings also let you select a microphone, language, model, download location, voice cleanup, and auto-submit.

    Voice cleanup

    Recording settings include an optional algorithmic cleanup chain (on by default):

    highpass=f=120, adeclick, dynaudnorm, alimiter
    

    It strips rumble/hum, mic pops, and levels volume before transcription, for every recorder backend. Disable it per-machine with the Voice cleanup toggle; tune the Rumble filter cutoff (40–500 Hz, default 120 Hz) if a specific room needs it. Details and measurements: docs/voice-cleanup.md.

    Models

    The picker includes 19 whisper.cpp GGML choices and 68 ASR GGUF models from Handy's transcribe.cpp catalog. Pick one model at a time; only that model is downloaded.

    Useful whisper.cpp starting points:

    Model Size Notes
    Whisper Tiny Q5_1 31 MB fastest, lowest accuracy
    Whisper Base Q5_1 57 MB small multilingual option
    Whisper Small Q5_1 181 MB compact Small
    Whisper Small 465 MB default, multilingual
    Whisper Medium Q4_1 469 MB Handy-compatible quantization
    Whisper Medium Q5_0 514 MB higher-quality multilingual
    Whisper Turbo Q5_0 547 MB compact large-v3 Turbo
    Whisper Turbo 1.5 GB large, faster than full large
    Whisper Large Q5_0 1.0 GB accurate, slower

    The managed transcribe-cpp sidecar provides these Handy catalog families. Diarization and VAD assets are excluded because they do not produce text through the transcription runtime.

    Model Size Notes
    Parakeet, Nemotron, GigaAM varies English, multilingual, and Russian models
    Moonshine varies English plus language-specific variants
    Canary, Cohere, SenseVoice, Fun-ASR varies multilingual and specialized recognizers
    Whisper, Breeze, Voxtral, Qwen, Granite varies general-purpose and language-specialized models

    Model downloads support resume, retry, progress, and SHA256 verification. Sidecar catalog URLs are pinned to an upstream Hugging Face revision and their LFS SHA-256 digest, so a model cannot be activated until its expected artifact verifies.

    [!NOTE] Large models can require multiple gigabytes of disk space and significant memory. Start with Whisper Small, GigaAM V3 for Russian, or Parakeet for a fast GGUF option.

    Troubleshooting

    Run diagnostics first:

    npx @safwan-o/opencode-voice doctor
    
    • Engine not found in registry: transcribe-cpp: the installed plugin expects a release registry that does not yet include the sidecar. Update the plugin after its matching Engine Release is published, or locally import a built sidecar with opencode-voice engine import transcribe-cpp <path>.
    • Could not start recorder on Windows: open /voice-settings and select the enumerated microphone rather than relying on the system default. The error includes the exact ffmpeg command and stderr for diagnosis.
    • Hold-to-talk is unavailable: current OpenCode releases do not expose terminal key-release events to TUI plugins. Use the default ctrl+space toggle mode instead.
    • Plugin failed in the TUI: version/entry mismatch. On 2.x, load it from cli.json as shown above — never from tui.json, and never via opencode plugin add (that registers server plugins; this package is CLI-only). On 1.x, install with opencode plugin @safwan-o/opencode-voice.
    • Empty transcriptions on noisy mics: enable Voice cleanup in Recording settings (on by default); if a specific room defeats it, raise the Rumble filter cutoff.

    Platform Status

    Platform Status
    Linux one-command engine/model install; recording uses arecord, ffmpeg, or sox
    macOS one-command engine/model install; recording uses ffmpeg AVFoundation until the native recorder sidecar ships
    Windows one-command engine/model/recorder install; recording uses DirectShow through a managed cached ffmpeg.exe, with system/bundled ffmpeg fallback

    Architecture

    This is a dual-version TUI plugin: OpenCode V2 loads Plugin.define({ id, setup }) (@opencode/plugin/tui, ./tui entry); OpenCode V1 loads { id, tui() } (main entry, index.js). Requires Node 22+ and OpenCode >= 1.17.4 (engines). Shared logic lives in lib/ as plain JS consumed by both adapters.

    • V2 entry: src/tui.ts (setup, keymap layer, delivery)
    • src/voice.ts — V2 typed facade over the shared controller (lib/voice-controller.js)
    • src/dialogs.ts — settings pickers on promise dialog APIs
    • src/settings.ts — V2 types + re-export of shared settings (lib/settings.js)
    • src/clipboard.ts — OS clipboard writers (wl-copy/pbcopy/clip/xclip/xsel)
    • src/enhance.ts — voice-cleanup chain builder (canonical logic in lib/enhance.js)
    • src/formatters.ts, src/languages.ts — display helpers, Whisper language list
    • runtime settings live in TUI plugin storage (ctx.storage.store on V2, api.kv on V1)

    Files:

    • src/tui.ts - V2 plugin entrypoint, commands, keymap layer, delivery
    • index.js - V1 plugin entrypoint, commands, dialogs, keymap layer, delivery (live load path on OpenCode 1.x)
    • lib/settings.js - shared settings normalization, hotkey migration, cutoff parsing
    • lib/voice-controller.js - shared record/transcribe controller with single-flight phase lock
    • lib/models.js - model registry, cache paths, default settings
    • lib/download.js - resumable model download and SHA256 verification
    • lib/engine.js - recorder selection, managed Windows recorder install, runtime-routed transcription, pre-transcribe cleanup
    • lib/enhance.js - cleanup filter chain builder + cutoff normalization
    • lib/engines.js - managed native engine download, status, import, and removal
    • lib/handy-model-catalog.js - pinned Handy GGUF model metadata
    • test/ - hermetic unit tests plus binary-gated fixture/integration tests (test/fixtures/ is generated, gitignored)
    • docs/ - design notes and measurements (voice-cleanup.md, v2-inventory.md, loader notes)
    • bin/opencode-voice.js - install wrapper and diagnostics CLI
    • sidecar/ - Rust transcribe-cpp command-line runtime for GGUF models

    Voice input needs native audio and STT binaries. The JS plugin manages OpenCode UI, settings, model downloads, and delivery (clipboard/submit on V2, in-place append/submit on V1). The managed Rust sidecar provides the transcribe.cpp runtime for supported GGUF model families.

    Roadmap

    • Rust recorder sidecar with cpal and VAD
    • streaming transcription for sidecar models
    • Windows recorder stability and UX polish

    Development

    Run checks:

    npm run check
    npm pack --dry-run
    cargo check --manifest-path sidecar/Cargo.toml
    

    The JavaScript plugin has no frontend build step. The optional GGUF runtime is a Rust sidecar built by the Engine Release workflow.

    Local development cannot point the v2.0.1 TUI at a checkout (file specs are silently ignored by its loader) — verify through the hermetic test suite and npm pack --dry-run instead:

    git clone https://github.com/safwan-o/opencode-voice.git opencode-voice
    cd opencode-voice
    npm install
    npm run check
    

    See CONTRIBUTING.md for the full workflow.

    Project Status

    This is an independent OpenCode plugin. It is not built by the OpenCode team and is not affiliated with OpenCode.

    Credits

    • OpenCode wordmark SVG adapted from the public OpenCode repository. The voice mark was added for this plugin.
    • Handy inspired the local-first model experience and supplies the curated transcribe.cpp model catalog this plugin pins and verifies.
    • Local Whisper transcription uses whisper.cpp.
    • Handy GGUF transcription runs through the Rust transcribe-cpp binding and its upstream transcribe.cpp runtime.
    • Recording and conversion rely on FFmpeg where the platform recorder requires it; managed distribution uses ffmpeg-static.
    • Model artifacts are hosted by Hugging Face. Individual model creators and licenses remain those declared by each upstream model repository; their pinned source URLs are retained in lib/handy-model-catalog.js.

    OpenCode Website | Docs | Discord

    同类生态推荐