opencode-speech-to-textSpeech-to-text plugin for opencode — push-to-talk dictation via sox + Groq Whisper (Ctrl+')
0
385
近 7 天 175
32.0
生态多维模型
2 个月前
2026-06-04
快速安装与配置
opencode.json写入当前项目的 opencode.json,只对这个仓库生效。
opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-speech-to-text@0.1.1"]
}写入 ~/.config/opencode/opencode.json,对所有项目生效。
~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-speech-to-text@0.1.1"]
}若你要在本地改造这个插件,先装到项目里再从本地路径引用。
shell
pnpm add -D opencode-speech-to-textopencode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。
Push-to-talk speech-to-text for opencode — record with Ctrl+', transcribe with Groq Whisper, and drop the cleaned text straight into your prompt.
Features
- Push-to-talk recording via sox (Ctrl+' toggles start/stop)
- Groq Whisper transcription (OpenAI and custom endpoints too)
- LLM cleanup pass — fixes punctuation and coding homophones (pie→pi, Jason→JSON, cash→cache, …)
- Languages — auto-detect or pick one of several
- Mic picker + persistent preferences across sessions
Prerequisites
- sox —
brew install sox - GROQ_API_KEY — get one at console.groq.com
Install
The simplest way — let opencode wire up both targets and install it:
opencode plugin opencode-speech-to-text -gf
This is a TUI plugin, so its real entry must be registered in
tui.json (not opencode.json). To wire it manually, add it to the
plugin array in tui.json:
// ~/.config/opencode/tui.json
{
"$schema": "https://opencode.ai/tui.json",
"plugin": ["opencode-speech-to-text"],
}
Restart opencode after installing — TUI plugins load at startup.
Usage
Press Ctrl+' to start recording, speak, then press Ctrl+'
again to stop. The transcription is cleaned up and appended to your
prompt. You can also run /voice (works in any terminal, even when
Ctrl+' can't be transmitted).
Change the shortcut
Ctrl+' needs a terminal that transmits modified keys (kitty/CSI-u
protocol). If yours doesn't, override the key via plugin options:
// tui.json
{ "plugin": [["opencode-speech-to-text", { "keybind": "ctrl+g" }]] }
Commands
| Command | Description |
|---|---|
/voice |
Toggle recording (start/stop) |
/voice-cancel |
Cancel the current recording |
/voice-provider |
Switch transcription provider |
/voice-model |
Pick the Whisper model |
/voice-mic |
Pick the audio input device |
/voice-language |
Set language or auto-detect |
Providers
| Provider | API key env | Default model |
|---|---|---|
| Groq | GROQ_API_KEY |
whisper-large-v3 |
| OpenAI | OPENAI_API_KEY |
whisper-1 |
| Custom | (configurable) | whisper-large-v3 |
Cleanup runs on Groq Llama 3.3 70B (falls back to 3.1 8B) and also
uses GROQ_API_KEY.
Development
pnpm install
pnpm run build
pnpm run lint
License
MIT