Files
Skald-Circle/docs/plugins/whisper_local.md
T
dguiducci 4b1affa600
Nightly Build / build (push) Successful in 7m1s
plugins: merge the user Plugins page into per-plugin sidebar pages
The generic per-user #plugins page is gone: a plugin with per-user
settings hosts them in its own web_pages() sidebar page instead
(Telegram's pairing page is new; Honcho's opt-in page already existed).
The admin catalog moves from #plugin-catalog to #plugins (old hash
redirected), and user_config_schema is removed from the Plugin trait,
the API DTOs and both plugins — the my-config endpoint, the
plugin_user_configs store and the update_user_config hook stay, now
driven by each plugin's own page fragment.
2026-07-28 20:48:03 +01:00

38 lines
2.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Whisper Local
- **Plugin id:** `whisper_local`
- **Category:** Speech-to-text, local
- **Runs:** on this machine, in-process (via `whisper.cpp`, Metal-accelerated on Apple Silicon) — no cloud, no API key
## What it does
Transcribes voice messages entirely on-device using [whisper.cpp](https://github.com/ggerganov/whisper.cpp). Nothing leaves the machine — the right choice when privacy matters more than raw speed, or when there's no budget for a cloud transcription API.
The model (roughly 13 GB depending on size) is loaded into memory only when first needed ("lazy" loading) and unloaded again after a configurable idle period to free RAM — unless eager loading is turned on.
## Requirements
- A GGML `.bin` Whisper model file, downloaded manually onto the host filesystem — this plugin does not fetch it automatically. Example:
```
curl -L -o models/ggml-large-v3.bin \
https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.bin
```
Smaller/faster models (e.g. `ggml-medium.bin`, `ggml-small.bin`) trade accuracy for speed and size — point at any of them.
- `ffmpeg` installed on the host (used to convert incoming audio to the 16 kHz mono format Whisper needs).
## Enabling & configuring (admin)
1. Download a model file first (see above) and note its path.
2. Plugins page → **Whisper Local** → enable, then **Configure**.
3. Fields:
- **`model`** (required) — path to the `.bin` file, e.g. `models/ggml-large-v3.bin`.
- **`language`** — a BCP-47 code (`it`, `en`, …) or `auto` for automatic detection (default `auto`).
- **`load_at_startup`** — load the model into memory as soon as the plugin starts, instead of on first use (default off).
- **`idle_timeout_secs`** — unload the model from memory after this many seconds of inactivity, `0` = never unload (default `1200`, 20 minutes).
## Notes
- The model occupies roughly 13 GB of RAM while loaded.
- The very first transcription after an idle unload is slower, since the model has to reload.
- No external service and no per-use cost — the trade-off is local CPU/GPU time and disk space for the model file.