Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Rewrites written prose for the
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-speak-summary-4c81947be0f9 ,按照其中的说明把「speak-summary」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Turn written text into audio someone will actually want to listen to.
This skill is deliberately a terminal step in a chain. Another skill (or you)
produces the text; this one makes it listenable. It pairs naturally with
roundup, daily-prep, meeting-minutes, or any summarisation work.
Everything runs locally on CPU. No text is sent to a cloud speech service, which matters when the content is confidential, and it means the skill works in a headless cloud agent or CI container just as well as on a laptop.
The synthesis engine is Kyutai pocket-tts,
a small neural TTS model designed to run on CPUs.
The bundled script installs it automatically into a cached virtualenv on first use, so usually you need do nothing. To install it explicitly:
pip install pocket-tts # any platform
brew install pocket-tts # macOS, if preferred
pocket-tts requires Python >=3.10 and <3.15. The script searches for a
compatible interpreter rather than assuming python3 is one — worth knowing if
you are on a very new Python, where installation would otherwise fail.
You also need an encoder. ffmpeg is strongly preferred (brew install ffmpeg
or apt-get install -y ffmpeg); on macOS the script falls back to the built-in
afconvert and emits .m4a instead of .mp3.
The first run downloads the model (~1GB) from Hugging Face. After that it is fully offline and synthesises roughly 6x faster than real-time.
Do not feed written text straight into the synthesiser. Prose that reads well on screen is tiring to listen to. Rewriting it first is what separates a useful audio digest from an unlistenable one.
Produce a spoken script that:
Write this spoken script to its own .txt file. Keep the original written
version with its links intact — the audio is a companion to it, not a
replacement. The user will want to click through later.
./scripts/tts.sh <input.txt> <output.mp3> [voice.safetensors]
The script strips any residual markdown, splits the text on sentence boundaries into ~600 character chunks (quality degrades on long single inputs), synthesises each chunk, and concatenates the result into a mono MP3 at 96kbps — small enough to sync to a phone, good enough for speech.
Environment overrides:
| Variable | Purpose |
|---|---|
SPEAK_TTS_BIN | Path to a specific pocket-tts binary; skips all auto-detection. |
SPEAK_TTS_HOME | Where to create/find the cached virtualenv. Default ~/.cache/speak-summary/venv. |
The default English voice is alba. To use a different one, pocket-tts
supports voice cloning from a short clean audio sample:
pocket-tts export-voice --help
Pass the resulting .safetensors file as the third argument to the script.
Only clone a voice you have the rights to use. Do not clone a real person's voice — colleague, customer, or public figure — without their explicit consent.
~/Music/Briefings/ unless the user says otherwise; it is easy to point a phone or podcast app at.<subject>-<YYYY-MM-DD>.mp3.afplay <path> on macOS, ffplay -nodisp -autoexit <path> elsewhere.Aim for 4–6 minutes for a routine digest, which is roughly 600–900 spoken words at a natural pace. If the source would run past about 10 minutes, say so and offer either a tighter edit or a split into multiple files — attention drops off sharply beyond that for informational audio.
The natural pattern is gather → summarise → speak:
roundup → speak-summary — a spoken version of the status briefing.daily-prep → speak-summary — tomorrow's schedule, listened to tonight.meeting-minutes → speak-summary — catch up on a meeting you missed.When invoked as part of a chain, do not re-summarise. The upstream skill owns what to say; this skill owns how it sounds. Take its output, rewrite it for the ear, and synthesise.
To run unattended (a briefing waiting before breakfast), schedule the upstream skill with a workflow and have it finish by calling this one.
Audio cuts off mid-sentence. A chunk exceeded the model's comfortable length. Shorten the sentences in the spoken script.
Words mispronounced. Spell them phonetically in the input — "Kubernetes" as "koo-ber-net-eez". This is a normal part of preparing a spoken script.
First run is slow. That is the one-off model download. Later runs start in about a second.
pocket-tts not found after install. The virtualenv may be stale, or your
python3 may be outside the supported 3.10–3.14 range. Delete
~/.cache/speak-summary/venv and re-run, or point SPEAK_TTS_BIN at a known binary.