Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via reference_id. Use when the user prefers fish
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-fish-audio-tts-8762019d7210 ,按照其中的说明把「fish-audio-tts」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Requires FISH_AUDIO_API_KEY in .env (create one at https://fish.audio/go-api/api-keys/).
Create voice models in the fish.audio playground and pass their id as reference_id to reuse a cloned voice.
Single synchronous call returning raw audio bytes:
POST https://api.fish.audio/v1/tts
Authorization: Bearer ${FISH_AUDIO_API_KEY}
Content-Type: application/json
model: <backend model> # HTTP header selects the backend, e.g. s1
The backend model is chosen with the model HTTP header, not a body field. In OpenMontage this maps to the tool's model input.
model is required — there is no default. Pass one of:
s2.1-pro — latest generation. Best quality: inline emotion tags, 80+ languages, multi-speaker. Hero narration.s2.1-pro-free — promotional free access to s2.1-pro. Drafts, samples, and validation runs at $0 during the promo window only. Per the fish.audio announcement: free through August 31, 2026, subject to Fair Use, no SLA/latency guarantee, requests may be retained, and commercial use is restricted. Never route production or client narration through it.s2-pro — first S2 generation. Stable high quality with emotion-tag support.s1 — previous flagship. Kept for compatibility with existing integrations.Billing is per UTF-8 byte of input text (not per character). CJK text and emoji cost 3-4x an ASCII character of the same visible length. Current list pricing: s1 / s2-pro / s2.1-pro = $15 per 1M bytes, s2.1-pro-free = $0 during the promo window only (the tool's estimate_cost() switches to the paid s2.1-pro rate after August 31, 2026). Verify current pricing at https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits before large batches.
s2-pro / s2.1-pro / s2.1-pro-free interpret inline emotion tags embedded in the text:
[laugh], [whispers] change the delivery mid-sentence."That's hilarious [laugh] but let me explain seriously."s1 does not interpret emotion tags — they may be read out as plain text, so strip them when targeting s1.reference_id. The selector's generic voice_id is accepted as an alias when reference_id is absent.reference_id, fish.audio uses its default voice for the chosen model.Inline on-the-fly cloning (uploading reference audio + text per request) is not supported by this tool — create a voice model in the playground first.
Generate with the TTS selector:
from tools.audio.tts_selector import TTSSelector
result = TTSSelector().execute({
"preferred_provider": "fish_audio",
"text": "Here's why compound interest quietly beats every get-rich-quick scheme.",
"model": "s1",
"reference_id": "<playground voice model id>",
"output_path": "projects/my-video/assets/audio/narration.mp3",
})
Or call the provider directly:
from tools.audio.fish_audio_tts import FishAudioTTS
result = FishAudioTTS().execute({
"text": "Short sample line for approval.",
"model": "s1",
"reference_id": "<playground voice model id>",
"output_path": "projects/my-video/assets/audio/fish_sample.mp3",
})
The provider writes the audio to output_path and returns data.output plus the resolved model and reference_id.
latency: normal (default, best quality), balanced (a little faster), or low (fastest, slight quality cost).normalize: default true; keep it on so numbers, dates, and currency read naturally.prosody: optional { "speed": 1.0, "volume": 0 } to nudge pace/loudness.mp3_bitrate: 128 is a good default; raise to 192 for music-bed-heavy mixes.temperature: default 0.7. Raise toward 0.9 for more expressive reads (recommended when leaning on emotion tags); lower for a steadier, more predictable delivery.top_p / repetition_penalty: usually leave at the defaults (0.7 / 1.2).model + reference_id before a full paid narration.s2.1-pro-free (promo-window $0; non-commercial drafts only) and upgrade the final to s2.1-pro.401 Unauthorized: wrong or missing FISH_AUDIO_API_KEY.402 / payment errors: account credit exhausted.404 / bad voice: the reference_id is wrong or not owned by this account.text is non-empty and normalize is not stripping the whole input.Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names.