-
Notifications
You must be signed in to change notification settings - Fork 0
Text Editor
The Text Editor opens when you click the ✏ edit button on any playlist item. It lets you read and modify the text for that item, and apply AI-powered writing tools before it goes to the TTS engine.

Changes made in the Text Editor only affect what KoKoFish speaks — your original file on disk is never modified.
Click the ✏ edit icon on any playlist item in the Speech Lab. The editor opens with the full text of that item loaded and ready to edit.
The right side of the editor has tabs for each AI tool. All AI functions require the LLM model to be downloaded (see LLM Models).
Available on Fish-Speech engines only.
Insert emotion and effect tags directly into the text to control how Fish-Speech delivers specific lines.
Emotion Tags (insert at start of a sentence or phrase):
| Tag | Effect |
|---|---|
(excited) |
Excited, high energy |
(happy) |
Happy, upbeat |
(satisfied) |
Content, calm satisfaction |
(confident) |
Assertive, steady |
(gentle) |
Soft, warm |
(serious) |
Measured, authoritative |
(sad) |
Melancholy, downcast |
(angry) |
Sharp, forceful |
(nervous) |
Hesitant, tense |
(fearful) |
Scared, shaky |
(surprised) |
Startled, rising inflection |
(confused) |
Uncertain, questioning |
Voice Effect Tags:
| Tag | Effect |
|---|---|
(laugh) |
Inserts a laugh |
[whisper] |
Speaks the following text in a whisper |
[breath] |
Inserts a breath sound |
(sigh) |
Inserts a sigh |
Suggest Tags — Rule-based, instant. Scans the text for punctuation and keywords to insert appropriate tags automatically. No AI model required.
Generate Tags — AI-powered. Reads the full context and places tags where the emotion genuinely fits the narrative. Requires Qwen or another LLM to be downloaded.
Kokoro users: Kokoro does not support inline tags. The Tags tab is replaced with a Kokoro info panel — use the voice dropdown and blend slider in the Speech Lab to control voice style instead.
Rewrites the text to sound more natural when spoken aloud.

The enhancement is engine-aware — it applies a different style depending on which TTS engine is active:
| Engine | Enhancement Style |
|---|---|
| Kokoro | Plain text only. Adds punctuation, breaks long sentences, expands abbreviations. No tags. |
| Fish-Speech 1.4 | Conservative tagging. One emotion tag per sentence maximum. |
| S1 Mini | Light tagging. Tags only where clearly needed — "when in doubt, omit." |
| S1 Full | Richer tagging. More expressive direction, [breath] support, full emotional range. |
What the AI does regardless of engine:
- Adds commas and em-dashes for natural pauses
- Breaks very long sentences into shorter ones
- Expands abbreviations (e.g. "Dr." → "Doctor", "vs." → "versus")
- Spells out numbers where appropriate for speech
- Preserves all meaning, names, and fictional terminology
After enhancement, a preview panel shows the suggested changes. Click Accept to apply or Discard to keep the original.
Rewrites the text in a different tone while keeping the core meaning intact.

Available Tones:
| Tone | Description |
|---|---|
| Neutral | Balanced, uncolored prose |
| Casual / Conversational | Relaxed, everyday language |
| Formal / Professional | Polished, structured |
| Dramatic / Cinematic | Heightened, evocative |
| Energetic / Upbeat | Fast-paced, enthusiastic |
| Calm / Soothing | Gentle, measured |
| Humorous / Playful | Light-hearted, fun |
| Narrative / Storytelling | Immersive, story-voice |
| Tense / Suspenseful | Edge-of-seat, urgent |
Select a tone, click Preview Rewrite, review the result, then Accept or Discard.
Translates the text into another language using the local AI model.

Supported Languages: Japanese, Spanish, French, German, Hindi, Italian, Portuguese, Korean, Russian, Arabic, Mandarin Chinese, and more.
Translation Tones: Natural, Formal, Casual, Professional
Fish-Speech tags like (laugh) and [whisper] are preserved through translation unchanged.
Important: Translation quality depends on which LLM model you have selected. Qwen 0.5B is the default and works well for short passages. For long chapters or higher accuracy, a larger model like Gemma 3 1B or 4B will produce noticeably better results. See LLM Models.
See also: Translation Guide for how translation interacts with Kokoro vs Fish-Speech engines.
| Button | Action |
|---|---|
| Save | Commits edits to the playlist item. The original file on disk is untouched. |
| Cancel | Discards all changes and closes the editor. |
A live character count is shown as you type.
- The AI tools work on whatever text is currently in the editor — you can manually edit first, then run enhancement on your edited version.
- You can run multiple tools in sequence: Translate → Enhance → Tone, for example.
- For audiobooks with many chapters, consider using Assisted Flow in the playlist instead — it applies enhancement automatically to each item before playback without opening the editor.
Getting Started
Labs
Features
Reference