Skip to content

Commit a907427

Browse files
Merge branch 'main' into fix/whisper-empty-non-speech-tokens
2 parents 2880fa7 + aef6ebc commit a907427

75 files changed

Lines changed: 6664 additions & 322 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/tests.yml‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -43,6 +43,11 @@ jobs:
4343
- name: Install all deps
4444
run: pip install -e ".[all,dev]"
4545

46+
- name: Install ffmpeg for server audio tests
47+
run: |
48+
brew install ffmpeg
49+
ffmpeg -version
50+
4651
- name: Test DSP imports without TTS/STT
4752
run: |
4853
python -c "from mlx_audio.dsp import stft, mel_filters; print('DSP OK')"

‎CONTRIBUTIONS.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -29,3 +29,10 @@ This file acknowledges the original authors and contributors of models ported to
2929
- **Copyright**: NVIDIA Corporation
3030
- **License**: [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/)
3131
- **MLX Port**: [@ARahim3](https://github.com/ARahim3)
32+
33+
## rumik-oss 1 (Text-to-Speech)
34+
35+
- **Original**: [rumik-ai/rumik-oss-1](https://huggingface.co/rumik-ai/rumik-oss-1)
36+
- **Copyright**: rumik.ai
37+
- **License**: [CC-BY-NC 4.0 with Cohere Labs acceptable-use addendum](https://cohere.com/cohere-labs-cc-by-nc-license) (weights); Mimi codec CC-BY-4.0
38+
- **MLX Port**: Suryansh Shakya ([@nullHawk](https://github.com/nullHawk))

‎README.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -132,6 +132,7 @@ for result in model.generate(
132132
| **Ming Omni TTS (Dense)** | Lightweight dense Ming Omni variant for voice cloning and style control | EN, ZH | [mlx-community/Ming-omni-tts-0.5B-bf16](https://huggingface.co/mlx-community/Ming-omni-tts-0.5B-bf16) |
133133
| **KugelAudio** | SOTA 7B AR+Diffusion TTS for European languages | EN, DE, FR, ES, IT, PT, NL, PL, RU, UK, + 14 more | [kugelaudio/kugelaudio-0-open](https://huggingface.co/kugelaudio/kugelaudio-0-open) |
134134
| **Voxtral TTS** | Mistral's 4B multilingual TTS (20 voices, 9 languages) | EN, FR, ES, DE, IT, PT, NL, AR, HI | [mlx-community/Voxtral-4B-TTS-2603-mlx-bf16](https://huggingface.co/mlx-community/Voxtral-4B-TTS-2603-mlx-bf16) |
135+
| **rumik-oss 1** | 3B expressive multilingual Indic TTS with 22-language support, description-conditioned delivery and inline vocalizations | 22 Indic languages + EN | [rumik-ai/rumik-oss-1](https://huggingface.co/rumik-ai/rumik-oss-1), [8bit](https://huggingface.co/rumik-ai/rumik-oss-1-mlx-8bit), [4bit](https://huggingface.co/rumik-ai/rumik-oss-1-mlx-4bit) |
135136
| **VoxCPM2** | 2B tokenizer-free TTS with 48kHz output, voice design, voice cloning, and continuation | 30 languages | [bf16](https://huggingface.co/mlx-community/VoxCPM2-bf16), [8bit](https://huggingface.co/mlx-community/VoxCPM2-8bit), [4bit](https://huggingface.co/mlx-community/VoxCPM2-4bit) |
136137
| **LongCat-AudioDiT** | SOTA diffusion TTS in waveform latent space with voice cloning | ZH, EN | [mlx-community/LongCat-AudioDiT-1B-bf16](https://huggingface.co/mlx-community/LongCat-AudioDiT-1B-bf16) |
137138
| **MeloTTS** | Lightweight VITS2-based TTS with streaming | EN (more coming) | [mlx-community/MeloTTS-English-MLX](https://huggingface.co/mlx-community/MeloTTS-English-MLX) |
@@ -178,7 +179,9 @@ See the model READMEs for API details, streaming examples, and conversion steps.
178179
| Model | Description | Use Case | Repo |
179180
|-------|-------------|----------|------|
180181
| **SAM-Audio** | Text-guided source separation | Extract specific sounds | [mlx-community/sam-audio-large](https://huggingface.co/mlx-community/sam-audio-large) |
182+
| [**DialogueSidon**](mlx_audio/sts/models/dialogue_sidon/README.md) | Two-speaker separation and restoration | Separate dialogue into speaker tracks | [mlx-community/DialogueSidon](https://huggingface.co/mlx-community/DialogueSidon) (FP32), [mlx-community/DialogueSidon-bf16](https://huggingface.co/mlx-community/DialogueSidon-bf16) (BF16) |
181183
| **Liquid2.5-Audio*** | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | [mlx-community/LFM2.5-Audio-1.5B-8bit](https://huggingface.co/mlx-community/LFM2.5-Audio-1.5B-8bit) |
184+
| **MiMo-Audio** | English/Chinese TTS, ASR, audio understanding and dialogue; Base few-shot speech tasks | Speech interactions and audio completion | [Instruct](https://huggingface.co/XiaomiMiMo/MiMo-Audio-7B-Instruct), [Base](https://huggingface.co/XiaomiMiMo/MiMo-Audio-7B-Base), [audio tokenizer](https://huggingface.co/XiaomiMiMo/MiMo-Audio-Tokenizer), [guide](mlx_audio/sts/models/mimo_audio/README.md) |
182185
| **MossFormer2 SE** | Speech enhancement | Noise removal | [starkdmi/MossFormer2_SE_48K_MLX](https://huggingface.co/starkdmi/MossFormer2_SE_48K_MLX) |
183186
| **DeepFilterNet (1/2/3)** | Speech enhancement | Noise suppression | [mlx-community/DeepFilterNet-mlx](https://huggingface.co/mlx-community/DeepFilterNet-mlx) |
184187
| **NemotronLabs VoiceChat** | Full-duplex speech-to-speech with streaming transcription and function calling | Real-time voice conversation | [mlx-community/NemotronLabs-VoiceChat-11B-4bit](https://huggingface.co/mlx-community/NemotronLabs-VoiceChat-11B-4bit) |

‎docs/guides/nemotron-live-input.md‎

Lines changed: 48 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,48 @@
1+
# Nemotron live input
2+
3+
Nemotron implements the shared [`StreamingSession`](streaming-stt.md) protocol,
4+
also used by Voxtral Realtime. Sessions use greedy
5+
decoding (`temperature=0`), mono float32 PCM at `session.input_sample_rate`, and
6+
independent frontend, encoder and RNNT state.
7+
8+
```python
9+
session = model.create_streaming_session()
10+
session.feed(samples)
11+
print("".join(session.step(max_decode_tokens=8)), end="")
12+
session.close()
13+
while not session.done:
14+
print("".join(session.step(max_decode_tokens=8)), end="")
15+
```
16+
17+
Call `step` regularly from one decoder thread. Its budget counts joint evaluations,
18+
including blanks and special tokens. Each step ingests at most one native audio
19+
chunk and drains existing encoded frames before ingesting more. Packet boundaries
20+
do not reset frontend or decoder state. Text is emitted as deltas, not cumulative
21+
transcripts. No transcript history is retained by the session.
22+
23+
`feed` copies input and rejects non-finite or non-mono data. Pending audio is
24+
limited to 30 seconds; exceeding that limit raises `BufferError` without dropping
25+
previously queued audio. This is a pending-input limit, not a session duration
26+
limit. Producers must pace input and consumers must keep stepping. `close` is
27+
idempotent and signals input completion, not immediate decoder completion.
28+
`cancel` discards pending work without emitting a successful final; `reset`
29+
starts independent state. Both must run with no concurrent decoder step.
30+
31+
The generic server selects models through `/v1/realtime?model=<model-id>` and
32+
accepts session updates and PCM16 audio append/commit events. Use model-rate PCM
33+
to avoid conflating session behavior with packet-wise server resampling. This
34+
change does not add a Nativ route, authentication policy or new server protocol.
35+
36+
Validation covers tiny-model frontend/encoder behavior, real tiny-decoder parity,
37+
and the actual ASGI WebSocket handler (partial before commit, completion and a
38+
second turn). A local multilingual checkpoint check used five paced runs each of
39+
Portuguese and English synthesized speech, with 20 ms / 16 kHz chunks. Text matched
40+
the existing streaming decoder after trimming outer whitespace. Both paths made
41+
the same Portuguese recognition error; this is parity evidence, not an accuracy
42+
benchmark.
43+
44+
For those ten local session runs, first-partial p50/p95 was 1.912/3.433 seconds and
45+
close-to-final p50/p95 was 0.267/0.676 seconds (maximum 0.921 seconds). Quantiles use
46+
linear interpolation. These small-sample measurements are not WebSocket latency
47+
or Nativ/OpenClaw end-to-end proof. Long-session memory-residency measurements,
48+
broader speech samples and production integration remain outstanding.

‎docs/guides/quantization.md‎

Lines changed: 23 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -78,6 +78,29 @@ python -m mlx_audio.convert \
7878
--upload-repo username/Kokoro-82M-4bit
7979
```
8080

81+
## MiMo-Audio
82+
83+
Base and Instruct can load the original checkpoints directly. To reduce memory, convert either checkpoint with the standard converter:
84+
85+
```bash
86+
python -m mlx_audio.convert \
87+
--hf-path XiaomiMiMo/MiMo-Audio-7B-Instruct \
88+
--mlx-path ./MiMo-Audio-7B-Instruct-4bit \
89+
-q --q-bits 4 --q-group-size 64
90+
```
91+
92+
MiMo's quantization policy covers the global Qwen2 backbone and text head. Acoustic patch transformers, speech embeddings and speech heads retain their floating-point precision. The separate audio tokenizer is not quantized. Consequently, total memory use is larger than a uniformly quantized 7B text model. Both source and converted weights use the same [offline APIs](../models/sts/mimo-audio.md).
93+
94+
To save the standalone tokenizer in MLX convolution layouts:
95+
96+
```bash
97+
python -m mlx_audio.codec.models.mimo_audio_tokenizer.convert \
98+
--hf-path XiaomiMiMo/MiMo-Audio-Tokenizer \
99+
--mlx-path ./MiMo-Audio-Tokenizer-mlx
100+
```
101+
102+
Its converter preserves stored precision by default and accepts `--dtype` and `--revision`. The resulting directory contains weights, config, and a model card. The language-model converter records the separate codec dependency and its pinned revision in `config.json`; use `audio_tokenizer_path` or the CLI's `--audio-tokenizer-path` to choose a local codec. No upload is required for local inference.
103+
81104
## Conversion Options Reference
82105

83106
| Flag | Description |

‎docs/guides/streaming-stt.md‎

Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,49 @@
1+
# Adding a live-input STT model
2+
3+
`mlx_audio.stt.streaming.StreamingSession` is the shared structural protocol
4+
used by the realtime server. Nemotron and Voxtral Realtime both implement it.
5+
No base class, registration step, or model-name branch in the server is needed.
6+
7+
Implement `create_streaming_session(...) -> StreamingSession` on the model.
8+
Keep model-specific construction and options there; the session presents:
9+
10+
- `input_sample_rate: int`: native input rate.
11+
- `feed(samples: np.ndarray) -> None`: queue mono float32 PCM; no inference.
12+
- `step(*, max_decode_tokens: int = 4) -> list[str]`: advance on one model
13+
executor and return append-only text deltas, possibly an empty list.
14+
- `close() -> None`: signal end of input, without discarding the pending tail.
15+
- `done: bool`: true once decoding has finished.
16+
17+
`feed` and `close` may run on a producer thread while one consumer calls `step`.
18+
Do not modify submitted arrays while queued, or feed after close/completion.
19+
Keep calling `step` after close until done. An empty result is not completion.
20+
Models may also finish early (for example EOS or a model-specific token cap).
21+
The decode budget is not a time limit: model-specific encoder work can still
22+
make a step expensive. Document input limits and errors in the model guide.
23+
Reset/cancel are optional implementation extensions, not part of this contract.
24+
25+
```python
26+
from mlx_audio.stt.streaming import StreamingSession
27+
28+
session: StreamingSession = model.create_streaming_session(temperature=0.0)
29+
for samples in audio_chunks:
30+
if session.done:
31+
break
32+
session.feed(samples)
33+
for delta in session.step():
34+
emit(delta)
35+
session.close()
36+
while not session.done:
37+
for delta in session.step():
38+
emit(delta)
39+
```
40+
41+
This example executes sequentially; a live server schedules the producer and
42+
consumer independently. Factory options remain model-specific; the existing
43+
server forwards transcription delay only when the factory declares it.
44+
45+
Add a small model fixture to `mlx_audio/stt/tests/test_streaming_session.py`
46+
to exercise the same server/session contract, plus model-specific decoder
47+
parity tests. Runtime `isinstance` checks only member presence; behavioral tests
48+
are necessary to prove the contract. The shared test uses tiny random weights
49+
and deterministic token selection, not production checkpoints or downloads.

‎docs/models/index.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -30,6 +30,7 @@ Generate natural-sounding speech from text. Multiple models with multilingual su
3030
| **KittenTTS** | Compact KittenTTS 0.8 models for edge-friendly TTS | EN | [nano](https://huggingface.co/mlx-community/kitten-tts-nano-0.8) / [micro](https://huggingface.co/mlx-community/kitten-tts-micro-0.8) / [mini](https://huggingface.co/mlx-community/kitten-tts-mini-0.8) |
3131
| **Qwen3-TTS** | Alibaba's multilingual TTS with voice design | ZH, EN, JA, KO, + more | [mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16](https://huggingface.co/mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16) |
3232
| **Voxtral TTS** | Mistral's 4B multilingual TTS (20 voices, 9 languages) | EN, FR, ES, DE, IT, PT, NL, AR, HI | [mlx-community/Voxtral-4B-TTS-2603-mlx-bf16](https://huggingface.co/mlx-community/Voxtral-4B-TTS-2603-mlx-bf16) |
33+
| **rumik-oss 1** | 3B expressive multilingual Indic TTS with 22-language support, description-conditioned delivery and inline vocalizations | 22 Indic languages + EN | [rumik-ai/rumik-oss-1](https://huggingface.co/rumik-ai/rumik-oss-1), [8bit](https://huggingface.co/rumik-ai/rumik-oss-1-mlx-8bit), [4bit](https://huggingface.co/rumik-ai/rumik-oss-1-mlx-4bit) |
3334
| **VoxCPM2** | 2B tokenizer-free TTS with 48kHz output, voice design, voice cloning, and continuation | 30 languages | [bf16](https://huggingface.co/mlx-community/VoxCPM2-bf16), [8bit](https://huggingface.co/mlx-community/VoxCPM2-8bit), [4bit](https://huggingface.co/mlx-community/VoxCPM2-4bit) |
3435
| **CSM / MisoTTS** | Sesame-style conversational speech models with voice cloning | EN | [mlx-community/csm-1b](https://huggingface.co/mlx-community/csm-1b), [MisoTTS bf16](https://huggingface.co/mlx-community/MisoLabs-MisoTTS-bf16), [MisoTTS 8bit](https://huggingface.co/mlx-community/MisoLabs-MisoTTS-8bit) |
3536
| **Dia** | Dialogue-focused TTS | EN | [mlx-community/Dia-1.6B-fp16](https://huggingface.co/mlx-community/Dia-1.6B-fp16) |

‎docs/models/sts/dialogue-sidon.md‎

Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
# DialogueSidon
2+
3+
Separate a recording of two speakers into individual audio tracks on Apple
4+
Silicon. DialogueSidon also restores degraded speech and outputs two mono WAV
5+
files at **24 kHz**, with the same duration as the input.
6+
7+
Follow the [MLX Audio installation instructions](../../getting-started/installation.md)
8+
before running the examples below.
9+
10+
## Supported models
11+
12+
| Precision | Repository ID | Weight size |
13+
|-----------|---------------|-------------|
14+
| FP32 | [mlx-community/DialogueSidon](https://huggingface.co/mlx-community/DialogueSidon) | 1.78 GB |
15+
| BF16 | [mlx-community/DialogueSidon-bf16](https://huggingface.co/mlx-community/DialogueSidon-bf16) | 889 MB |
16+
17+
Either repository ID works in the examples below. Choose BF16 for a smaller
18+
download and lower weight memory use. You can also load a local model directory.
19+
20+
## Command line
21+
22+
```bash
23+
python -m mlx_audio.sts.generate \
24+
--model mlx-community/DialogueSidon \
25+
--audio dialogue.wav --output-path separated.wav \
26+
--num-steps 30 --seed 0
27+
```
28+
29+
This writes `separated_speaker_1.wav` and `separated_speaker_2.wav`.
30+
31+
## Python
32+
33+
```python
34+
from mlx_audio.sts import load
35+
from mlx_audio.audio_io import write
36+
37+
model = load("mlx-community/DialogueSidon")
38+
result = model.separate("dialogue.wav", num_steps=30, seed=0)
39+
for i, speaker in enumerate(result.speakers, 1):
40+
write(f"speaker_{i}.wav", speaker, result.sample_rate)
41+
```
42+
43+
`result.speakers` contains two tracks with shape `[2, samples]`, and
44+
`result.sample_rate` is `24000`.
45+
46+
For a NumPy or MLX array, provide its sample rate:
47+
48+
```python
49+
from mlx_audio.audio_io import read
50+
51+
audio, sample_rate = read("dialogue.wav")
52+
result = model.separate(audio, sample_rate=sample_rate, seed=0)
53+
```
54+
55+
Arrays must use `[samples]` or `[samples, channels]` layout. Stereo and other
56+
multichannel inputs are averaged to mono before separation.
57+
58+
## Options and long recordings
59+
60+
| Python option | CLI option | Default | Purpose |
61+
|---------------|------------|---------|---------|
62+
| `num_steps` | `--num-steps` | `30` | Fewer steps run faster but can change output quality. |
63+
| `seed` | `--seed` | Random | Set a seed to repeat a run with the same input, model, and settings. |
64+
| `chunk_seconds` | `--chunk-seconds` | `20.0` | Audio processed at once; smaller chunks reduce memory use. |
65+
| `overlap_seconds` | `--overlap-seconds` | `5.0` | Overlap between chunks, in seconds. Must be less than the chunk duration. |
66+
67+
Long recordings are processed automatically in overlapping chunks. To process
68+
an entire file at once in Python, use `chunk_seconds=None`; this uses more
69+
memory.
70+
71+
The model supports exactly two speakers and processes recorded audio offline.
72+
The tracks are numbered rather than labeled with speaker identities; a speaker
73+
may switch tracks after a long silence. Because the model also restores speech,
74+
adding the tracks together may not reproduce the input exactly.
75+
76+
## License and attribution
77+
78+
Model weights use **CC-BY-NC-4.0**. Original model by Wataru Nakata, Yuki Saito,
79+
Kazuki Yamauchi, Emiru Tsunoo, and Hiroshi Saruwatari (SaruLab).
80+
See the [original model card](https://huggingface.co/sarulab-speech/DialogueSidon).
81+
The upstream [Sidon code](https://github.com/sarulab-speech/Sidon) is MIT licensed.

0 commit comments

Comments
 (0)