|
| 1 | +# DialogueSidon |
| 2 | + |
| 3 | +Separate a recording of two speakers into individual audio tracks on Apple |
| 4 | +Silicon. DialogueSidon also restores degraded speech and outputs two mono WAV |
| 5 | +files at **24 kHz**, with the same duration as the input. |
| 6 | + |
| 7 | +Follow the [MLX Audio installation instructions](../../getting-started/installation.md) |
| 8 | +before running the examples below. |
| 9 | + |
| 10 | +## Supported models |
| 11 | + |
| 12 | +| Precision | Repository ID | Weight size | |
| 13 | +|-----------|---------------|-------------| |
| 14 | +| FP32 | [mlx-community/DialogueSidon](https://huggingface.co/mlx-community/DialogueSidon) | 1.78 GB | |
| 15 | +| BF16 | [mlx-community/DialogueSidon-bf16](https://huggingface.co/mlx-community/DialogueSidon-bf16) | 889 MB | |
| 16 | + |
| 17 | +Either repository ID works in the examples below. Choose BF16 for a smaller |
| 18 | +download and lower weight memory use. You can also load a local model directory. |
| 19 | + |
| 20 | +## Command line |
| 21 | + |
| 22 | +```bash |
| 23 | +python -m mlx_audio.sts.generate \ |
| 24 | + --model mlx-community/DialogueSidon \ |
| 25 | + --audio dialogue.wav --output-path separated.wav \ |
| 26 | + --num-steps 30 --seed 0 |
| 27 | +``` |
| 28 | + |
| 29 | +This writes `separated_speaker_1.wav` and `separated_speaker_2.wav`. |
| 30 | + |
| 31 | +## Python |
| 32 | + |
| 33 | +```python |
| 34 | +from mlx_audio.sts import load |
| 35 | +from mlx_audio.audio_io import write |
| 36 | + |
| 37 | +model = load("mlx-community/DialogueSidon") |
| 38 | +result = model.separate("dialogue.wav", num_steps=30, seed=0) |
| 39 | +for i, speaker in enumerate(result.speakers, 1): |
| 40 | + write(f"speaker_{i}.wav", speaker, result.sample_rate) |
| 41 | +``` |
| 42 | + |
| 43 | +`result.speakers` contains two tracks with shape `[2, samples]`, and |
| 44 | +`result.sample_rate` is `24000`. |
| 45 | + |
| 46 | +For a NumPy or MLX array, provide its sample rate: |
| 47 | + |
| 48 | +```python |
| 49 | +from mlx_audio.audio_io import read |
| 50 | + |
| 51 | +audio, sample_rate = read("dialogue.wav") |
| 52 | +result = model.separate(audio, sample_rate=sample_rate, seed=0) |
| 53 | +``` |
| 54 | + |
| 55 | +Arrays must use `[samples]` or `[samples, channels]` layout. Stereo and other |
| 56 | +multichannel inputs are averaged to mono before separation. |
| 57 | + |
| 58 | +## Options and long recordings |
| 59 | + |
| 60 | +| Python option | CLI option | Default | Purpose | |
| 61 | +|---------------|------------|---------|---------| |
| 62 | +| `num_steps` | `--num-steps` | `30` | Fewer steps run faster but can change output quality. | |
| 63 | +| `seed` | `--seed` | Random | Set a seed to repeat a run with the same input, model, and settings. | |
| 64 | +| `chunk_seconds` | `--chunk-seconds` | `20.0` | Audio processed at once; smaller chunks reduce memory use. | |
| 65 | +| `overlap_seconds` | `--overlap-seconds` | `5.0` | Overlap between chunks, in seconds. Must be less than the chunk duration. | |
| 66 | + |
| 67 | +Long recordings are processed automatically in overlapping chunks. To process |
| 68 | +an entire file at once in Python, use `chunk_seconds=None`; this uses more |
| 69 | +memory. |
| 70 | + |
| 71 | +The model supports exactly two speakers and processes recorded audio offline. |
| 72 | +The tracks are numbered rather than labeled with speaker identities; a speaker |
| 73 | +may switch tracks after a long silence. Because the model also restores speech, |
| 74 | +adding the tracks together may not reproduce the input exactly. |
| 75 | + |
| 76 | +## License and attribution |
| 77 | + |
| 78 | +Model weights use **CC-BY-NC-4.0**. Original model by Wataru Nakata, Yuki Saito, |
| 79 | +Kazuki Yamauchi, Emiru Tsunoo, and Hiroshi Saruwatari (SaruLab). |
| 80 | +See the [original model card](https://huggingface.co/sarulab-speech/DialogueSidon). |
| 81 | +The upstream [Sidon code](https://github.com/sarulab-speech/Sidon) is MIT licensed. |
0 commit comments