Skip to content
Merged
Show file tree
Hide file tree
Changes from 19 commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
71785d3
Add MLX DeepFilterNet CLI and parity fixes
kylehowells Mar 4, 2026
1e5b953
Add noise-reduction overlay and abs-diff comparison plots
kylehowells Mar 4, 2026
eda6dd1
Match libDF ISTFT normalization for closer PyTorch parity
kylehowells Mar 4, 2026
a129feb
DeepFilterNet MLX parity fixes and benchmark tooling
kylehowells Mar 5, 2026
73f4705
Speed up MLX DeepFilterNet inference hot paths
kylehowells Mar 5, 2026
7c897c4
Add native MLX support for DeepFilterNet v1/v2/v3
kylehowells Mar 5, 2026
6780c88
refactor DeepFilterNet into STS and add true streaming support
kylehowells Mar 5, 2026
2ec2cff
deepfilternet: support model-path loading with config-driven version
kylehowells Mar 5, 2026
0a19b28
deepfilternet: speed up mlx path and simplify model loading
kylehowells Mar 9, 2026
fe2556e
deepfilternet: consolidate model selection to --model
kylehowells Mar 9, 2026
1c1e793
style: run black and isort across deepfilternet changes
kylehowells Mar 9, 2026
ec3a079
docs: add DeepFilterNet attribution to contributions
kylehowells Mar 9, 2026
8a0582a
docs: update DeepFilterNet model link in README
kylehowells Mar 10, 2026
97c3966
deepfilternet: cleanup imports, add docstrings, move converter
kylehowells Mar 10, 2026
9fd175f
deepfilternet: add sample noisy audio for testing
kylehowells Mar 10, 2026
80f54ec
deepfilternet: move dev scripts into model package and update benchmark
kylehowells Mar 10, 2026
3356d5b
style: format benchmark.py with black
kylehowells Mar 10, 2026
85b6d09
deepfilternet: add end-to-end integration tests
kylehowells Mar 10, 2026
0585e08
deepfilternet: simplify from_pretrained to use HuggingFace-first loading
kylehowells Mar 10, 2026
af8fa82
Address PR review: remove dev scripts, add generic STS generate CLI, …
kylehowells Mar 11, 2026
9c2860e
Add PyTorch reference target audio and parity test
kylehowells Mar 11, 2026
c950c22
sts generate: simplify CLI with lazy imports and remove dead code
kylehowells Mar 11, 2026
6db73c5
style: run pre-commit formatter
kylehowells Mar 11, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CONTRIBUTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,10 @@ This file acknowledges the original authors and contributors of models ported to
- **Copyright**: Speech Lab, Alibaba Group
- **License**: Apache License 2.0
- **MLX Port**: Dmitry Starkov ([@starkdmi](https://github.com/starkdmi))

## DeepFilterNet (Speech Enhancement)

- **Original**: [Rikorose/DeepFilterNet](https://github.com/Rikorose/DeepFilterNet)
- **Copyright**: Hendrik Schröter and contributors
- **License**: MIT / Apache-2.0
- **MLX Port**: Kyle Howells ([@kylehowells](https://github.com/kylehowells))
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,8 +119,9 @@ See the [Sortformer README](mlx_audio/vad/models/sortformer/README.md) for API d
| Model | Description | Use Case | Repo |
|-------|-------------|----------|------|
| **SAM-Audio** | Text-guided source separation | Extract specific sounds | [mlx-community/sam-audio-large](https://huggingface.co/mlx-community/sam-audio-large) |
| **Liquid2.5-Audio*** | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | [mlx-community/LFM2.5-Audio-1.5B-8bit](https://huggingface.co/mlx-community/LFM2.5-Audio-1.5B-8bit)
| **Liquid2.5-Audio*** | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | [mlx-community/LFM2.5-Audio-1.5B-8bit](https://huggingface.co/mlx-community/LFM2.5-Audio-1.5B-8bit) |
| **MossFormer2 SE** | Speech enhancement | Noise removal | [starkdmi/MossFormer2_SE_48K_MLX](https://huggingface.co/starkdmi/MossFormer2_SE_48K_MLX) |
| **DeepFilterNet (1/2/3)** | Speech enhancement | Noise suppression | [iky1e/DeepFilterNet3-MLX](https://huggingface.co/iky1e/DeepFilterNet3-MLX) |

## Model Examples

Expand Down
7 changes: 7 additions & 0 deletions examples/deepfilternet.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
"""Example script for DeepFilterNet speech enhancement."""

from mlx_audio.sts.deepfilternet import main

if __name__ == "__main__":
main()
Binary file added examples/denoise/noisey_audio_10s.wav
Binary file not shown.
21 changes: 20 additions & 1 deletion mlx_audio/sts/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,11 @@
from .models.deepfilternet import (
DeepFilterNet2Config,
DeepFilterNet3Config,
DeepFilterNetConfig,
DeepFilterNetModel,
DeepFilterNetStreamer,
DeepFilterNetStreamingConfig,
)
from .models.mossformer2_se import (
MossFormer2SE,
MossFormer2SEConfig,
Expand All @@ -11,7 +19,11 @@
SeparationResult,
save_audio,
)
from .voice_pipeline import VoicePipeline

try:
from .voice_pipeline import VoicePipeline
except ImportError:
VoicePipeline = None

__all__ = [
"SAMAudio",
Expand All @@ -21,6 +33,13 @@
"save_audio",
"SAMAudioConfig",
"VoicePipeline",
# DeepFilterNet
"DeepFilterNetModel",
"DeepFilterNetConfig",
"DeepFilterNet2Config",
"DeepFilterNet3Config",
"DeepFilterNetStreamer",
"DeepFilterNetStreamingConfig",
# MossFormer2 SE
"MossFormer2SE",
"MossFormer2SEConfig",
Expand Down
90 changes: 90 additions & 0 deletions mlx_audio/sts/deepfilternet.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
"""CLI entrypoint for DeepFilterNet speech enhancement."""

from __future__ import annotations

import argparse
from pathlib import Path

from mlx_audio.sts.models.deepfilternet import DeepFilterNetModel


def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Denoise audio with DeepFilterNet (MLX)"
)
parser.add_argument("input", help="Input noisy audio file (48 kHz)")
parser.add_argument("-o", "--output", help="Output denoised audio path")
parser.add_argument(
"-m",
"--model",
type=str,
default=None,
help=(
"Model name or path. Can be a Hugging Face repo id or local model directory."
),
)
parser.add_argument(
"--stream",
action="store_true",
help="Run chunked streaming enhancement (stateful process_chunk/flush path)",
)
parser.add_argument(
"--chunk-ms",
type=float,
default=480.0,
help="Chunk size for --stream mode in milliseconds (default: 480)",
)
parser.add_argument(
"--pad-end-frames",
type=int,
default=3,
help="Extra zero frames to flush streaming tail (default: 3)",
)
parser.add_argument(
"--no-delay-compensation",
action="store_true",
help="Disable initial delay compensation crop in streaming output",
)
return parser


def main() -> None:
args = build_parser().parse_args()

in_path = Path(args.input).expanduser().resolve()
if not in_path.exists():
raise FileNotFoundError(f"Input file not found: {in_path}")

out_path = (
Path(args.output).expanduser().resolve()
if args.output
else in_path.with_stem(in_path.stem + "_enhanced_mlx")
)

model = DeepFilterNetModel.from_pretrained(model_name_or_path=args.model)
if args.stream:
sr = model.config.sample_rate
chunk_samples = max(model.config.hop_size, int(sr * args.chunk_ms / 1000.0))
try:
model.enhance_file_streaming(
in_path,
out_path,
chunk_samples=chunk_samples,
pad_end_frames=max(0, args.pad_end_frames),
compensate_delay=(not args.no_delay_compensation),
)
except NotImplementedError as exc:
raise NotImplementedError(
f"Streaming mode is unavailable for {model.model_version}: {exc}"
) from exc
else:
model.enhance_file(in_path, out_path)

print(f"Input : {in_path}")
print(f"Model : {model.model_version}")
print(f"Mode : {'streaming' if args.stream else 'offline'}")
print(f"Output: {out_path}")


if __name__ == "__main__":
main()
15 changes: 15 additions & 0 deletions mlx_audio/sts/models/__init__.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,13 @@
# Copyright (c) 2025 Prince Canuma and contributors (https://github.com/Blaizzy/mlx-audio)

from .deepfilternet import (
DeepFilterNet2Config,
DeepFilterNet3Config,
DeepFilterNetConfig,
DeepFilterNetModel,
DeepFilterNetStreamer,
DeepFilterNetStreamingConfig,
)
from .lfm_audio import (
ChatState,
GenerationConfig,
Expand All @@ -25,6 +33,13 @@
"Batch",
"save_audio",
"SAMAudioConfig",
# DeepFilterNet
"DeepFilterNetModel",
"DeepFilterNetConfig",
"DeepFilterNet2Config",
"DeepFilterNet3Config",
"DeepFilterNetStreamer",
"DeepFilterNetStreamingConfig",
# MossFormer2 SE
"MossFormer2SE",
"MossFormer2SEConfig",
Expand Down
46 changes: 46 additions & 0 deletions mlx_audio/sts/models/deepfilternet/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# DeepFilterNet (MLX)

DeepFilterNet speech enhancement in pure MLX with support for model versions 1, 2, and 3.

## Quick Start

```python
from mlx_audio.sts.models.deepfilternet import DeepFilterNetModel

model = DeepFilterNetModel.from_pretrained()
model.enhance_file("noisy.wav", "clean.wav")
```

Or load from a local model directory (must contain `config.json` and weights):

```python
model = DeepFilterNetModel.from_pretrained("./models/MyDeepFilterNet")
```

Or load from a Hugging Face repo id:

```python
model = DeepFilterNetModel.from_pretrained("iky1e/DeepFilterNet3-MLX")
```

Streaming/chunked mode (true per-hop stateful processing for DF2/DF3):

```python
streamer = model.create_streamer(pad_end_frames=3, compensate_delay=True)
out_1 = streamer.process_chunk(chunk_a)
out_2 = streamer.process_chunk(chunk_b)
out_tail = streamer.flush()
```

## Model Selection

Model architecture is selected from `config.json` (`model_version`).

## Example Script

```bash
python examples/deepfilternet.py examples/denoise/noisey_audio_10s.wav
python examples/deepfilternet.py examples/denoise/noisey_audio_10s.wav --model ./models/DeepFilterNet3
python examples/deepfilternet.py examples/denoise/noisey_audio_10s.wav --model iky1e/DeepFilterNet3-MLX
python examples/deepfilternet.py examples/denoise/noisey_audio_10s.wav --stream
```
19 changes: 19 additions & 0 deletions mlx_audio/sts/models/deepfilternet/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
"""DeepFilterNet speech enhancement model for MLX."""

from .config import DeepFilterNet2Config, DeepFilterNet3Config, DeepFilterNetConfig
from .model import DeepFilterNetModel
from .streaming import DeepFilterNetStreamer, DeepFilterNetStreamingConfig

Model = DeepFilterNetModel
ModelConfig = DeepFilterNetConfig

__all__ = [
"DeepFilterNetModel",
"DeepFilterNetConfig",
"DeepFilterNet2Config",
"DeepFilterNet3Config",
"DeepFilterNetStreamer",
"DeepFilterNetStreamingConfig",
"Model",
"ModelConfig",
]
Loading