Skip to content

Add DeepFilterNet speech enhancement model - #100

Merged
lucasnewman merged 36 commits into
Blaizzy:mainfrom
kylehowells:main
Apr 22, 2026
Merged

lucasnewman merged 36 commits into
Blaizzy:mainfrom
kylehowells:main

Conversation

@kylehowells

@kylehowells kylehowells commented Mar 16, 2026 •

Copy link
Copy Markdown
Contributor

Swift version of the python mlx-audio change Blaizzy/mlx-audio#561


Summary

  • Adds DeepFilterNet (V1/V2/V3) real-time speech enhancement to mlx-audio-swift
  • Default model: mlx-community/DeepFilterNet-mlx (same as the Python mlx-audio repo)
  • Supports offline batch enhancement and low-latency streaming
  • GRU layers optimized with Accelerate (vDSP_mmul) for CPU inference, avoiding Metal kernel dispatch overhead in the sequential hidden-state loop
  • Includes CLI integration (--mode short / --mode stream), CI coverage, and comprehensive test suite
  • Doc comments on all public API surfaces

Files added

File Purpose
DeepFilterNetModel.swift Model class, loading, public API
DeepFilterNetStreamer.swift Stateful hop-by-hop streaming pipeline
DeepFilterNetForward.swift Network forward pass (V1/V2/V3)
DeepFilterNetLayers.swift Conv2d, ConvTranspose, BatchNorm, GRU, Linear
DeepFilterNetDSP.swift Feature normalization, ERB filterbank, weight caches
DeepFilterNetConfig.swift Model configuration from config.json
README.md Model docs with usage examples, architecture, performance

Performance (Apple M1 Max, 48kHz mono)

Swift vs Python (offline, DeepFilterNet V3)

Audio Python (deepFilter CLI) Swift Speedup
10s 0.284s 0.19s 1.5x faster
52s 1.22s 0.55s 2.2x faster

Offline

Audio Time Real-time Factor
10s 0.23s ~43x RT
52s 0.55s ~95x RT

Streaming

Metric Value
Per-hop latency (true GPU) ~4.8ms
Real-time budget 10ms (480 samples @ 48kHz)
Headroom ~52%
Peak memory ~80MB RSS
Algorithmic latency ~30ms (2-frame lookahead + processing)

Per-hop stage breakdown (streaming, stage eval)

Stage Per-hop %
STFT analysis 0.31ms 8%
Feature extraction 0.36ms 10%
Encoder + embedding GRU 1.11ms 30%
ERB decoder 0.91ms 24%
DF decoder 0.68ms 18%
ISTFT synthesis 0.30ms 8%

Loading

Supports the mlx-community/DeepFilterNet-mlx repo layout (v1/v2/v3 subfolders) and flat model directories:

// Default: V3 from mlx-community
let model = try await DeepFilterNetModel.fromPretrained()

// Specific version
let model = try await DeepFilterNetModel.fromPretrained(subfolder: "v2")

// Local path
let model = try await DeepFilterNetModel.fromPretrained("/path/to/model")

Test plan

  • Offline enhancement produces correct output (golden-audio spectral comparison test, gated by MLXAUDIO_DFN_MODEL_DIR)
  • Streaming output matches offline output (correlation > 0.999)
  • Config parsing, streaming config defaults, and error handling unit tests
  • STS.loadModel routing test
  • CI job added for DeepFilterNet tests
  • Bit-identical output verified before/after file-split refactoring

🤖 Generated with Claude Code Opus 4.6

@kylehowells
kylehowells marked this pull request as draft March 16, 2026 13:07
@kylehowells
kylehowells marked this pull request as ready for review March 16, 2026 16:59
@kylehowells

Copy link
Copy Markdown
Contributor Author

This is the swift version of the noise reduction added to the python repo.

I actually have an ever faster version as it's own repo https://github.com/kylehowells/DeepFilterNet-mlx but that uses some tricks which I didn't think would really fit in with this repo. Namely custom metal kernels, instead of just using swift-mlx building blocks, and also moving some of the small steps back to the CPU and doing them with Accelerate instead (weirdly found some of the small steps the overhead of sending a job to the. GPU slower than just having the CPU do the work there and then).

This version is the most optimised I could get it just using swift MLX kernels and techniques.

Comment thread Sources/MLXAudioCore/AudioUtils.swift
Comment thread Sources/MLXAudioCore/ModelUtils.swift
Comment thread .github/workflows/tests.yaml Outdated
- name: Run ${{ matrix.name }} tests
run: xcodebuild test-without-building -scheme MLXAudio-Package -destination 'platform=macOS' -only-testing:${{ matrix.suite }} CODE_SIGNING_ALLOWED=NO

sts-deepfilternet:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we just add the STS tests to the existing suite? Spinning up a new runner isn't ideal.

diff --git a/.github/workflows/tests.yaml b/.github/workflows/tests.yaml
index 13a4a3a..21827e7 100644
--- a/.github/workflows/tests.yaml
+++ b/.github/workflows/tests.yaml
@@ -21,6 +21,8 @@ jobs:
             suite: MLXAudioTests/MLXAudioTTSTests
           - name: STT
             suite: MLXAudioTests/MLXAudioSTTTests
+          - name: STS
+            suite: MLXAudioTests/MLXAudioSTSTests
           - name: VAD
             suite: MLXAudioTests/MLXAudioVADTests

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, that should be fixed now.

@lucasnewman

Copy link
Copy Markdown
Collaborator

This is the swift version of the noise reduction added to the python repo.

I actually have an ever faster version as it's own repo https://github.com/kylehowells/DeepFilterNet-mlx but that uses some tricks which I didn't think would really fit in with this repo. Namely custom metal kernels, instead of just using swift-mlx building blocks, and also moving some of the small steps back to the CPU and doing them with Accelerate instead (weirdly found some of the small steps the overhead of sending a job to the. GPU slower than just having the CPU do the work there and then).

This version is the most optimised I could get it just using swift MLX kernels and techniques.

@kylehowells Thanks! The model looks good overall. I left a couple of questions around changes in the common code.

@beshkenadze

Copy link
Copy Markdown
Contributor

@kylehowells @lucasnewman FYI I built my own version that's working on ANE

kylehowells and others added 22 commits March 23, 2026 23:14
…kups

Cache transposed GRU weights at model init to avoid per-step transpositions.
Extract StreamGRU/StreamGRULayer structs so the streamer holds pre-resolved
weight references instead of dictionary lookups each frame. Pre-cache
additional hot-path weights (enc_df_fc_emb, df_dec_skip, df_dec_out). Add
inference sub-stage profiling (encode/emb/erb/df) and a fast path that
bypasses sample accumulation for exact hop-size chunks.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add an opt-in compiled infer path for Swift DeepFilterNet streaming (DFN_STREAM_COMPILE=1) and plumb streamer config support.

Result: parity was unchanged (identical output / same PyTorch metrics), but performance regressed on 10s/10ms-hop streaming runs.

Measured mean: compile off 3.63s (RTF 0.36) vs compile on 6.14s (RTF 0.61), ~1.69x slower with compile.

Likely causes: compile setup overhead, extra per-hop state packing/unpacking for functionalized state, and insufficient kernel fusion for this graph shape.
Revert the compiled streaming infer experiment and restore the eager streaming fast path as the only mode.

Reason: compile mode was parity-safe but consistently slower in benchmark runs, so keeping it adds complexity with no runtime benefit.
…uning

- Replace per-hop ring-history reconstruction with fixed rolling tensor histories for encoder/DF paths\n- Add rolling DF low-band spec history update in streamer\n- Keep streaming parity to PyTorch streaming reference\n- Set default stream materialization interval to 512 hops (env override still supported)
- Vectorize deep-filter assignment across time (loop over DF order only)\n- Use ERB filterbank matmul in offline features path (same as streaming)\n- Preserve streaming behavior and parity\n- Improves offline runtime while increasing offline parity vs PyTorch reference
- Add V1 config fields (convKEnc, convKDec, gruGroups, groupShuffle, etc.)
- Implement V1 forward pass: encoder, ERB decoder, DF decoder with alpha
- Implement GroupedGRU with CPU-side sequential loop using vDSP_mmul
  for hidden-state projection (avoids ~60K tiny GPU dispatches per run)
- Batch input projections on GPU, run recurrent gates on CPU via Accelerate
- Add V1 convkxf with auto-detection of normal vs transposed conv
- Add V1 GroupedLinear with packed einsum path
- Add exact sequential EMA normalization (bandMeanNormExact, bandUnitNormExact)
- Guard streaming mode against V1 (not yet supported)
- Add version-mapping config test
- Add .xcode_derived/ and default.profraw to .gitignore

Benchmarked on 10.6s audio: Swift 0.24s vs Python 0.55s (2.3x faster).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace the per-timestep GPU GRU loop in pytorchGRULayer with a
CPU-side implementation using vDSP_mmul. Batch the input projection
as a single GPU matmul, then run the sequential hidden-state loop
on CPU via Accelerate, avoiding thousands of tiny GPU dispatches.

Same approach already used for V1's groupedGRUV1Packed.

Before → after (release, Apple Silicon):
  V2 10s:  1.09s → 0.17s  (6.4x)
  V2 52s:  2.81s → 0.57s  (4.9x)
  V3 10s:  0.61s → 0.19s  (3.2x)
  V3 52s:  3.23s → 0.55s  (5.9x)
  V3 2hr: 66s (~100x real-time)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Break DeepFilterNetModel.swift into well-organized files by concern:
- DeepFilterNetModel.swift (448 lines): error types, config, model class, loading, public API
- DeepFilterNetStreamer.swift (897 lines): streaming pipeline as top-level class
- DeepFilterNetForward.swift (777 lines): network forward pass (V1/V2/V3)
- DeepFilterNetLayers.swift (455 lines): conv2d, convTranspose, batchNorm, GRU, linear
- DeepFilterNetDSP.swift (433 lines): feature normalization, ERB, weight cache builders

Also adds benchmark scripts (speed, accuracy, latency, memory) used to verify
zero regression: all accuracy metrics are bit-identical before/after.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Expand DeepFilterNet README from 32 to ~120 lines: add HuggingFace model
link, architecture description, Swift/CLI usage examples, streaming config
reference, performance table, and license section.

Update main README STS table with proper HuggingFace repo link.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Document all public types, properties, and methods with /// comments:
- DeepFilterNetModel: class, properties, fromPretrained, fromLocal,
  enhance, createStreamer, enhanceStreaming (sync + async)
- DeepFilterNetStreamer: class, init, processChunk, flush, reset,
  profilingSummary
- DeepFilterNetStreamingConfig: all properties and init
- DeepFilterNetStreamingChunk: all properties
- DeepFilterNetConfig: struct-level doc
- DeepFilterNetError: enum-level doc

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Keep scripts locally for development use but exclude from the repository.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Align with the Python mlx-audio repo's default model. The mlx-community
repo contains v1/, v2/, v3/ subfolders. Add subfolder parameter to
fromPretrained/fromLocal to select the version (defaults to "v3").

Handles both subfolder layout (mlx-community) and flat layout (direct
model dir with config.json at top level).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
kylehowells and others added 2 commits March 23, 2026 23:15
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Consolidate the standalone sts-deepfilternet CI job into the existing
test matrix as an STS entry, avoiding a redundant runner.

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
lucasnewman
lucasnewman previously approved these changes Mar 23, 2026

@lucasnewman lucasnewman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good!

@kylehowells

kylehowells commented Mar 23, 2026 •

Copy link
Copy Markdown
Contributor Author

The force push was a rebase of the latest from main, as I wasn't able to get to updating this/replying this week and main moved on, so pulled in those changes and did a rebase so the merge was clean. Should hopefully all be good to go now. Though let me know if there are any other issues.

kylehowells and others added 2 commits March 24, 2026 22:18
…dditions

Conflicts resolved:
- .github/workflows/tests.yaml: kept our test matrix strategy
- .gitignore: combined both sets of entries

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
PR Blaizzy#121 intentionally removed the matrix strategy for Metal VM
stability. The merge conflict resolution re-added it but the run
step already uses -skip-testing instead of -only-testing:$matrix,
so the matrix was dead code. Remove it to stay aligned with main.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@kylehowells

Copy link
Copy Markdown
Contributor Author

Updated to fix merge conflicts from PRs #121–#123 landing on main.

Two files conflicted:

  • .github/workflows/tests.yaml — Dropped the test matrix strategy to match main. PR feat: add MLXAudioG2P — neural ByT5 + dictionary lexicons #121 intentionally removed it for Metal VM stability, and the single-job run with -skip-testing:SmokeTests already covers all suites including STS.
  • .gitignore — Combined entries from both sides.

Everything else auto-merged cleanly (G2P, StyleTTS2, KittenTTS don't overlap with DeepFilterNet).

@beshkenadze

Copy link
Copy Markdown
Contributor

@lucasnewman @kylehowells I’d like to point out that this implementation works well only for offline processing. Perhaps we should reflect this in the file or module name?

@lucasnewman

Copy link
Copy Markdown
Collaborator

@kylehowells This generally looks good -- it looks like the README mentions streaming support but the code doesn't. Can we rationalize that and either implement streaming or clarify the docs?

@beshkenadze

Copy link
Copy Markdown
Contributor

@lucasnewman I can provide a streaming solution, but it will use a different implementation, since the current one relies on batching.

@kylehowells

Copy link
Copy Markdown
Contributor Author

Thanks @lucasnewman I’ll update it with the streaming support hopefully tomorrow.

kylehowells and others added 5 commits March 26, 2026 19:18
Adds mlx-audio-swift-demo, a CLI tool that plays audio through a real-time
processing pipeline with DeepFilterNet noise reduction applied inline
hop-by-hop (10ms frames). Backpressure from AVAudioPlayerNode paces the
processing thread to real-time, simulating live audio from a phone call.

Controls: Space=play/pause, e=toggle effect, q=quit. Audio loops continuously.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This reverts commit 7ed285c07820614c15dfaf37f2bc90941414dc9f.
This reverts commit 17bed24656bac5eb2d5265901587622edd6ef7f2.
@kylehowells

Copy link
Copy Markdown
Contributor Author

@lucasnewman had a chance today to sit down properly and go through this, and I don't understand @beshkenadze's comment about lack of streaming support.

I've implemented DeepFilterNetStreamer.swift for that purpose. It's public, matches similar approaches to the APIs of other components in the project, it's mentioned in the README.md file and has code examples. I've benchmarked the per hop latency.

I thought from the above comments it got removed by mistake during cleanup or something, but it's all there.

let streamer = model.createStreamer(
    config: DeepFilterNetStreamingConfig(
        padEndFrames: 3,
        compensateDelay: true
    )
)

// Feed chunks as they arrive (any size, internally buffered to hop-sized frames)
let out1 = try streamer.processChunk(chunk1)
let out2 = try streamer.processChunk(chunk2)
// ...
let tail = try streamer.flush()  // flush remaining samples

In case something was broken I've also built a test app which lets you toggle applying the audio effect in realtime to an audio stream that is playing, so you can hear the effect and how much latency is involved.

The CLI tool is in the Git history I just pushed. I added it then reverted, to the diff shouldn't be any different, but you can run the streaming test tool by undoing the "revert commit".

mlx-audio-swift-demo --audio Tests/media/noisy_audio_10s.wav

I've tested and everything is working, streaming API, offline API.

Also ran some latency measuring unit tests to double check, here's the results + Opus writeup of where that time is being spent.


  • Model: DeepFilterNet3 (48kHz, hop=480 samples)
  • Audio: 10 seconds of noisy speech (480,000 samples)
  • Budget: 10ms per hop (480 samples / 48,000 Hz)

Results

Test 1: Real-Time Streaming (materializeEveryHops=1)

Forces eval() after every single hop to simulate true real-time output — the enhanced audio is fully materialised and available to the caller after each 10ms frame.

Metric Value
min 7.23ms
avg 7.68ms
median 7.42ms
p95 8.41ms
p99 8.86ms
max 115.8ms (JIT compilation spike on first eval)
Over 10ms budget 2/990 hops (0.2%)
RTF 0.77x (1.3x real-time)

Test 2: Batched Streaming (materializeEveryHops=512)

Allows MLX to build up a lazy computation graph across 512 hops (~5.1 seconds) before forcing evaluation. Still processes hop-by-hop through the streamer, but defers GPU synchronisation.

Metric Value
Total wall time 7.60s for 10s audio
Per-hop (amortised) 2.42ms
RTF 0.76x (1.3x real-time)

Test 3: Offline enhance()

Processes the entire 10-second file in a single batched forward pass (no streaming state, no per-hop eval).

Metric Value
Total wall time 0.357s for 10s audio
RTF 0.036x (28x real-time)

Where the Time Goes

The per-stage profiling breakdown for real-time mode (materializeEveryHops=1):

Stage Time % of Total What it does
materialize 5.12s 67.0% eval() — GPU synchronisation barrier. Forces MLX to flush its lazy computation graph and wait for the GPU to finish.
infer 2.31s 30.2% Neural network forward pass (encoder → GRU embedding → ERB decoder → DF decoder).
features 0.11s 1.4% ERB energy and unit-norm feature extraction from the STFT frame.
synthesis 0.05s 0.7% Inverse FFT + overlap-add to reconstruct the time-domain output.
analysis 0.05s 0.6% Windowed FFT of the input hop.

Inference sub-stages

Sub-stage Time % of Infer What it does
infer.erb 0.78s 35.2% ERB mask decoder (transposed conv + GRU)
infer.df 0.61s 27.6% Deep filter coefficient decoder (GRU + conv)
infer.enc 0.57s 25.9% Encoder convolutions (ERB path + DF path)
infer.emb 0.23s 10.6% GRU embedding layer

The Key Insight: The Model Only Takes ~2.4ms Per Hop

When materialize overhead is removed (Test 2), the per-hop cost drops to 2.42ms — well under the 10ms budget with 7.5ms of headroom.

The neural network inference itself is fast. The bottleneck in real-time mode is the GPU synchronisation cost of calling eval() every single hop. Each eval() call forces a round-trip to the Metal GPU: submit the computation graph, wait for it to complete, and copy the result back to the CPU. For a tiny workload like a single 480-sample hop, the fixed overhead of this synchronisation dominates the actual compute.

This means there is substantial room to batch additional GPU work alongside DeepFilterNet within the same eval() call. Since the model itself only uses ~2.4ms of GPU compute per hop, you could fit other lightweight per-frame processing (level metering, VAD, spectral visualisation, etc.) into the remaining ~7.5ms budget without impacting real-time performance — as long as everything shares the same eval() barrier rather than adding separate ones.


@kylehowells

Copy link
Copy Markdown
Contributor Author

Looking through the original rust API and how it does streaming I see some things I left out around configuration and control of the denoising, so adding those settings in now and will push an update to the PR with the extra control/config options exposed, but the API itself won't change (just new config settings).

@beshkenadze

Copy link
Copy Markdown
Contributor

@kylehowells Sorry, I must have misread the code and mistaken buffering for fixed batching!
My apologies for that. Once the result is finalized, I’ll test how it works in real time with a microphone and get back to you with a report.

@beshkenadze

beshkenadze commented Mar 27, 2026 •

Copy link
Copy Markdown
Contributor

@kylehowells I ran some bench against my implementation.
The issue is a latency of up to 106 ms, but overall it works very well (I tested it on your MLX version of the code).

Hardware specs: MBP 14 M1 Max 32Gb

image

@beshkenadze

Copy link
Copy Markdown
Contributor

Hey @kylehowells, I really like your implementation. Do you plan to continue working on this PR?

@kylehowells

Copy link
Copy Markdown
Contributor Author

@beshkenadze yes, apologies for the delay. I started deep diving into the streaming latency and then got busy with work and wasn’t able to work on this for a while.

Do you want to hold out until the streaming latency is better, or merge as is, and follow up with streaming improvements as subsequent PRs?

@beshkenadze

Copy link
Copy Markdown
Contributor

@kylehowells no worries.
Let's ask @lucasnewman 🙂

@lucasnewman

Copy link
Copy Markdown
Collaborator

I'm inclined to just merge and we can send follow ups if it's working as expected.

@lucasnewman
lucasnewman merged commit 5ddd0f6 into Blaizzy:main Apr 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants