Skip to content

fix(vibevoice-asr): add audio resampling and normalization to preprocessing - #510

Merged
Blaizzy merged 3 commits into
Blaizzy:mainfrom
bellkjtt:fix/vibevoice-asr-preprocessing
Feb 19, 2026
Merged

Blaizzy merged 3 commits into
Blaizzy:mainfrom
bellkjtt:fix/vibevoice-asr-preprocessing

Conversation

@bellkjtt

Copy link
Copy Markdown
Contributor

Summary

  • Add audio normalization (-25 dB FS) to match the official VibeVoice PyTorch pipeline
  • Add sampling_rate parameter to generate() and stream_transcribe() so that array inputs at non-24 kHz rates are resampled correctly
  • Previously, passing a raw waveform array (e.g. loaded via librosa at 16 kHz) resulted in poor transcription because the audio was fed to the model without resampling or normalization

Details

The official VibeVoice ASR pipeline (PyTorch) applies two critical preprocessing steps that were missing in the MLX implementation:

  1. Resampling to 24 kHz - The model expects 24 kHz input. When users load audio at a different sample rate and pass it as an array, _preprocess_audio now resamples via resample_audio().
  2. Loudness normalization - The official pipeline normalizes audio to -25 dB FS with clipping avoidance. A new _normalize_audio() static method replicates this behavior.

Test plan

  • Test with audio loaded at 16 kHz via librosa - should now produce correct transcriptions
  • Test with audio file path (str) - should behave identically to before
  • Test with audio already at 24 kHz - should apply normalization without resampling
  • Compare output quality against the official PyTorch VibeVoice ASR demo

…essing

The VibeVoice ASR model requires 24 kHz input audio normalized to -25 dB FS,
matching the official PyTorch pipeline. Previously, when users passed raw
waveform arrays (e.g. loaded via librosa at 16 kHz), the audio was used
as-is without resampling or normalization, causing poor transcription quality.

Changes:
- Add `_normalize_audio()` static method that replicates the official
  VibeVoice AudioNormalizer (target -25 dB FS with clipping avoidance)
- Add `sampling_rate` parameter to `_preprocess_audio()`, `generate()`,
  and `stream_transcribe()` so callers can specify the input sample rate
- When an array is passed with a non-24kHz sample rate, resample via
  `resample_audio()` before encoding
- Always normalize array inputs to -25 dB FS for consistent loudness

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@bellkjtt
bellkjtt force-pushed the fix/vibevoice-asr-preprocessing branch from 96b18ff to af25c90 Compare February 18, 2026 04:40
bellkjtt and others added 2 commits February 18, 2026 13:53
Change generate() and stream_transcribe() defaults to match the
official VibeVoice ASR Gradio demo parameters:
- top_p: 0.95 -> 1.0 (no filtering)
- top_k: 25 -> 0 (disabled)
- min_p: 0.02 -> 0.0 (disabled)
- repetition_penalty: 1.2 -> 1.0 (no penalty)

This produces pure greedy decoding (temperature=0 with no additional
sampling filters), which matches the official demo and improves
speaker diarization accuracy.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@Blaizzy Blaizzy left a comment •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks!

@Blaizzy
Blaizzy merged commit f89c122 into Blaizzy:main Feb 19, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants