Skip to content

Fix(ESP32): Correct PDM DAC clock configuration - #81

Closed
Meshwa428 wants to merge 2 commits into
earlephilhower:masterfrom
Meshwa428:master
Closed

Meshwa428 wants to merge 2 commits into
earlephilhower:masterfrom
Meshwa428:master

Conversation

@Meshwa428

Copy link
Copy Markdown

Description

This pull request resolves a critical bug in the ESP32PDMAudio class where audio playback was occurring at an incorrect, sped-up rate.

The Problem

Users of the ESP32PDMAudio class with a single-pin PDM DAC setup experienced audio playing back much faster than the original source. The root cause was the use of the default I2S PDM configuration (I2S_PDM_TX_CLK_DEFAULT_CONFIG), which is intended for PDM codecs. This configuration sets a clock rate proportional to the source sample rate.

For a single-pin PDM DAC output, a fixed, high-frequency clock is required, which the hardware then modulates based on the incoming PCM data. The incorrect clocking scheme led to the hardware consuming audio samples too quickly.

The Solution

This PR corrects the issue by switching the ESP32PDMAudio class to use the specific PDM DAC configurations provided by the ESP-IDF.

Specifically, it replaces:

  • I2S_PDM_TX_CLK_DEFAULT_CONFIG with I2S_PDM_TX_CLK_DAC_DEFAULT_CONFIG
  • I2S_PDM_TX_SLOT_DEFAULT_CONFIG with I2S_PDM_TX_SLOT_DAC_DEFAULT_CONFIG

These changes ensure the PDM hardware is configured correctly for a DAC-style output, providing a stable clock and resolving the playback speed issue.

The ESP32PDMAudio class was using the default I2S PDM configurations intended for a full PDM codec, not a single-pin PDM DAC output. This resulted in an incorrect clock rate and sped-up audio playback.

This commit switches to using the `I2S_PDM_TX_CLK_DAC_DEFAULT_CONFIG` and `I2S_PDM_TX_SLOT_DAC_DEFAULT_CONFIG` macros, which correctly configure the hardware for DAC mode. This ensures the PDM clock is stable and independent of the source sample rate, resolving the playback speed issue.
@earlephilhower

Copy link
Copy Markdown
Owner

Very nice, thanks!

@earlephilhower

Copy link
Copy Markdown
Owner

Does this only work @ 44.1khz? I get good results using PlayAACFromROM modified to have a ESP32PDMOutput (44.1kHZ) but when I go to SerialSpeak which runs at 22050 I get very loud noise and nothing intelligible.

If I had to hazard a guess, it's like either the high and low bytes of the samples got reversed in the PDM unit (i.e. playing 0x1882 as 0x8218) or there's some massive gain multiplier in-line (i.e. it's double-integrating the PDM bits or something).

You can try yourself if you take master and apply #82 (already merged, and a definite logic bug unrelated to actual IDF clocking) and run

#include <BackgroundAudioSpeech.h>

#include <libespeak-ng/voice/en_029.h>
#include <libespeak-ng/voice/en_gb_scotland.h>
#include <libespeak-ng/voice/en_gb_x_gbclan.h>
#include <libespeak-ng/voice/en_gb_x_gbcwmd.h>
#include <libespeak-ng/voice/en_gb_x_rp.h>
#include <libespeak-ng/voice/en.h>
#include <libespeak-ng/voice/en_shaw.h>
#include <libespeak-ng/voice/en_us.h>
#include <libespeak-ng/voice/en_us_nyc.h>
BackgroundAudioVoice v[] = {
  voice_en_029,
  voice_en_gb_scotland,
  voice_en_gb_x_gbclan,
  voice_en_gb_x_gbcwmd,
  voice_en,
  voice_en_shaw,
  voice_en_us,
  voice_en_us_nyc
};
#include <ESP32PDMAudio.h>
ESP32PDMAudio audio(5);
BackgroundAudioSpeech BMP(audio);

void setup() {
  Serial.begin(115200);
  // We need to set up a voice before any output
  BMP.setVoice(v[0]);
  delay(3000);
  Serial.printf("Ready!\r\n");
  BMP.begin();
  BMP.speak("I have awoken.  Tremble in fear!");
}

void loop() {
}

Reapply this patch and you'll get the screeching...

@Meshwa428

Copy link
Copy Markdown
Author

Yes I am working on that part, will update soon

@Meshwa428

Copy link
Copy Markdown
Author

are you okay with adding an extra function for setting upsampling ? like BMP.setInternalUpsampling(2);

@earlephilhower

Copy link
Copy Markdown
Owner

We really need to avoid anything like that. It's a change to all the generators, potentially, and not necessarily a constant.

For example, say I have 2 WAVs in a row: one at 44.1khz, one at 16khz. I couldn't even know, at the app level, when the 1st WAV's last sample went out to do that call.

It would also mess up how things are currently only a 1-line change to go from PWM to I2S to PDM.

@earlephilhower

Copy link
Copy Markdown
Owner

If PDM-DAC can only work with a fixed frequency, then the PDM object should do the resampling of inputs.

But what's strange is that the original clocking mode that's in there now, as long as you take into account the 1.25x speedup, does work for lower and higher sampling rates just fine. So I'm sure there's a way (but probably poorly documented...the IDF has a lot of text but it still leaves a lot of questions!)...

@Meshwa428

Copy link
Copy Markdown
Author

Ok so i tried fixing it and from our perspective Espeak is the only one being at the ~22k frequency. Then we can upsample it by detecting if the output engine is PDM or not ???

something similar to this seems feasible ??

    bool begin() {
        if (_playing || !_voice || !_voiceLen) {
            return false;
        }

        // [FIX] Automatically detect if the output is ESP32 PDM and enable up-sampling if needed.
        // This check is safe because the type is known within the library ecosystem.
        #ifdef ESP32
        if (dynamic_cast<ESP32PDMAudio*>(_out) != nullptr) {
            _upSample = 2;
        }
        #endif

        espeak_EnableSingleStep();
        espeak_InstallDict(__espeakng_dict, __espeakng_dictlen);
        espeak_InstallPhonIndex(_phonindex, sizeof(_phonindex));
        espeak_InstallPhonTab(_phontab, sizeof(_phontab));
        espeak_InstallPhonData(_phondata, sizeof(_phondata));
        espeak_InstallIntonations(_intonations, sizeof(_intonations));
        espeak_InstallVoice(_voice, _voiceLen);

        int samplerate = espeak_Initialize(AUDIO_OUTPUT_SYNCH_PLAYBACK, 20, nullptr, 0);
        espeak_SetVoiceByFile("INTERNAL");
        espeak_SetSynthCallback(_speechCB);

        _out->setBuffers(5, framelen);
        _out->onTransmit(&_cb, (void *)this);
        _out->setBitsPerSample(16);
        _out->setStereo(true);
        // [FIX] Set the hardware output frequency to the potentially up-sampled rate.
        _out->setFrequency(samplerate * _upSample);
        _out->begin();

        uint16_t zeros[32] __attribute__((aligned(4))) = {};
        while (_out->availableForWrite() > 32) {
            _out->write((uint8_t *)zeros, sizeof(zeros));
        }

        _playing = true;

        return true;
    }

@earlephilhower

Copy link
Copy Markdown
Owner

eSpeak is not the only one,. WAV could also be some random value from 8K to 48K. I also don't think you can dynamic cast, we don't generally have RTTI enabled in the compiler because it's very expensive WRT RAM/CPU.

While I appreciate the effort, I think maybe you're overthinking it.

The 1.25x scale here seems to work at 44.1 and 22050, so if we need to bust open everything to use the DAC mode, it doesn't seem worth it to me...

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants