Skip to content

Latest commit

 

History

History
132 lines (92 loc) · 5.64 KB

File metadata and controls

132 lines (92 loc) · 5.64 KB

Quality control

AuK returns audio, not a confidence. If a generation collapses to silence, clips, drifts off the requested duration, or simply ignores the instruction, nothing upstream notices. On a 40-minute recording split into 170 model calls, "nothing notices" is not an acceptable position.

Every result is measured against what the operation promised to do.

Measurement

openspeech.qc.measure() reports, with no reference needed:

Metric What it catches
duration, sample_rate Truncation, wrong-length generation
peak_dbfs, true_peak_dbfs Clipping, inter-sample overshoot
rms_dbfs, lufs Level and perceived loudness
dc_offset A bias the model sometimes introduces
clipped_ratio How much of the file is pinned at full scale
nonfinite_samples NaN/Inf from a diverged sample
speech_ratio, silence_ratio A generation that collapsed to silence
spectral_flatness Noise-like output where speech was expected
f0_median_hz The pitch actually produced

compare() adds the reference-relative half: duration_ratio, level_delta_db, lufs_delta, timbre_similarity (cosine of mean log-mel spectra), pitch_delta_semitones, flatness_delta.

Pitch estimation

estimate_f0 runs autocorrelation over voiced frames with a normalised peak test and parabolic interpolation. On synthetic harmonic signals it lands within 0.1 Hz of the true f0, and it measures a known two-semitone shift as 2.00 semitones.

That accuracy is the point: without it, "raise the pitch by 2 semitones" is an instruction you send and hope about.

The gates

from openspeech.qc import evaluate
report = evaluate(op, result_audio, source_audio, expected_seconds=7.0)
report.verdict   # Verdict.PASS | WARN | FAIL

Integrity checks run on everything: finite samples, non-empty, not silent, no clipping, no DC offset.

Intent checks depend on the operation, and this is where the real value is:

Operation Check Tolerance
volume_edit Measured level change matches the requested dB ±4 dB, error
pitch_edit Measured f0 change matches the requested semitones ±1.2 st, warning
speed_edit Duration ratio matches 1/multiplier ±15%, error
enhance_speech, improve_quality, extract_vocals Output is not noisier than the input flatness ≤ +0.08, warning
Voice-preserving operations Spectral similarity to the source ≥0.55, warning

So a volume edit that did not land is caught, by measurement:

[error] volume_applied: level moved +4.0 dB, asked for +10.0 dB

Tolerances are deliberately loose. The model is generative, and a gate that fires on normal variation is a gate people switch off.

Intent checks are skipped when an integrity check already failed — there is no point asking whether silence has the right pitch.

Repair

What happens next depends on the kind of fault. Getting this distinction right saves real money.

Retry — for stochastic faults

A collapse to silence, a diverged sample, a wildly wrong duration: that is a bad draw. Retry with a different seed and more sampling steps.

retry_settings(report, attempt=1, policy)
# {'seed': 7919, 'nfe': 48}

Exact DSP repair — for arithmetic

A gain that landed 6 dB short is not a bad draw, it is a number. Re-rolling the model costs GPU time and usually lands somewhere else wrong. OpenSpeech closes the gap exactly:

repair: applied the residual +6.0 dB the model left on the table
        (asked +10.0 dB, got +4.0 dB)

retry_settings returns None when every failure is one DSP can fix, so those attempts are never spent.

Repairs currently applied:

Fault Repair
DC offset Subtract the mean. Exact.
Clipping / true-peak overshoot Iterative true-peak limiter to the ceiling
Volume shortfall Apply the residual gain
Length overrun/shortfall ≤5% Trim or pad — only for operations that promised equal length

That last restriction matters: trimming a content edit would cut off the words it was told to add. The condition is on the operation's declared duration policy, not on how close the number looks.

What is not repaired

A model that substituted the voice, smeared the spectrum, or ignored the instruction cannot be fixed downstream. Those are reported, the step is marked, and the run's verdict reflects it. The gate's job is to be honest, not to make every result look green.

Policy

from openspeech.qc import RepairPolicy

policy = RepairPolicy(
    max_attempts=3,        # total model attempts, including the first
    escalate_nfe=True,     # more ODE steps on each retry
    vary_seed=True,        # a different draw, not the same one again
    dsp_corrections=True,  # apply the exact fixes
    accept_warnings=True,  # do not burn a retry on a warning
)

accept_warnings=False makes the pipeline stricter at the cost of more model calls.

Reading a report

result = process(audio, "boost the volume by 10 dB", engine=engine)
result.verdict                       # worst verdict across all steps

for step in result.steps:
    print(step.op, step.verdict, step.attempts, step.engine_calls)
    for repair in step.repairs:
        print("  ", repair)
    for entry in step.quality:
        for check in entry["checks"]:
            if not check["passed"]:
                print("  ", check["severity"], check["name"], check["detail"])

--report out.json on the CLI writes the whole structure, including per-chunk quality entries, to disk. It is JSON-serialisable by design so it can be stored, diffed across runs, or shipped to a monitoring system.