Skip to content

The /ɪz/ possessive hint splits the word in the audio, and reaches the subtitles #197

Description

@chrusch

Version: abogen 1.3.1 (pip), Python 3.12.13, macOS. Measured with the
kokoro that ships with it, British bm_george and American am_fenrir.

Summary

apply_phoneme_hints replaces the ‹IZ› marker with " iz" — a
space and a syllable — so a sibilant singular possessive is handed to
kokoro as two words. Two consequences, one of which I think is the more
serious:

  1. The audio breaks the word. Church's is spoken as "Church · iz",
    with an audible gap and a stress on the second part. This is what a
    listener notices first.
  2. The caption shows the pronunciation script. The subtitle is
    written from the post-normalization text, so the .srt reads
    the Church iz spiritual economy where the book reads
    the Church's spiritual economy.

And the finding that makes both avoidable: the hint is not needed.
kokoro already phonemizes the plain possessive to exactly the /ɪz/ the
marker is reaching for. The hint replaces a correct pronunciation with a
broken one.

Evidence

Phonemes, straight from KPipeline (lang_code="b", bm_george), for
the six words a real book actually hit:

input phonemes
the Church's reply ðə ʧˈɜːʧɪz ɹɪplˈI
the Church iz reply ðə ʧˈɜːʧ ˈɪz ɹɪplˈI
the Strauss's reply ðə stɹˈWsɪz ɹɪplˈI
the Strauss iz reply ðə stɹˈWs ˈɪz ɹɪplˈI
the James's reply ðə ʤˈAmzɪz ɹɪplˈI
the James iz reply ðə ʤˈAmz ˈɪz ɹɪplˈI

The raw possessive is already …ɪz as one word. The hinted form is two
words, and ˈɪz carries its own primary stress — which is the gap you
hear. Same for boss's, Ahaz's, Augustus's; same in American
(lang_code="a", am_fenrir): ðə ʧˈɜɹʧᵻz becomes ðə ʧˈɜɹʧ ˈɪz. In
one case the hint also perturbs the preceding article — the Ahaz's
gives ðə, the Ahaz iz gives ði.

Reproduction:

from kokoro import KPipeline
pipe = KPipeline(lang_code="b")
ph = lambda t: " ".join(s.phonemes for s in pipe(t, voice="bm_george", speed=1.0))

ph("the Church's reply")    # 'ðə ʧˈɜːʧɪz ɹɪplˈI'   <- already the target
ph("the Church iz reply")   # 'ðə ʧˈɜːʧ ˈɪz ɹɪplˈI' <- what abogen sends

And the caption half, through abogen's own API, all defaults:

from abogen.kokoro_text_normalization import ApostropheConfig, normalize_for_pipeline
from abogen.normalization_settings import build_apostrophe_config, get_runtime_settings

s = get_runtime_settings()
cfg = build_apostrophe_config(settings=s, base=ApostropheConfig())
normalize_for_pipeline("He does not contradict the Church's spiritual economy.",
                       config=cfg, settings=s)
# -> 'He does not contradict the Church iz spiritual economy.'

Cause

ApostropheConfig defaults to sibilant_possessive_mode = "mark" and
add_phoneme_hints = True, so the possessive becomes base + ‹IZ›
(kokoro_text_normalization.py:1533-1543), and apply_phoneme_hints
substitutes " iz" (:1997-2002, called at :2363). The leading space
is what makes it a separate token to the phonemizer.

The caption follows because emit_text normalizes first
(webui/conversion_runner.py:1841), feeds the normalized text to kokoro
(:1864), and writes each segment's graphemes — kokoro's record of
what it was given — as the subtitle text (:1873, :1898-1904).

Suggested fix

Default sibilant_possessive_mode to "keep". It fixes both
problems at once, adds no code, and by the table above loses nothing:
the plain possessive already produces the intended /ɪz/ in every case
measured, in both language codes.

Two narrower changes work less well, and are worth ruling out
explicitly:

  • Drop the space (text.replace(iz_marker, "iz")). Fixes the audio
    for Church and boss — phonemes identical to the raw possessive —
    but not for Strauss, James, Ahaz or Augustus, where
    Straussiz and friends phonemize differently again. And the caption
    still reads Churchiz.
  • sibilant_possessive_mode = "approx" (writes es). Matches the
    raw possessive's phonemes in 5 of the 6 words above, failing on the
    proper noun Ahaz. The caption then reads Churches, which is a real
    word but the wrong one.

If the hint is needed for some phonemizer or lang_code I have not
measured, then keeping it for audio while writing subtitles from the
un-normalized text would at least confine the damage to the audio. The
chunk payload already carries what that needs: chunk_text attaches
display_text and original_text beside text (chunking.py,
_attach_display_text). The honest caveat is that one chunk can produce
several kokoro segments, so the cue/segment mapping has to be decided
first — it is not a one-line swap.

Impact

Measured across one 543-page book's two audiobook editions, 160
chapters: 99 occurrences in 44 chapters (13 in 8 for the main
edition, 86 in 36 for the paraphrase) — Strauss iz, Saint James iz,
Augustus iz, Ahaz iz, Church iz. Each is both an audible break in
the narration and a wrong word in the caption, and they ride into the
MP4s along with the muxed subtitle track.

Nothing in the output flags it: every marker is substituted, so there is
no stray ‹…› to notice, and the source text is clean at every stage.


Sibling of #196, found in the same audiobook run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions