Version: abogen 1.3.1 (pip), Python 3.12.13, macOS. Measured with the
kokoro that ships with it, British bm_george and American am_fenrir.
Summary
apply_phoneme_hints replaces the ‹IZ› marker with " iz" — a
space and a syllable — so a sibilant singular possessive is handed to
kokoro as two words. Two consequences, one of which I think is the more
serious:
- The audio breaks the word.
Church's is spoken as "Church · iz",
with an audible gap and a stress on the second part. This is what a
listener notices first.
- The caption shows the pronunciation script. The subtitle is
written from the post-normalization text, so the .srt reads
the Church iz spiritual economy where the book reads
the Church's spiritual economy.
And the finding that makes both avoidable: the hint is not needed.
kokoro already phonemizes the plain possessive to exactly the /ɪz/ the
marker is reaching for. The hint replaces a correct pronunciation with a
broken one.
Evidence
Phonemes, straight from KPipeline (lang_code="b", bm_george), for
the six words a real book actually hit:
| input |
phonemes |
the Church's reply |
ðə ʧˈɜːʧɪz ɹɪplˈI |
the Church iz reply |
ðə ʧˈɜːʧ ˈɪz ɹɪplˈI |
the Strauss's reply |
ðə stɹˈWsɪz ɹɪplˈI |
the Strauss iz reply |
ðə stɹˈWs ˈɪz ɹɪplˈI |
the James's reply |
ðə ʤˈAmzɪz ɹɪplˈI |
the James iz reply |
ðə ʤˈAmz ˈɪz ɹɪplˈI |
The raw possessive is already …ɪz as one word. The hinted form is two
words, and ˈɪz carries its own primary stress — which is the gap you
hear. Same for boss's, Ahaz's, Augustus's; same in American
(lang_code="a", am_fenrir): ðə ʧˈɜɹʧᵻz becomes ðə ʧˈɜɹʧ ˈɪz. In
one case the hint also perturbs the preceding article — the Ahaz's
gives ðə, the Ahaz iz gives ði.
Reproduction:
from kokoro import KPipeline
pipe = KPipeline(lang_code="b")
ph = lambda t: " ".join(s.phonemes for s in pipe(t, voice="bm_george", speed=1.0))
ph("the Church's reply") # 'ðə ʧˈɜːʧɪz ɹɪplˈI' <- already the target
ph("the Church iz reply") # 'ðə ʧˈɜːʧ ˈɪz ɹɪplˈI' <- what abogen sends
And the caption half, through abogen's own API, all defaults:
from abogen.kokoro_text_normalization import ApostropheConfig, normalize_for_pipeline
from abogen.normalization_settings import build_apostrophe_config, get_runtime_settings
s = get_runtime_settings()
cfg = build_apostrophe_config(settings=s, base=ApostropheConfig())
normalize_for_pipeline("He does not contradict the Church's spiritual economy.",
config=cfg, settings=s)
# -> 'He does not contradict the Church iz spiritual economy.'
Cause
ApostropheConfig defaults to sibilant_possessive_mode = "mark" and
add_phoneme_hints = True, so the possessive becomes base + ‹IZ›
(kokoro_text_normalization.py:1533-1543), and apply_phoneme_hints
substitutes " iz" (:1997-2002, called at :2363). The leading space
is what makes it a separate token to the phonemizer.
The caption follows because emit_text normalizes first
(webui/conversion_runner.py:1841), feeds the normalized text to kokoro
(:1864), and writes each segment's graphemes — kokoro's record of
what it was given — as the subtitle text (:1873, :1898-1904).
Suggested fix
Default sibilant_possessive_mode to "keep". It fixes both
problems at once, adds no code, and by the table above loses nothing:
the plain possessive already produces the intended /ɪz/ in every case
measured, in both language codes.
Two narrower changes work less well, and are worth ruling out
explicitly:
- Drop the space (
text.replace(iz_marker, "iz")). Fixes the audio
for Church and boss — phonemes identical to the raw possessive —
but not for Strauss, James, Ahaz or Augustus, where
Straussiz and friends phonemize differently again. And the caption
still reads Churchiz.
sibilant_possessive_mode = "approx" (writes es). Matches the
raw possessive's phonemes in 5 of the 6 words above, failing on the
proper noun Ahaz. The caption then reads Churches, which is a real
word but the wrong one.
If the hint is needed for some phonemizer or lang_code I have not
measured, then keeping it for audio while writing subtitles from the
un-normalized text would at least confine the damage to the audio. The
chunk payload already carries what that needs: chunk_text attaches
display_text and original_text beside text (chunking.py,
_attach_display_text). The honest caveat is that one chunk can produce
several kokoro segments, so the cue/segment mapping has to be decided
first — it is not a one-line swap.
Impact
Measured across one 543-page book's two audiobook editions, 160
chapters: 99 occurrences in 44 chapters (13 in 8 for the main
edition, 86 in 36 for the paraphrase) — Strauss iz, Saint James iz,
Augustus iz, Ahaz iz, Church iz. Each is both an audible break in
the narration and a wrong word in the caption, and they ride into the
MP4s along with the muxed subtitle track.
Nothing in the output flags it: every marker is substituted, so there is
no stray ‹…› to notice, and the source text is clean at every stage.
Sibling of #196, found in the same audiobook run.
Version: abogen 1.3.1 (pip), Python 3.12.13, macOS. Measured with the
kokoro that ships with it, British
bm_georgeand Americanam_fenrir.Summary
apply_phoneme_hintsreplaces the‹IZ›marker with" iz"— aspace and a syllable — so a sibilant singular possessive is handed to
kokoro as two words. Two consequences, one of which I think is the more
serious:
Church'sis spoken as "Church · iz",with an audible gap and a stress on the second part. This is what a
listener notices first.
written from the post-normalization text, so the
.srtreadsthe Church iz spiritual economywhere the book readsthe Church's spiritual economy.And the finding that makes both avoidable: the hint is not needed.
kokoro already phonemizes the plain possessive to exactly the /ɪz/ the
marker is reaching for. The hint replaces a correct pronunciation with a
broken one.
Evidence
Phonemes, straight from
KPipeline(lang_code="b",bm_george), forthe six words a real book actually hit:
the Church's replyðə ʧˈɜːʧɪz ɹɪplˈIthe Church iz replyðə ʧˈɜːʧ ˈɪz ɹɪplˈIthe Strauss's replyðə stɹˈWsɪz ɹɪplˈIthe Strauss iz replyðə stɹˈWs ˈɪz ɹɪplˈIthe James's replyðə ʤˈAmzɪz ɹɪplˈIthe James iz replyðə ʤˈAmz ˈɪz ɹɪplˈIThe raw possessive is already
…ɪzas one word. The hinted form is twowords, and
ˈɪzcarries its own primary stress — which is the gap youhear. Same for
boss's,Ahaz's,Augustus's; same in American(
lang_code="a",am_fenrir):ðə ʧˈɜɹʧᵻzbecomesðə ʧˈɜɹʧ ˈɪz. Inone case the hint also perturbs the preceding article —
the Ahaz'sgives
ðə,the Ahaz izgivesði.Reproduction:
And the caption half, through abogen's own API, all defaults:
Cause
ApostropheConfigdefaults tosibilant_possessive_mode = "mark"andadd_phoneme_hints = True, so the possessive becomes base +‹IZ›(
kokoro_text_normalization.py:1533-1543), andapply_phoneme_hintssubstitutes
" iz"(:1997-2002, called at:2363). The leading spaceis what makes it a separate token to the phonemizer.
The caption follows because
emit_textnormalizes first(
webui/conversion_runner.py:1841), feeds the normalized text to kokoro(
:1864), and writes each segment'sgraphemes— kokoro's record ofwhat it was given — as the subtitle text (
:1873,:1898-1904).Suggested fix
Default
sibilant_possessive_modeto"keep". It fixes bothproblems at once, adds no code, and by the table above loses nothing:
the plain possessive already produces the intended /ɪz/ in every case
measured, in both language codes.
Two narrower changes work less well, and are worth ruling out
explicitly:
text.replace(iz_marker, "iz")). Fixes the audiofor
Churchandboss— phonemes identical to the raw possessive —but not for
Strauss,James,AhazorAugustus, whereStraussizand friends phonemize differently again. And the captionstill reads
Churchiz.sibilant_possessive_mode = "approx"(writeses). Matches theraw possessive's phonemes in 5 of the 6 words above, failing on the
proper noun
Ahaz. The caption then readsChurches, which is a realword but the wrong one.
If the hint is needed for some phonemizer or lang_code I have not
measured, then keeping it for audio while writing subtitles from the
un-normalized text would at least confine the damage to the audio. The
chunk payload already carries what that needs:
chunk_textattachesdisplay_textandoriginal_textbesidetext(chunking.py,_attach_display_text). The honest caveat is that one chunk can produceseveral kokoro segments, so the cue/segment mapping has to be decided
first — it is not a one-line swap.
Impact
Measured across one 543-page book's two audiobook editions, 160
chapters: 99 occurrences in 44 chapters (13 in 8 for the main
edition, 86 in 36 for the paraphrase) —
Strauss iz,Saint James iz,Augustus iz,Ahaz iz,Church iz. Each is both an audible break inthe narration and a wrong word in the caption, and they ride into the
MP4s along with the muxed subtitle track.
Nothing in the output flags it: every marker is substituted, so there is
no stray
‹…›to notice, and the source text is clean at every stage.Sibling of #196, found in the same audiobook run.