Version: abogen 1.3.1 (pip), Python 3.12.13, macOS.
Summary
normalize_for_pipeline duplicates everything from an opening quotation
mark that has no matching closing mark to the end of the text. Because
emit_text calls normalize_for_pipeline on every chunk before handing
it to kokoro, the duplicated text is synthesized — the audiobook
says it twice, and the caption cue carries both copies.
Unclosed quotation marks are not an edge case here: sentence and
paragraph chunking both produce them routinely (details below).
Reproduction
from abogen.kokoro_text_normalization import ApostropheConfig, normalize_for_pipeline
from abogen.normalization_settings import build_apostrophe_config, get_runtime_settings
s = get_runtime_settings()
cfg = build_apostrophe_config(settings=s, base=ApostropheConfig())
normalize_for_pipeline("“Hello.", config=cfg, settings=s)
# -> '“Hello.“Hello.'
Only the tail from the unmatched mark is duplicated, and every mark in
_QUOTE_PAIRS is affected:
"«Bonjour." -> '«Bonjour.«Bonjour.'
"He said “hi”. Then “bye." -> 'He said “hi”. Then “bye.“bye.'
"A stray “ mark, and words." -> 'A stray “mark, and words.“mark, and words.'
"“Balanced.”" -> '“Balanced.”' # correct
"No quotes here." -> 'No quotes here.' # correct
Why unclosed marks are common, not exceptional
Both chunk levels in abogen/chunking.py produce them from ordinary,
correctly-punctuated prose.
Sentence level — a quotation of several sentences opens on the first
and closes on the last, so every chunk but the last has an open mark:
from abogen.chunking import chunk_text
para = ("“And it came to pass, when Jesus had finished all these sayings. "
"Then assembled together the chief priests. "
"But they said, not on the feast-day.”")
for c in chunk_text(chapter_index=0, chapter_title="T", text=para, level="sentence"):
print(c["normalized_text"])
# “And it came to pass, when Jesus had finished all these sayings.“And it came to pass, when Jesus had finished all these sayings.
# Then assembled together the chief priests.
# But they said, not on the feast-day.”
Paragraph level — the standard typographic convention for a
quotation spanning paragraphs is an opening mark on each paragraph and a
closing mark only on the last, so every paragraph but the last is
unclosed and duplicated:
text = ("“The first paragraph of the quotation runs on.\n\n"
"“The second paragraph continues it.\n\n"
"“And the third closes it.”")
# paragraphs 1 and 2 come back doubled; paragraph 3 is correct
Cause
abogen/kokoro_text_normalization.py, _normalize_all_caps_quotes
(line 1136). When the scan finds no closing mark it appends the
remainder and breaks without advancing index, so the tail block
below the loop appends the same remainder a second time:
if cursor >= length:
builder.append(text[index:]) # line 1157 — the remainder
break # line 1158 — index not advanced
...
if index < length:
builder.append(text[index:]) # line 1170 — the remainder again
The reachability path is
emit_text → normalize_for_pipeline (webui/conversion_runner.py:1841)
→ _normalize_all_caps_quotes (kokoro_text_normalization.py:2360,
enabled by default via normalization_caps_quotes).
Suggested fix
One line — mark the remainder as consumed before breaking:
if cursor >= length:
builder.append(text[index:])
index = length
break
return "".join(builder) in place of the break would do the same.
Skipping caps normalization inside the unclosed remainder looks
deliberate and correct (there is no delimited passage to normalize);
only the second append is wrong.
Impact
Found by ear in a generated audiobook. Across one 543-page book's two
editions: 38 chapters affected, 72 duplicated caption cues, ~1585 words
spoken twice — every scripture passage the book quotes, since each opens
a multi-sentence quotation. Nothing in the output flags it; the WAV, the
SRT, the M4B and the uploaded video all agree with each other, and only
listening reveals it.
Possibly related, much smaller — say the word and I will split it out
_cleanup_spacing treats the straight " as a closing mark only
(kokoro_text_normalization.py:691), so a space is inserted after an
opening one:
'"Hello there," he said.' -> '" Hello there," he said.'
'“Hello there,” he said.' -> '“Hello there,” he said.' # correct
Cosmetic for TTS, but it makes straight-quoted text render oddly in the
captions.
Version: abogen 1.3.1 (pip), Python 3.12.13, macOS.
Summary
normalize_for_pipelineduplicates everything from an opening quotationmark that has no matching closing mark to the end of the text. Because
emit_textcallsnormalize_for_pipelineon every chunk before handingit to kokoro, the duplicated text is synthesized — the audiobook
says it twice, and the caption cue carries both copies.
Unclosed quotation marks are not an edge case here: sentence and
paragraph chunking both produce them routinely (details below).
Reproduction
Only the tail from the unmatched mark is duplicated, and every mark in
_QUOTE_PAIRSis affected:Why unclosed marks are common, not exceptional
Both chunk levels in
abogen/chunking.pyproduce them from ordinary,correctly-punctuated prose.
Sentence level — a quotation of several sentences opens on the first
and closes on the last, so every chunk but the last has an open mark:
Paragraph level — the standard typographic convention for a
quotation spanning paragraphs is an opening mark on each paragraph and a
closing mark only on the last, so every paragraph but the last is
unclosed and duplicated:
Cause
abogen/kokoro_text_normalization.py,_normalize_all_caps_quotes(line 1136). When the scan finds no closing mark it appends the
remainder and breaks without advancing
index, so the tail blockbelow the loop appends the same remainder a second time:
The reachability path is
emit_text→normalize_for_pipeline(webui/conversion_runner.py:1841)→
_normalize_all_caps_quotes(kokoro_text_normalization.py:2360,enabled by default via
normalization_caps_quotes).Suggested fix
One line — mark the remainder as consumed before breaking:
return "".join(builder)in place of thebreakwould do the same.Skipping caps normalization inside the unclosed remainder looks
deliberate and correct (there is no delimited passage to normalize);
only the second append is wrong.
Impact
Found by ear in a generated audiobook. Across one 543-page book's two
editions: 38 chapters affected, 72 duplicated caption cues, ~1585 words
spoken twice — every scripture passage the book quotes, since each opens
a multi-sentence quotation. Nothing in the output flags it; the WAV, the
SRT, the M4B and the uploaded video all agree with each other, and only
listening reveals it.
Possibly related, much smaller — say the word and I will split it out
_cleanup_spacingtreats the straight"as a closing mark only(
kokoro_text_normalization.py:691), so a space is inserted after anopening one:
Cosmetic for TTS, but it makes straight-quoted text render oddly in the
captions.