Skip to content

Pasted text uses the temp file's random ID as the chapter title (narrated aloud, written into m4b tags) #194

Description

@cajones86

Pasted text uses the temp file's random ID as the chapter title (narrated aloud, written into m4b tags)
Version: abogen 1.3.1 (web UI) Platform: Windows 11, Python 3.12, CPU Output format: m4b

Summary
When you paste text into the wizard instead of uploading a file, the resulting chapter is titled with a 32-character random hex ID, e.g. b80d30a55ce94e8693cacca878e6b6c6. That ID is not just cosmetic:

it is spoken aloud by the narrator as the chapter intro (it comes out as a garbled "bdaceecaccaebc"),
it is written into the m4b as the chapter title, and
it becomes the album tag.
The title the user typed is applied to the title tag only.

Steps to reproduce
Open the web UI, choose the paste-text option.
Paste any text with no headings (so no chapter split is detected).
Set Title to Paste Test.
Convert to m4b.
Expected
The chapter title and album tag use the title the user supplied (Paste Test).

Actual
$ ffprobe -v error -print_format json -show_chapters
-show_entries format_tags=title,album Paste_Test.m4b
{
"chapters": [
{ "tags": { "title": "b80d30a55ce94e8693cacca878e6b6c6" } },
{ "tags": { "title": "Outro" } }
],
"format": { "tags": { "title": "Paste Test",
"album": "b80d30a55ce94e8693cacca878e6b6c6" } }
}
And in the job log, the ID is narrated:

Chapter 1 title · 15/135: bdaceecaccaebc.
Root cause
wizard_text() in abogen/webui/routes/main.py writes the pasted text to a temp file named after a random UUID:

file_path = temp_dir / f"{uuid.uuid4().hex}.txt" # line 216
extract_from_path() dispatches .txt to _extract_plaintext(), which passes the file stem as the default title (abogen/text_extractor.py, lines 101-104):

def _extract_plaintext(path: Path) -> ExtractionResult:
...
return _extract_from_string(raw, default_title=path.stem)
With no headings, _split_chapters() falls back to that default for the single chapter (lines 146-150), and _build_metadata_payload() uses the same value for ALBUM (line 202).

wizard_text() then patches only one of the three uses:

extraction = extract_from_path(file_path)

Override title since text extraction might not find one

extraction.metadata["title"] = title # line 225 - chapter title and album still hold the UUID
So the UUID survives into the chapter title and the album tag.

Suggested fix
Repair the remaining fallback uses in wizard_text(), right after the existing title override:

     extraction = extract_from_path(file_path)
     # Override title since text extraction might not find one
     extraction.metadata["title"] = title
  •    # Pasted text is stored as a randomly named temp file, and plain-text
    
  •    # extraction falls back to the file stem whenever it finds no headings.
    
  •    # That leaves the random id as the chapter title and album name, where
    
  •    # it gets narrated aloud and written into the m4b tags.
    
  •    placeholder = file_path.stem
    
  •    if extraction.metadata.get("album") == placeholder:
    
  •        extraction.metadata["album"] = title
    
  •    for chapter in extraction.chapters:
    
  •        if chapter.title == placeholder:
    
  •            chapter.title = title
    

This is deliberately conservative: it only replaces values that still equal the temp stem, so a title parsed out of the text itself (via the TITLE: metadata prefix or a detected heading) is left untouched.

An alternative worth considering is threading the caller's title into _extract_from_string() as the default_title instead of patching afterwards, which would fix this at the source for any caller that knows the real title.

Verified
With the patch applied, same repro:

{
"chapters": [
{ "tags": { "title": "Paste Test" } },
{ "tags": { "title": "Outro" } }
],
"format": { "tags": { "title": "Paste Test", "album": "Paste Test" } }
}
and the log now reads Chapter 1 title · 11/135: Paste Test.

Re-ran an EPUB conversion afterwards to confirm no regression - chapter titles and tags unchanged.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions