Skip to content

Assign label_seq to every subchain of a multi-chain PDB entity - #255

Open
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/pdb-parser-multichain-subchains
Open

Assign label_seq to every subchain of a multi-chain PDB entity#255
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/pdb-parser-multichain-subchains

Conversation

@fnachon

@fnachon fnachon commented Jul 10, 2026

Copy link
Copy Markdown

Fixes #236.

When a PDB COMPND record declares an entity spanning multiple chains (e.g. CHAIN: A, B), the parser only processed entity.subchains[0], so residues in every subchain after the first were left with an empty label_seq. That marks them is_present=False, so they're dropped during CIF serialization while feature.pkl still has them, producing a residue-count mismatch that crashes downstream steps (inverse fold, etc.) that load both and expect consistent dimensions. Now assigns label_seq to every subchain of the entity, not just the first.

fnachon and others added 9 commits January 10, 2026 15:54
Changes made to run without errors on the Mac MPS device: torch.autocast, number of devices and workers to use on M1-5 chips, workaround for CUDA-specific code, handling of float64 incompatibilities for MPS.
Replace hardcoded torch.autocast("cuda") with device-agnostic
device_type=tensor.device.type in confidence_utils, inverse_fold,
and writer modules introduced in the upstream merge.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Python pickle does not preserve RDKit atom-level SetProp values. When
PyTorch DataLoader spawns worker processes (default num_workers=1 on
macOS), self.canonicals is pickled and all atom 'name' properties are
lost, causing KeyError in process_atom_features.

Fix: load all required molecules directly from the moldir zip inside
each get_sample() / get_feat() call instead of using the pickled
self.canonicals. The moldir zip handle is cached per-process by
_get_zipfile(), so there is no repeated I/O overhead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ders

- Disable pin_memory on MPS (unsupported, causes UserWarning)
- Enable persistent_workers when num_workers > 0 (avoids repeated
  worker init overhead and the PL suggestion warning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously only entity.subchains[0] was processed, so when a COMPND
record declared an entity spanning multiple chains (e.g. "CHAIN: A,
B"), residues in every subchain after the first were left with an
empty label_seq. That marks them is_present=False, so they get
dropped during CIF serialization while feature.pkl still has them,
causing a residue-count mismatch that crashes downstream steps
(inverse fold, etc.) that load both and expect consistent dimensions.

Fixes HannesStark#236

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: PDB parsing drops residues from secondary subchains in multi-chain entities

1 participant