Training code for the SEGMENT submission: Qwen3.6-27B with a rank-8 LoRA adapter, trained on 23,454 examples at 128 frames per clip.
annotations/ our frame-level annotations + the VQA pairs generated from them
frame_data/ question generation from foreign-object annotations
segment_data/ question derivation from already-labelled object tracks
orena_sft/ shared modules (system prompt, frame-track dataset builder)
segment_track/ SEGMENT dataset builder, collator, trainer, preflight
slurm/ the two jobs that run the pipeline
The annotation data is not committed. It comes in two halves.
Public — HeiCo and hernia questions, the hernia annotation sheets, and the hernia frames. Published as a Hugging Face dataset:
python annotations/fetch.pywhich downloads
Machine-Learning-Oncology/orena-segment-annotations
into annotations/public/.
Private — the LapChole questions and sheets. Not hosted: the LapChole-FOCUS data usage
agreement forbids publishing any derivative until the organisers release that dataset.
Available on request as a small archive; unpack it into annotations/private/:
mkdir -p annotations/private
tar xzf orena-segment-lapchole-annotations_*.tar.gz \
-C annotations/private --strip-components=1Either half is optional — the pipeline uses whatever is present. The expected layout is:
annotations/
├── public/ vqa/*.jsonl, raw/hernia_mesh_annotations/, frames/, SOURCES.md
└── private/ vqa/lapchole.jsonl, raw/lapchole_annotations/, raw/crop_boxes.json
See annotations/README.md for the schemas and the reasoning behind the split.
- Python 3.12, 8× H100 80 GB (single node) for training; CPU only for data preparation.
pip install -r requirements.txt— installtorchfrom the CUDA 12.8 index first.Qwen/Qwen3.6-27Bcached in$HF_HOME(~56 GB). Training runs withHF_HUB_OFFLINE=1.- The FOCUS dataset root, containing
heico/,lapchole/andhernia_yt/. - Foreign-object annotation CSVs (
lapchole_annotations/, hernia sheets).
Both jobs are configured through environment variables. No paths are hardcoded.
| variable | required | meaning |
|---|---|---|
PYTHON |
yes | interpreter of the environment built from requirements.txt |
ORENA_DATA_ROOT |
yes | directory holding heico/, lapchole/, hernia_yt/ |
HF_HOME |
training | cache containing Qwen/Qwen3.6-27B |
ORENA_ANNOT_DIR |
no | directory holding crop_boxes.json (default annotations/private/raw) |
LAPCHOLE_CSV_DIR |
no | LapChole sheets (default annotations/private/raw/lapchole_annotations) |
MESH_CSV_DIR |
no | hernia sheets (default annotations/public/raw/hernia_mesh_annotations) |
ORENA_ANNOT_FRAMES |
prep | where extracted annotation frames are written |
EXPORT |
no | dataset output directory (default ./sft_export_128) |
N_FRAMES |
no | frames per clip (default 128, must be even) |
RUN_NAME |
no | checkpoint directory name |
NPROC |
no | GPUs per node (default 8) |
There are two routes. Use route A unless you are regenerating the annotations from raw sheets — it needs no annotation sheets, no LapChole videos and no question generators.
export PYTHON=/path/to/venv/bin/python
export ORENA_DATA_ROOT=/path/to/orena
python annotations/fetch.py --private # or without --private
# 1. build the official export at 128 frames from YOUR copy of FOCUS
$PYTHON segment_track/build_segment_sft_dataset.py \
--datasets heico lapchole --root-dir "$ORENA_DATA_ROOT" \
--num-frames 128 --out-dir sft_export_128
cat sft_export_128/{train,eval,test}.jsonl > sft_export_128/all_gold.jsonl
# 2. merge the released questions into it
$PYTHON segment_track/integrate_released_annotations.py \
--export sft_export_128 --annotations annotations --n-frames 128HeiCo and LapChole questions inherit their clip from the matching window in your export,
which is why those frames are never redistributed. Hernia questions compute indices from the
window and use the frames that ship with the release. 160 LapChole questions annotate windows
absent from the official export; they carry explicit frames_indices and need
--extra-frames <dir> pointing at frames extracted from the LapChole videos, or they are
skipped with a warning.
export PYTHON=/path/to/venv/bin/python
export ORENA_DATA_ROOT=/path/to/orena
sbatch slurm/1_prepare_data.sbatchRuns five steps and writes $EXPORT/train_128.jsonl:
| step | output | rows |
|---|---|---|
| 1 | official export rebuilt at 128 frames → all_gold.jsonl |
20,000 |
| 2 | derived from labelled tracks → segment_derived_all.jsonl |
4,202 |
| 3 | lapchole gallstone / AHA → lapchole_segment.jsonl |
627 |
| 4 | hernia / mesh → mesh_segment.jsonl |
440 |
| 5 | merge, deduplicate, shuffle → train_128.jsonl |
23,454 |
Step 5 asserts that every row carries exactly N_FRAMES frames and fails otherwise.
Individual generators can also be run directly:
$PYTHON -m segment_data.derive_segment --out segment_derived_all.jsonl
$PYTHON -m frame_data.lapchole_segment_questions --n-frames 128 --extract
$PYTHON -m frame_data.mesh_segment_questions --n-frames 128
$PYTHON segment_track/build_segment_sft_dataset.py --datasets heico lapchole --num-frames 128export PYTHON=/path/to/venv/bin/python
export HF_HOME=/path/to/hf_cache
sbatch slurm/2_train.sbatch| base model | Qwen/Qwen3.6-27B |
| adapter | LoRA r=8, α=16, dropout 0.05 |
| target modules | q_proj k_proj v_proj o_proj gate_proj up_proj down_proj (vision tower frozen) |
| batch | per-device 1 × grad-accum 4 × 8 GPUs = effective 32 |
| schedule | lr 1e-4, bf16, seed 42, 2 epochs = 1,466 steps |
| checkpoints | every 366 steps → checkpoints/$RUN_NAME/ |
| runtime | ~11 h 30 m; 73.8 GB peak allocated per GPU |
The job verifies the training file exists and that every clip has N_FRAMES frames, then
runs segment_track/preflight.py, which gates triton>=3.7.1 and flash-linear-attention.
The run saves every 366 steps and the final checkpoint, checkpoint-1466 (epoch 2.0), is
the one submitted.
| checkpoint | epoch | eval_loss |
|---|---|---|
checkpoint-366 |
0.50 | 0.2612 |
checkpoint-732 |
1.00 | 0.1897 |
checkpoint-1098 |
1.50 | 0.1438 |
checkpoint-1464 |
2.00 | 0.1171 |
checkpoint-1466 |
2.00 | 0.1170 |
Merge the adapter from checkpoints/$RUN_NAME/checkpoint-1466 into the base model, then
build the container. N_FRAMES in the container's frame sampler must equal the value used
for training.
N_FRAMESmust be even — the model fuses frames pairwise (temporal_patch_size=2).- Frames are sampled on a fixed 15 fps grid, so changing
N_FRAMESneeds no re-extraction. segment_track/build_segment_sft_dataset.pydefaults--root-dirto a path that may not exist on your cluster; setORENA_DATA_ROOTor pass--root-direxplicitly.
We are grateful to Dr. Todd S. Harris of California Hernia Specialists, who kindly allowed us
to use twelve of his publicly available laparoscopic hernia-repair videos in this work. His
permission covers non-commercial research use and the release of the annotated frames; the
original videos remain his and are not redistributed. The video titles and links are listed
in annotations/public/SOURCES.md.