Yi Tang1,*
Xinyi Shang1,2,*
Jiacheng Cui1
Sondos Mahmoud Bsharat1
Jiacheng Liu1
Xiaohan Zhao1
Tran Dinh Tien1
Ahmed Elhagry1
Salwa K. Al Khatib1
Tianjun Yao1
Yonina C. Eldar3
Jing-Hao Xue2
Hao Li1
Salman Khan1
Zhiqiang Shen1,†
1Mohamed bin Zayed University of Artificial Intelligence 2University College London 3Weizmann Institute of Science
* Equal contribution | † Corresponding author
Predicted tampered pixels — Ours vs. PIXAR — on four held-out (OOD) generators: GPT-Image-2.0, Gemini-3.1, FLUX.2, Seedream 4.5.
TL;DR We study domain generalization for pixel-level image tampering detection across modern VLM generators. A simple training recipe — balanced real/tampered mini-batch sampling, late injection of a small companion source, and a low constant learning rate — improves cross-generator localization by up to 26.1% / 26.8% relative gIoU / cIoU at 13B (+21.4% / +21.1% at 7B) over the prior state of the art (PIXAR), while training on only 19.2% of its data.
- [2026-07] 🚀 Code for training, dataset construction, and evaluation is released.
Modern VLMs (ChatGPT, Gemini, Qwen-Image, …) differ sharply in architecture, editing pipeline, and post-processing, so a tampering detector trained on one generator often overfits to model-specific artifacts and degrades on unseen generators. We treat each generator as a domain and target cross-generator generalization for pixel-level tampering localization.
Our framework leaves the segmentation-based detector unchanged and improves only the training recipe, with three simple, compatible strategies:
- Balanced mini-batch sampling — every step sees a fixed real:tampered ratio, preventing the optimizer from collapsing onto clean-image priors or tampering artifacts.
- Late injection — train to convergence on a large base source (Qwen-Image), then inject a small companion source (Gemini-2.5) so emerging-domain signal is absorbed without dominating early feature learning.
- Low constant learning rate — a conservative constant
2e-5(warmup-then-hold) schedule preserves transferable forensic cues during adaptation.
The detector is built on the PIXAR architecture: a LLaVA (LLaMA-2 + CLIP ViT-L/14) backbone with LoRA and a SAM ViT-H mask decoder. Three task tokens drive three prediction heads, and the language model additionally generates a natural-language description of the edit:
| Token | Head | Output |
|---|---|---|
[CLS] |
classification | real vs. tampered |
[OBJ] |
multi-label recognition | tampered object categories (80 COCO classes + background) |
[SEG] |
SAM prompt | pixel-level tampering mask |
Training optimizes five losses — semantic (multi-label), BCE, DICE, classification, and text. The recipe weights λ_dice=1.0, λ_sem=0.5, λ_text=3.0 are set in scripts/train.sh (the semantic loss is implemented as --obj_loss_weight in train_PIXAR.py).
Balanced sampling stabilizes optimization: the gradient norm of the [CLS] head stays smooth and flat instead of fluctuating under naive sampling, and training settles into a wider, flatter loss basin.
Left: [CLS]-head gradient norm, random (red) vs. balanced (green). Middle & right: loss landscape, PIXAR vs. Ours.
Cross-generator performance, averaged over four held-out OOD generators (GPT-Image-2.0, Gemini-3.1, FLUX.2, Seedream 4.5):
| Method | Pixel Recall | Pixel F1 | gIoU | cIoU | Binary Acc. |
|---|---|---|---|---|---|
| PIXAR-7B | 25.8 | 28.5 | 0.159 | 0.166 | 69.6 |
| Ours-7B | 45.1 | 33.3 | 0.193 | 0.201 | 79.6 |
| PIXAR-13B | 33.5 | 31.0 | 0.176 | 0.183 | 59.0 |
| Ours-13B | 62.2 | 37.4 | 0.222 | 0.232 | 84.0 |
Ours-7B is trained on 73,353 tampered images (70,000 Qwen-Image + 3,353 Gemini-2.5) — 19.2% of PIXAR's 380K — yet improves every averaged metric at both scales. See the paper for per-generator and in-domain results.
Python 3.10. The base requirements install a CUDA 11.7 (cu117) PyTorch; fix/fix.sh then migrates the environment to CUDA 12.1 (cu121 PyTorch, deepspeed>=0.14.4, matching bitsandbytes/libstdc++). Run it after installing:
conda create -n pixar python=3.10 -y && conda activate pixar
pip install -r requirements.txt
bash fix/fix.sh # migrate to CUDA 12.1 (run fix/fix_again.sh if issues persist)Place pretrained weights under pretrains/ (or symlink it):
| Weight | Source |
|---|---|
pretrains/PIXAR-7B |
base detector for the main recipe — PIXAR-7B (PIXAR); PIXAR-13B from the same release |
pretrains/sam_vit_h_4b8939.pth |
SAM ViT-H |
openai/clip-vit-large-patch14 |
auto-downloaded |
The code reads data/, pretrains/, and writes to outputs/ as relative directories — create them or symlink to your storage.
The training set is built on top of the public PIXAR benchmark — download it first, then run the preprocessing scripts. Full details and options are in docs/DATA.md.
1. Download the PIXAR benchmark (preprocessed at τ = 0.05; Shang et al., 2026) and place it under data/.
2. Split into train / validation:
python preprocess/split_pixar.py \
--src data/PIXAR_preprocessed/test_full_0.05/full_0.05 \
--dst data/pixar_0.053. Build the headline training set:
python preprocess/make_qg_70k_subsets.py \
--variant 3353x1 \
--qwen_src data/PIXAR_preprocessed/train_0.05/ours_0.05 \
--gemini_src data/pixar_0.05 \
--out_dir data
# -> data/pixar_qg_70k_3353x1 (73,353 train tampered; validation symlinked from pixar_0.05)make_qg_70k_subsets.py also builds the base-source size variants used in the data-scale study — pass --variant 30k_3353x1, 150k_3353x1, or 380k_3353x1.
conda activate pixar
bash scripts/train.sh # PIXAR init, full recipeRecipe: PIXAR init, data/pixar_qg_70k_3353x1, lr 2e-5 constant, 4 epochs × 2500 steps, real:tampered = 1:1, with late injection of Gemini-2.5 from epoch 1 ({"0-1":{"gemini":0},"1-end":{"gemini":2}}). Then merge the checkpoint into a standalone HuggingFace model:
bash scripts/merge.sh --exp_name ours --base_model pretrains/PIXAR-7B
# -> outputs/merged/oursbash scripts/eval.sh --model outputs/merged/ours --dataset_dir data/pixar_0.05metrics.json reports overall accuracy, pixel-level gIoU / cIoU / F1, and a per-generator breakdown under per_model_metrics. The paper's headline numbers are the mean over the four OOD generators; see docs/EVAL.md for the exact field-to-table mapping and the OOD-average recipe.
python inference.py \
--version outputs/merged/ours \
--vision_pretrained pretrains/sam_vit_h_4b8939.pth \
--seg_prompt_mode fuse \
--image_paths path/to/image.png \
--output_dir example_outputsThe same command runs any merged checkpoint — just point --version at it.
| Model | Link |
|---|---|
| Ours-7B | coming soon |
| Ours-13B | coming soon |
This work builds on PIXAR, SIDA, LISA, LLaVA, and SAM. We thank the authors for releasing their code and models.
@article{tang2026simpledg,
title={Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs},
author={Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li, Salman Khan, and Zhiqiang Shen},
year={2026},
eprint={2607.18230},
archivePrefix={arXiv},
primaryClass={cs.CV},
}
