To ensure broad coverage and reproducibility, we curate 32 tabular datasets: 24 well-established benchmarks from prior work and 8 recent additions published post-2025. Every dataset is drawn from trusted, verifiable sources, remains publicly accessible, and can be reliably reproduced, supporting both the integrity and extensibility of our evaluation framework.