micromamba activate <environment>In order to re-produce the MIA methods, you first need a synthetic dataset. You can refer to Blue team home page for detailed instructions.
-
In single-cell RNA-seq data, we focus on donor-level privacy. The dataset is initially split into donor-based train and test sets, making it feasible to assess privacy at this level. To set up a membership inference attack (MIA), we label samples (cells) from donors in the train set as 1 and those from donors in the test set as 0. After applying the MIA method, we aggregate the predictive scores per sample at the donor level and compute the AUROC and accuracy based on these donor-level average scores.
-
Given the high dimensionality of the scRNA-seq dataset, we performed subsampling to create a simplified demonstration. Specifically, we randomly selected 50 donors from both the train and test sets. The synthetic data generator (Poisson) was trained on the full train set, while the subsampled train set was used to guide the generation of synthetic data.
-
To facilitate reproducibility, we provide the following resources on the ELSA Benchmark website for download:
- the subsampled synthetic dataset:
distr_Poisson_subset_onek1k_annotated_synthetic.h5ad - MIA labels:
onek1k_subset_mia_labels.csv - the demonstration test dataset:
onek1k_subset_mia.h5ad
- the subsampled synthetic dataset:
-
Make sure to use the same donor subset and experimental setting to evaluate your own synthetic datasets.
-
We acknowledge that this demonstration may not fully capture real-world scenarios. Therefore, we strongly encourage participants to propose novel solutions for investigating donor-level privacy in synthetic single-cell datasets.
- We are particularly interested in approaches that are not only realistic but also memory- and compute-efficient. Your innovative methods and insights will be crucial in advancing this important area of research. 🚀
-
Assuming you generated a synthetic dataset using the baseline pipeline, we provide a guideline to run the modified version of GAN-leaks (
gan_leaks) provided by DOMIAS package. In this modification, we introduced batching to enable faster GPU compute. -
attack_modelparameter inside theconfig.yamlis by default set assc_domias_baselines. This configuration will rungan_leaksagainst the generator defined in thegenerator_config. -
Please ensure that
config.yamlis palced in the same directory where you are running your experiment. -
📈 Feel free to use relevant public datasets as a reference set in case your method depends on it.
In the below example, config.yaml runs sc_domias_baselines attack models on the synthetic dataset generated by distr_Poisson method for the given configuration subset on onek1k dataset.
Please modify the config according to your needs.
.
.
.
dataset_config:
name: "onek1k"
membership_label_col: "membership"
attack_model: "sc_domias_baselines"
generator_config:
model_name: "sc_dist"
experiment_name: "distr_Poisson_subset"
sc_domias_baselines_config:
taxonomy: "bb" # not used provided as an example Then use the below script to run the method,
# the above configuration assumes that the path to synthetic data is configured as below
# you can simply place the synthetic data you downloaded from Benchmark website under this directory
#synthetic_data_pth = /path/to/project/data_splits/onek1k/synthetic/sc_dist/distr_Poisson_subset
# /path/to/dataset/ >> path for the membership test dataset
# --mmb_labels_file /path/to/example/label/csv/ >> path for the membership test labels CSV file
python {src_dir}/mia/red_team.py run-singlecell-mia
/path/to/synthetic/data/ + "distr_Poisson_subset_onek1k_annotated_synthetic.h5ad"
/path/to/dataset/ + "onek1k_subset_mia.h5ad"
"experiment_sc_synthetic_data"
--mmb_labels_file /path/to/example/label/csv/ + "onek1k_subset_mia_labels.csv"-
The script will generate an
evaluation_results.csvunder/{home_dir}/results/mia/{dataset_name}/{attacker_name}/{generator_model}/{experiment_name}/{mia_experiment_name} -
In case
--mmb_labels_fileoptional argument is not available, the code only saves the predicted membership scores.
Classification metrics, accuracy, AUC, and AUPR, is utilized to evaluate attack performances. We report here the MIA performances against some of the baseline generators, using the default parameter values provided in the config.yaml.
| Method | Accuracy | AUCROC | Average Precision | TPR@FPR=0.1 | TPR@FPR=0.01 |
|---|---|---|---|---|---|
| gan_leaks | 0.50 | 0.5252 | 0.5636 | 0.18 | 0.04. |