Pesquisa | Portal Regional da BVS

Scalable and Privacy-Preserving Federated Principal Component Analysis.

Froelicher, David; Cho, Hyunghoon; Edupalli, Manaswitha; Sousa, Joao Sa; Bossuat, Jean-Philippe; Pyrgelis, Apostolos; Troncoso-Pastoriza, Juan R; Berger, Bonnie; Hubaux, Jean-Pierre.

Proc IEEE Symp Secur Priv ; 2023: 1908-1925, 2023 May.

Artigo em Inglês | MEDLINE | ID: mdl-38665901

RESUMO

Principal component analysis (PCA) is an essential algorithm for dimensionality reduction in many data science domains. We address the problem of performing a federated PCA on private data distributed among multiple data providers while ensuring data confidentiality. Our solution, SF-PCA, is an end-to-end secure system that preserves the confidentiality of both the original data and all intermediate results in a passive-adversary model with up to all-but-one colluding parties. SF-PCA jointly leverages multiparty homomorphic encryption, interactive protocols, and edge computing to efficiently interleave computations on local cleartext data with operations on collectively encrypted data. SF-PCA obtains results as accurate as non-secure centralized solutions, independently of the data distribution among the parties. It scales linearly or better with the dataset dimensions and with the number of data providers. SF-PCA is more precise than existing approaches that approximate the solution by combining local analysis results, and between 3x and 250x faster than privacy-preserving alternatives based solely on secure multiparty computation or homomorphic encryption. Our work demonstrates the practical applicability of secure and federated PCA on private distributed datasets.

Privacy-Preserving and Efficient Verification of the Outcome in Genome-Wide Association Studies.

Halimi, Anisa; Dervishi, Leonard; Ayday, Erman; Pyrgelis, Apostolos; Troncoso-Pastoriza, Juan Ramón; Hubaux, Jean-Pierre; Jiang, Xiaoqian; Vaidya, Jaideep.

Proc Priv Enhanc Technol ; 2022(3): 732-753, 2022.

Artigo em Inglês | MEDLINE | ID: mdl-36212774

RESUMO

Providing provenance in scientific workflows is essential for reproducibility and auditability purposes. In this work, we propose a framework that verifies the correctness of the aggregate statistics obtained as a result of a genome-wide association study (GWAS) conducted by a researcher while protecting individuals' privacy in the researcher's dataset. In GWAS, the goal of the researcher is to identify highly associated point mutations (variants) with a given phenotype. The researcher publishes the workflow of the conducted study, its output, and associated metadata. They keep the research dataset private while providing, as part of the metadata, a partial noisy dataset (that achieves local differential privacy). To check the correctness of the workflow output, a verifier makes use of the workflow, its metadata, and results of another GWAS (conducted using publicly available datasets) to distinguish between correct statistics and incorrect ones. For evaluation, we use real genomic data and show that the correctness of the workflow output can be verified with high accuracy even when the aggregate statistics of a small number of variants are provided. We also quantify the privacy leakage due to the provided workflow and its associated metadata and show that the additional privacy risk due to the provided metadata does not increase the existing privacy risk due to sharing of the research results. Thus, our results show that the workflow output (i.e., research results) can be verified with high confidence in a privacy-preserving way. We believe that this work will be a valuable step towards providing provenance in a privacy-preserving way while providing guarantees to the users about the correctness of the results.

RESUMO

RESUMO

ENVIAR RESULTADO:

SELEÇÃO DE REFERÊNCIAS

DETALHE DA PESQUISA