Synthetic Data: Can We Trust Statistical Estimators?

Decruyenaere, Alexander; Dehaene, Heidelinde; Rabaey, Paloma; Polet, Christiaan; Decruyenaere, Johan; Vansteelandt, Stijn; Demeester, Thomas

Synthetic Data: Can We Trust Statistical Estimators?

Alexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Christiaan Polet, Johan Decruyenaere, Stijn Vansteelandt, Thomas Demeester

NeurIPSW 2023

/neuripsw/2023/decruyenaere2023neuripsw-synthetic/

Abstract

The increasing interest in data sharing makes synthetic data appealing. However, the analysis of synthetic data raises a unique set of methodological challenges. In this work, we highlight the importance of inferential utility and provide empirical evidence against naive inference from synthetic data (that handles these as if they were really observed). We argue that the rate of false-positive findings (type 1 error) will be unacceptably high, even when the estimates are unbiased. One of the reasons is the underestimation of the true standard error, which may even progressively increase with larger sample sizes due to slower convergence. This is especially problematic for deep generative models. Before publishing synthetic data, it is essential to develop statistical inference tools for such data.

PDF NeurIPSW OpenReview Semantic Scholar

Cite

Text

Decruyenaere et al. "Synthetic Data: Can We Trust Statistical Estimators?." NeurIPS 2023 Workshops: DGM4H, 2023.

Markdown

[Decruyenaere et al. "Synthetic Data: Can We Trust Statistical Estimators?." NeurIPS 2023 Workshops: DGM4H, 2023.](https://mlanthology.org/neuripsw/2023/decruyenaere2023neuripsw-synthetic/)

BibTeX

@inproceedings{decruyenaere2023neuripsw-synthetic,
  title     = {{Synthetic Data: Can We Trust Statistical Estimators?}},
  author    = {Decruyenaere, Alexander and Dehaene, Heidelinde and Rabaey, Paloma and Polet, Christiaan and Decruyenaere, Johan and Vansteelandt, Stijn and Demeester, Thomas},
  booktitle = {NeurIPS 2023 Workshops: DGM4H},
  year      = {2023},
  url       = {https://mlanthology.org/neuripsw/2023/decruyenaere2023neuripsw-synthetic/}
}