A Novel Taxonomy for Navigating and Classifying Synthetic Data in Healthcare Applications

Stud Health Technol Inform. 2024 Nov 22:321:259-263. doi: 10.3233/SHTI241104.

Abstract

Data-driven technologies have improved the efficiency, reliability and effectiveness of healthcare services, but come with an increasing demand for data, which is challenging due to privacy-related constraints on sharing data in healthcare contexts. Synthetic data has recently gained popularity as potential solution, but in the flurry of current research it can be hard to oversee its potential. This paper proposes a novel taxonomy of synthetic data in healthcare to navigate the landscape in terms of three main varieties. Data Proportion comprises different ratios of synthetic data in a dataset and associated pros and cons. Data Modality refers to the different data formats amenable to synthesis and format-specific challenges. Data Transformation concerns improving specific aspects of a dataset like its utility or privacy with synthetic data. Our taxonomy aims to help researchers in the healthcare domain interested in synthetic data to grasp what types of datasets, data modalities, and transformations are possible with synthetic data, and where the challenges and overlaps between the varieties lie.

Keywords: Digital healthcare; Healthcare applications; Synthetic data; Synthetic data classification.

MeSH terms

  • Classification*
  • Datasets as Topic
  • Humans