On the use of variational autoencoders for biomedical data integration
Abstract Variational Autoencoders (VAEs) are a widely used framework to integrate diverse biomedical data modalities, create representations that capture the underlying structure of the datasets, and obtain insights about the relations between variables. Here we describe how this is achieved from an empirical point of view in our previously developed VAE-based framework MOVE, providing an intuitive perspective on the inner workings of multimodal VAEs in biomedical contexts. We explore how the models’ emerging dynamics shape their performance and how in silico perturbations can be leveraged to identify potential associations between variables. To do that, we extend our framework to handle perturbations of continuous variables, introduce a new approach to better capture associations between them, and create synthetic datasets to benchmark the proposed methods against well-defined ground truth associations. We finally showcase our findings in real biomedical scenarios, namely a multimodal dataset of inflammatory bowel disease and a dataset containing genetic knockdowns in K562 and RPE1 cells.
Authors
- Henry Webel (ORCID: https://orcid.org/0000-0001-8833-7617)
- Simon Rasmussen (ORCID: https://orcid.org/0000-0001-6323-9041)
- Ricardo Hernández Medina (ORCID: https://orcid.org/0000-0001-6373-2362)
- Marc Pielies Avelli
Institutions
- University of Copenhagen (DK)
- Novo Nordisk Foundation (DK)
- Technical University of Denmark (DK)
Publication Details
- Journal
- BMC Artificial Intelligence
- Published
- 2026-09-01
- DOI
- https://doi.org/10.1186/s44398-026-00034-9
- Citations
- 1
- Primary Topic
- Machine Learning in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 6.40
Funders
- Københavns Universitet
- Novo Nordisk
- Novo Nordisk Fonden
- Novo Nordisk Foundation Center for Basic Metabolic Research