COVID Data for Shared Learning (CDSL): From Pandemic Response to a Lasting Research Resource
Published in Healthcare & Nursing, Computational Sciences, and General & Internal Medicine
Sharing data when evidence was scarce
When COVID-19 reached hospitals in early 2020, clinicians and researchers were confronted with a new disease and limited evidence to guide decisions. At the same time, hospitals were generating large amounts of clinical data that could help researchers better understand the disease, provided those data could be shared rapidly and responsibly.
Against this backdrop, HM Hospitales launched COVID DATA SAVE LIVES, an initiative to make de-identified data from patients hospitalized with COVID-19 available to the international research community. The first version was released on 25 April 2020, when many publicly available COVID-19 datasets focused primarily on demographic or epidemiological information rather than combining detailed clinical records with medical imaging.
That initial release contained records from 2,310 hospitalization episodes, including admissions, diagnoses, treatments, intensive care, vital signs, laboratory results and medical imaging. Researchers submitted a request describing their proposed project, which was reviewed before access was granted. The dataset continued to grow as the pandemic evolved.
The data therefore came before the paper. As the resource evolved, the original files were restructured into a relational database, providing a more structured way to query, explore and analyse the clinical information. What began as an emergency data-sharing initiative would evolve over the following years into a more structured and documented research resource.
From early release to long-term reuse
Since its initial release, COVID Data for Shared Learning (CDSL) has grown into a multimodal, publicly accessible research resource comprising de-identified electronic health records (EHR) and chest imaging from 4,479 COVID-19 hospitalization episodes in Spain. It includes demographics, admissions, diagnoses, vital signs, laboratory results and treatments, together with 4,608 chest X-rays and more than 1.4 million CT slices.
Making the data available during the pandemic, however, was only the beginning. A dataset created under emergency conditions is not necessarily prepared for long-term reuse. Documentation, data processing and access mechanisms all need to remain understandable and sustainable. Four and a half years after the initial release, CDSL became available through PhysioNet in October 2024 [1], followed by its publication as a Data Descriptor in Scientific Data in September 2026 [2].
Although CDSL is publicly accessible through PhysioNet, access remains controlled to minimize potential re-identification risk: users must verify their identity and institutional affiliation, complete human-subjects research training, agree to the Data Use Agreement and submit their request for Contributor Review.
What happens behind a clinical dataset
Preparing CDSL for long-term sharing meant returning to almost every part of the dataset.
One of the most important tasks was de-identification. Structured identifiers were removed or replaced, dates were shifted while preserving the chronology of each hospitalization episode, and identifying metadata were removed from the original DICOM images. Free-text fields required particular attention: their content was translated from Spanish into English and manually reviewed both for semantic accuracy and for any residual identifying information.
The imaging data presented a different challenge: scale. More than 1.4 million CT slices, together with more than 4,600 chest X-rays, had to be processed. The original DICOM images were converted to JPG to reduce storage requirements and facilitate their use in standard image-processing and machine-learning workflows. Image conversion quality was assessed automatically, with samples also reviewed manually to verify image fidelity.
Working with real-world clinical data also meant dealing with heterogeneity. Records originated from different hospital information systems and had not been collected primarily for research. Rather than extensively transforming them into an artificially uniform dataset, we sought to preserve their correspondence with the original records while applying the changes needed for responsible sharing. Data completeness, plausible value ranges, duplicate records and linkage across tables were then assessed.
Finally, we wanted this process to be reproducible. Code for preprocessing, metadata extraction, image conversion and database deployment has been made publicly available, allowing researchers to understand how CDSL was constructed and reproduce key processing steps.
Why preserve a COVID-19 dataset now?
More than six years after its initial release, the value of CDSL extends beyond studying COVID-19 itself. It captures real-world hospital care during the early stages of an emerging disease and combines two types of information often studied separately: electronic health records and medical imaging. This creates opportunities for prognostic modelling, imaging analysis, multimodal machine learning and methodological research.
CDSL has already been reused in subsequent research, including studies on differentially private federated learning, multimodal predictive process monitoring and multimodal learning from irregularly sampled medical data [3-5]. These applications illustrate how a dataset originally shared in response to an immediate public-health need can later support new methodological questions and computational approaches.
Looking ahead
CDSL is now hosted on PhysioNet as a publicly accessible research resource, together with its documentation and accompanying source code. We hope it will continue to support research on multimodal clinical data and the development and evaluation of computational methods.
The journey of CDSL since its first release in April 2020 has reinforced a broader lesson for us: making data available is only the first step. Turning them into a lasting research resource requires a sustained commitment to privacy, documentation, validation and reproducibility. More than six years after its initial release, CDSL is designed to remain useful beyond the public-health emergency that led to its creation.
References
- Ritoré, Á., Oprescu, A. M., Estirado Bronchalo, A., & Armengol de la Hoz, M. Á. (2024). COVID Data for Shared Learning (CDSL): A comprehensive, multimodal COVID-19 dataset from HM Hospitales. PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/85z8-jq92
- Ritoré-Hidalgo, Á., Villar Fernández, J., Oprescu, A.M. et al. COVID Data for Shared Learning (CDSL): a multimodal publicly accessible dataset of electronic health records and chest imaging from hospitalized COVID-19 patients. Sci Data (2026). https://doi.org/10.1038/s41597-026-08184-1
- Nguyen, H.V.C., Nguyen, T.Q., Kravenkit, S. et al. Modality-aware differentially private federated learning for partially observed multimodal healthcare prediction. Sci Rep (2026). https://doi.org/10.1038/s41598-026-70939-y
- Pasquadibisceglie, V., Donadello, I., Appice, A., Lanz, O., Maggi, F. M., Fiameni, G., & Malerba, D. (2026). Multimodal predictive process monitoring and its application to explainable clinical pathways. Information Systems, 139, 102698. https://doi.org/10.1016/j.is.2026.102698
- Tölle, M., Scharaf, M., Fischer, S. et al. Visual prompt engineering for multimodal and irregularly sampled medical data. Commun Med 6, 474 (2026). https://doi.org/10.1038/s43856-026-01817-x
Follow the Topic
-
Scientific Data
A peer-reviewed, open-access journal for descriptions of datasets, and research that advances the sharing and reuse of scientific data.
Related Collections
With Collections, you can get published faster and increase your visibility.
10 years of the FAIR principles
Publishing Model: Open Access
Deadline: Nov 14, 2026
Data to support cancer treatment, diagnosis, and disease understanding
Publishing Model: Open Access
Deadline: Jan 02, 2027