COVID Data for Shared Learning (CDSL): From Pandemic Response to a Lasting Research Resource

CDSL began as an emergency data-sharing initiative during the COVID-19 pandemic. Four years later, its evolution into a documented, validated and publicly accessible research resource shows the sustained work needed to make shared clinical data reliable and reusable over time.
COVID Data for Shared Learning (CDSL): From Pandemic Response to a Lasting Research Resource

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Sharing data when evidence was scarce

When COVID-19 reached hospitals in early 2020, clinicians and researchers were confronted with a new disease and limited evidence to guide decisions. At the same time, hospitals were generating large amounts of clinical data that could help researchers better understand the disease, provided those data could be shared rapidly and responsibly.

Against this backdrop, HM Hospitales launched COVID DATA SAVE LIVES, an initiative to make de-identified data from patients hospitalized with COVID-19 available to the international research community. The first version was released on 25 April 2020, when many publicly available COVID-19 datasets focused primarily on demographic or epidemiological information rather than combining detailed clinical records with medical imaging.

That initial release contained records from 2,310 hospitalization episodes, including admissions, diagnoses, treatments, intensive care, vital signs, laboratory results and medical imaging. Researchers submitted a request describing their proposed project, which was reviewed before access was granted. The dataset continued to grow as the pandemic evolved.

The data therefore came before the paper. As the resource evolved, the original files were restructured into a relational database, providing a more structured way to query, explore and analyse the clinical information. What began as an emergency data-sharing initiative would evolve over the following years into a more structured and documented research resource.

From early release to long-term reuse

Since its initial release, COVID Data for Shared Learning (CDSL) has grown into a multimodal, publicly accessible research resource comprising de-identified electronic health records (EHR) and chest imaging from 4,479 COVID-19 hospitalization episodes in Spain. It includes demographics, admissions, diagnoses, vital signs, laboratory results and treatments, together with 4,608 chest X-rays and more than 1.4 million CT slices.

Making the data available during the pandemic, however, was only the beginning. A dataset created under emergency conditions is not necessarily prepared for long-term reuse. Documentation, data processing and access mechanisms all need to remain understandable and sustainable. Four and a half years after the initial release, CDSL became available through PhysioNet in October 2024 [1], followed by its publication as a Data Descriptor in Scientific Data in September 2026 [2].

Although CDSL is publicly accessible through PhysioNet, access remains controlled to minimize potential re-identification risk: users must verify their identity and institutional affiliation, complete human-subjects research training, agree to the Data Use Agreement and submit their request for Contributor Review.

What happens behind a clinical dataset

Preparing CDSL for long-term sharing meant returning to almost every part of the dataset.

One of the most important tasks was de-identification. Structured identifiers were removed or replaced, dates were shifted while preserving the chronology of each hospitalization episode, and identifying metadata were removed from the original DICOM images. Free-text fields required particular attention: their content was translated from Spanish into English and manually reviewed both for semantic accuracy and for any residual identifying information.

Behind CDSL: the main steps used to transform structured EHR data and radiological images into a de-identified research resource.

The imaging data presented a different challenge: scale. More than 1.4 million CT slices, together with more than 4,600 chest X-rays, had to be processed. The original DICOM images were converted to JPG to reduce storage requirements and facilitate their use in standard image-processing and machine-learning workflows. Image conversion quality was assessed automatically, with samples also reviewed manually to verify image fidelity.

Working with real-world clinical data also meant dealing with heterogeneity. Records originated from different hospital information systems and had not been collected primarily for research. Rather than extensively transforming them into an artificially uniform dataset, we sought to preserve their correspondence with the original records while applying the changes needed for responsible sharing. Data completeness, plausible value ranges, duplicate records and linkage across tables were then assessed.

Finally, we wanted this process to be reproducible. Code for preprocessing, metadata extraction, image conversion and database deployment has been made publicly available, allowing researchers to understand how CDSL was constructed and reproduce key processing steps.

Why preserve a COVID-19 dataset now?

More than six years after its initial release, the value of CDSL extends beyond studying COVID-19 itself. It captures real-world hospital care during the early stages of an emerging disease and combines two types of information often studied separately: electronic health records and medical imaging. This creates opportunities for prognostic modelling, imaging analysis, multimodal machine learning and methodological research.

CDSL has already been reused in subsequent research, including studies on differentially private federated learning, multimodal predictive process monitoring and multimodal learning from irregularly sampled medical data [3-5]. These applications illustrate how a dataset originally shared in response to an immediate public-health need can later support new methodological questions and computational approaches.

Looking ahead

CDSL is now hosted on PhysioNet as a publicly accessible research resource, together with its documentation and accompanying source code. We hope it will continue to support research on multimodal clinical data and the development and evaluation of computational methods.

The journey of CDSL since its first release in April 2020 has reinforced a broader lesson for us: making data available is only the first step. Turning them into a lasting research resource requires a sustained commitment to privacy, documentation, validation and reproducibility. More than six years after its initial release, CDSL is designed to remain useful beyond the public-health emergency that led to its creation.

References

  1. Ritoré, Á., Oprescu, A. M., Estirado Bronchalo, A., & Armengol de la Hoz, M. Á. (2024). COVID Data for Shared Learning (CDSL): A comprehensive, multimodal COVID-19 dataset from HM Hospitales. PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/85z8-jq92
  2. Ritoré-Hidalgo, Á., Villar Fernández, J., Oprescu, A.M. et al. COVID Data for Shared Learning (CDSL): a multimodal publicly accessible dataset of electronic health records and chest imaging from hospitalized COVID-19 patients. Sci Data (2026). https://doi.org/10.1038/s41597-026-08184-1
  3. Nguyen, H.V.C., Nguyen, T.Q., Kravenkit, S. et al. Modality-aware differentially private federated learning for partially observed multimodal healthcare prediction. Sci Rep (2026). https://doi.org/10.1038/s41598-026-70939-y
  4. Pasquadibisceglie, V., Donadello, I., Appice, A., Lanz, O., Maggi, F. M., Fiameni, G., & Malerba, D. (2026). Multimodal predictive process monitoring and its application to explainable clinical pathways. Information Systems, 139, 102698. https://doi.org/10.1016/j.is.2026.102698
  5. Tölle, M., Scharaf, M., Fischer, S. et al. Visual prompt engineering for multimodal and irregularly sampled medical data. Commun Med 6, 474 (2026). https://doi.org/10.1038/s43856-026-01817-x

Follow the Topic

Public Health
Life Sciences > Health Sciences > Public Health
COVID19
Life Sciences > Health Sciences > Clinical Medicine > Diseases > Respiratory Tract Diseases > COVID19
Data Analysis and Big Data
Mathematics and Computing > Statistics > Data Analysis and Big Data
Machine Learning
Mathematics and Computing > Computer Science > Artificial Intelligence > Machine Learning
Medical Imaging
Life Sciences > Health Sciences > Health Care > Medical Physics > Medical Imaging
Artificial Intelligence
Mathematics and Computing > Computer Science > Artificial Intelligence

Related Collections

With Collections, you can get published faster and increase your visibility.

10 years of the FAIR principles

This Scientific Data Collection invites researchers to submit manuscripts related to FAIR-aligned infrastructure, policy, or standardisation.

Publishing Model: Open Access

Deadline: Nov 14, 2026

Data to support cancer treatment, diagnosis, and disease understanding

This Scientific Data collection welcomes descriptions of any dataset relevant to cancer research, including ’omics data; in vitro, in vivo, in silico, and human studies; early‑phase drug discovery data; as well as larger cohort and population‑level datasets related to measuring outcomes and disease occurrence at large scales.

Publishing Model: Open Access

Deadline: Jan 02, 2027