Unlocking Data-driven Solutions for Oil Wells Monitoring

In this post, we share the rationale behind the design and curation of the 3W Dataset, a resource that has contributed to the advancement of oil well monitoring in various ways.
Unlocking Data-driven Solutions for Oil Wells Monitoring
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Certain types of events that occur in offshore oil wells during the production phase can lead to undesirable consequences for production, facilities, the environment, and the workers involved. This scenario is illustrated in Fig. 1.

Fig. 1 Diagram representing the considered scenario when designing the 3W Dataset 2.0.0.

We believe that at least part of these undesirable consequences can be avoided or mitigated when such events are detected quickly. We also believe that the vast amount and diversity of time-series data generated and archived over years from oil well monitoring is a valuable resource that can enable faster detections (commonly referred to as early detections).

However, we observe several challenges for digital solutions focused on early detections to become sufficiently robust, efficient, scalable, and integrable into existing oil well monitoring systems.

Since the beginning of this work, we have seen as a central point behind these challenges the lack of a representative, high-quality dataset, labeled by human experts and made publicly available. In light of these challenges, we established guidelines under which the 1st version of the 3W Dataset was created, made publicly available and detailed in this scientific article throughout 2019. The main guidelines are as follows:

  1. Inclusion of different types of industrial variables (pressures, temperatures, flow rates, valve states, etc., measured at different points in production systems). This guideline enables, for example, a substantially larger number of experiments dedicated to hypothesis studies regarding cause-and-effect relationships among these types of industrial variables and the types of events of interest;
  2. Preservation of characteristics of real industrial data (missing values, miscalibrated sensors, frozen sensors, outliers, etc.). This guideline primarily enables experimentation and development of more robust and resilient techniques capable of handling such characteristics, which is essential for any digital solution intended to be integrated into existing monitoring systems;
  3. Data labeled by human experts. The intention behind this guideline is to leverage the highly valuable tacit knowledge of human experts, both theoretical and operational, to establish gold-standard detection examples from which computational models can be developed;
  4. Licensing and providing the labeled data in an appropriate format and openly on the Internet. This guideline fosters the strengthening of the ecosystem and the virtuous cycle around the 3W Dataset, in which all parties involved benefit. In summary, students and professionals are trained using this resource; methodologies are developed and proposed by researchers; commercial solutions are developed by established companies or startups; oil operators evaluate promising models and commercial solutions;
  5. Encouraging collaboration on a global scale. This guideline stems from our belief that collaboration between academia and industry at a global level is essential for significant, applied, and cumulative technological advances. In this context, parties act according to their own businesses (motivations), and results are reused by other parties as inputs for further results. One example of an aspect we expect to evolve in this scenario concerns rare types of events. That is, when multiple oil operators collaborate with their own inputs (labeled data) related to rare types of events, they become less rare, and the consolidated input then enables more possibilities and more benefits for everyone involved.

Since the release of the 1st version of the 3W Dataset, it has been explored in more and more academic works, scientific investigations, and developments of innovative commercial digital solutions, some purely data-driven and others hybrid with phenomenology-based components. We have followed this evolution and see it as an important evidence that the lack of a representative, high-quality dataset, labeled by human experts and made publicly available is indeed the central point behind the challenges for digital solutions focused on early detections to become sufficiently robust, efficient, scalable, and integrable into existing oil well monitoring systems.

In order to strengthen the ecosystem and the virtuous cycle around the 3W Dataset, the 2nd version of the 3W Dataset was created, made publicly available, and detailed in this data article. The main evolutions between these versions are as follows:

  1. Time series began to be stored in Parquet files (more modern and preferable compared to CSV files);
  2. Labels were reviewed by human experts;
  3. Additional data related to the following were incorporated:
    • 20 industrial variables (total increased to 27);
    • 1 label (total increased to 2);
    • 24 real wells (total increased to 42);
    • 1 event type (total increased to 10).

We see the 3W Dataset as a living resource that will be continuously evolved.

We hope that this data article will be an effective means of discovering the 3W Dataset and a detailed source of information about its most recent version. These will be permanent contributions to strengthening both the ecosystem and the virtuous cycle around the 3W Dataset.

It is important to note that the 3W Dataset is one of the main resources of the 3W Project, an even broader and more ambitious initiative by Petrobras (Brazilian state-owned oil company). Both are based on the principle that significant scientific progress is significantly accelerated when data and tools are made openly available.

We also hope that the 3W Dataset and the 3W Project inspire similar efforts in other domains in which data-driven solutions are limited due to the unavailability of suitable datasets.

Please sign in or register for FREE

If you are a registered user on Research Communities by Springer Nature, please sign in

Follow the Topic

Data Science
Mathematics and Computing > Computer Science > Artificial Intelligence > Data Science
Machine Learning
Mathematics and Computing > Statistics > Statistics and Computing > Machine Learning
Time Series Analysis
Mathematics and Computing > Mathematics > Analysis > Dynamical Systems > Time Series Analysis

Related Collections

With Collections, you can get published faster and increase your visibility.

Computer vision in plant science and agriculture

This Scientific Data Collection invites Data Descriptors documenting the generation, curation, and validation of datasets that underpin computer vision applications across plant biology, crop science, and agricultural systems.

Publishing Model: Open Access

Deadline: Oct 10, 2026

Datasets in education

This Scientific Data Collection invites Data Descriptors that describe the generation, curation, and validation of open datasets related to educational systems, practices, and outcomes across diverse contexts and populations.

Publishing Model: Open Access

Deadline: Nov 19, 2026