From the Research Question to an Experimental Design

Week 2 focused on turning the research question into a testable experimental design. I examined longitudinal trajectories, missing information, modality reliability, and model comparisons while keeping the main hypothesis open to testing.

Published in Neuroscience

From the Research Question to an Experimental Design

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

During the second week, I focused less on adding new features and more on deciding how the main idea could actually be tested. Once I looked more closely at the structure of the project, it became clear that the different types of information cannot simply be treated as interchangeable. Blood biomarkers, MRI measurements, cognitive variables, genetic information, and demographic variables describe different aspects of a participant, and some of them change over time while others mainly describe baseline characteristics.

One decision that became clearer this week was the separation between information that changes over time and information that represents baseline susceptibility. Blood, MRI, and cognitive measurements are being treated as time-varying information, while genetic variables are kept as baseline information rather than being represented as trajectories. This distinction is already reflected in the implementation.

I also spent more time thinking about what “change” should actually mean. A single measurement that is outside an expected range may be useful, but it does not necessarily describe the same situation as a measurement that has gradually moved away from a person's own baseline. Because of this, the current analysis looks at several aspects of a participant's history, including deviation from baseline, slope, acceleration, persistence, and possible change points. These measurements are calculated separately before being combined at the modality level.

This helped narrow the main research question. Instead of asking only whether combining several modalities improves prediction, I am interested in whether a model performs differently when it knows which information is available, how complete that information is, and how the measurements have behaved over time.

The current fusion approach gives me a way to investigate this. Each modality receives a signal based on several characteristics of its trajectory. Its reliability is also influenced by how much information is available. The model then considers whether signals from different modalities show temporal agreement before calculating their combined contribution.

I do not want to interpret these weights as proof that one modality is objectively more important than another. At this stage, they are part of the proposed method. Their usefulness still needs to be demonstrated through experiments. A complicated weighting mechanism is not automatically better simply because it uses more information.

Missing information became another important part of the research design this week. In longitudinal data, missingness does not necessarily mean that an entire participant is absent. Someone could have several blood measurements but only a small number of MRI observations, while cognitive assessments might occur at different intervals. The current synthetic data generation reflects this by introducing missing observations separately across the different modalities.

This creates another question that can be tested directly: how does the model behave when one source of information becomes incomplete or unavailable? This may be more informative than looking only at performance when every modality is present.

I also clarified the purpose of the synthetic cohort. The current implementation creates 2,000 synthetic participants with repeated visits and several different trajectory patterns. These include stable trajectories as well as patterns in which blood, MRI, or cognitive information becomes informative at different times. There are also accelerating, delayed, discordant, multimodal, and transient patterns.

These patterns make it possible to test whether the computational framework responds differently when the timing and strength of signals change. However, they cannot be treated as evidence about real patients. The synthetic cohort was created to produce heterogeneous longitudinal trajectories for experimentation rather than to reproduce an actual clinical database.

That distinction is important for the rest of the project. A successful experiment on synthetic data would show that the proposed method behaves in a particular controlled environment. It would not establish clinical usefulness. Any later clinical interpretation would require appropriate longitudinal human datasets, patient-level separation, temporal validation, calibration, subgroup analysis, and independent validation.

Another part of this week's work was making the baseline comparisons more explicit. I do not want the final approach to be compared only against a weak or overly simple alternative. The planned comparisons move from single-modality models and conventional baselines toward static multimodal approaches, temporal approaches, adaptive fusion, and finally the complete proposed system. This makes it possible to investigate where any observed improvement actually comes from.

For example, if the complete model performs better, that result alone would not explain whether the improvement came from temporal information, adaptive weighting, missingness handling, reliability estimation, or another component. This is why ablation experiments are becoming an important part of the design. Individual components can be removed and evaluated separately rather than being assumed to contribute simply because they are included in the architecture.

I also paid more attention to how uncertainty should be represented. The current implementation produces both a score and an uncertainty estimate. The uncertainty increases when the available information is incomplete or when there is not enough longitudinal information. This is useful for the research question because a prediction based on limited observations should not automatically appear as reliable as one supported by a much richer history.

By the end of Week 2, the project had moved from a general idea about multimodal Alzheimer’s modelling toward a more defined experimental framework. There is now a clearer distinction between baseline information and changing measurements, a representation of individual trajectories, an adaptive way of combining modalities, explicit treatment of missing information, uncertainty estimation, and a set of comparisons that can be used to test the individual parts of the approach.

Most importantly, I am still treating the central idea as a hypothesis rather than a conclusion. The proposed approach may perform better than the alternatives, or some of its components may turn out not to provide a meaningful advantage. The experiments need to determine that.

For Week 3, the next step will be to establish the evaluation procedure in more detail. This means defining exactly what is being predicted, how participants are separated between training and evaluation, which metrics will be used, and how repeated experiments will be handled so that the conclusions do not depend on a single favourable run.

Follow the Topic

Neuroscience
Life Sciences > Biological Sciences > Neuroscience