Artificial Intelligence in Oncology - Supporting scientific research
Radboudumc
In short, Marianne Jonkers' research entails the following:
'The prognosis of patients with rare cancers is worse than that of patients with common cancers. An important reason is that for rare cancers it is more difficult to identify potential treatments, because only small data sets will typically be available for research. When using such small data sets to identify predictive factors, the risk of overfitting in the models used is high. This limits severely our ability to infer statistically significant patterns in these data, which might have suggested novel treatments. Pooling data from different institutions could alleviate the situation, but is in practice challenging due to regulatory and logistical problems.
Two complementary routes for confronting the challenges of small data sets for rare cancers are proposed. The first is to focus on more powerful techniques for inference that are better able to cope with small sample sizes, without overfitting. The second route is to design and improve machine learning algorithms that circumvent the need for data pooling at one location by `cycling’ around medical institutes with small data sets (federated learning). Data on salivary gland cancer will be analyzed with the proposed methods.
The project is a collaboration between: dr. Marianne Jonker, prof. Ton Coolen, prof. Kit Roes and prof. Carla van Herpen, all at the Radboud university medical center or Radboud university.'
Clinical data sets for rare cancers are usually small. This makes the accurate identification of prognostic factors for the disease course of patients with rare cancers challenging, and clinical outcome prediction for future patients with rare cancers unreliable. Researchers often try to enlarge their data sets by combining multiple data sets collected at different medical centers. However, merging data sets is frequently challenging or infeasible due to legal, ethical, or logistical barriers such as privacy regulations.
In this project, we have developed the Bayesian Federated Inference (BFI) methodology and supporting software, which enables researchers to reconstruct from local inferences made on separate data sets (e.g., across different medical centers), what would have been inferred if the data sets had been combined. This approach has proven to be highly accurate and broadly applicable. Extensive data analyses have been conducted, yielding excellent results. We are currently extending the methodology to Bayesian Federated Causal Inference in order to answer causal questions based on observational data. The latter is extremely important for rare cancer research as it is often challenging to perform clinical trials as the number of patients is limited. We expect to complete this within the next few months.
With the BFI methodology one can reconstruct from local inferences in separate data sets (centers) what would have been inferred had the data sets been merged. The methodology turns out to be very accurate and applicable in many situations. We have developed the methodology and software for multiple commonly used models (generalized linear models and time-to-event models), homogeneous and heterogeneous populations, and for prediction and association models. Data have been analyzed and the results are very good. The next step is to incorporate overfitting theory that overcomes overfitting in case of small datasets.
The first aim in this project is to develop and apply the theory of Bayesian Federated Inference (BFI) for parametric models and for models with a time-to-event outcome. We finished this for general parametric models and submitted our paper: Bayesian Federated Inference for Statistical Models (M.A. Jonker, H. Pazira, A.C.C. Coolen). The paper is still under review. The next step is to generalize the theory to the semi- parametric Cox PH model. We aim to finish this before the summer of 2023