miRNA Data t-SNE Explore XGBoost Classify Boxplot Validate Result t-SNE output Bladder Ca Controls Feature importance v5582 v1234
Use Case AI Biomarker Discovery

AI Biomarker Discovery Workflow

Dr Craig Davison
Dr Craig Davison, Customer Success
13 min read · Updated July 2024

Data type

miRNA sequencing

Method

t-SNE + XGBoost + Boxplot

Platform

Sonrai Discovery

Key Takeaways

  • This end-to-end AI biomarker discovery workflow takes miRNA sequencing data from raw upload through t-SNE exploration, XGBoost classification, and boxplot validation - all in a no-code environment
  • Applied to bladder cancer vs non-cancer controls, the XGBoost model achieved sensitivity and specificity of 1 - validated by an associated peer-reviewed publication
  • The workflow is fully configurable for different data types and disease contexts - demonstrating how translational analytics can move from raw data to decision-ready evidence without specialist bioinformatics expertise

End-to-End AI Biomarker Discovery

In the following use case, we'll showcase an illustrative workflow employing Sonrai Discovery for comprehensive AI biomarker discovery. This particular segment is dedicated to the analysis phase - but please contact us for a complete overview encompassing data management, secure collaboration, and reporting findings.

Our objective is to establish a classifier for identifying bladder cancer disease status using MicroRNA data. This workflow demonstrates how the right analytical approach - applied to the right data - can generate robust, publication-supported biomarker findings without requiring specialist bioinformatics expertise.

Full walkthrough of the AI biomarker discovery workflow in Sonrai Discovery.

Build Your Storyboard

We will develop a tailored workflow for this use case, configuring custom parameters including inclusion/exclusion criteria and fine-tuning parameters for various applications. After uploading, curating, and merging our data into an AI-ready format, our first step is constructing a storyboard:

  1. Choose pie charts and histograms to visualise the data and select the correct subcohort
  2. Select a t-SNE to visualise complex data as a scatterplot
  3. Train an XGBoost model and use a boxplot to validate the features identified

These applications can function independently but are also adaptable for use in sequence, with data seamlessly transferred between them.

Building the storyboard in Sonrai Discovery - selecting visualisations and configuring the cohort composition view.

Filter for Bladder Cancer and Non-Cancer Controls

As we can see in the initial view, we're looking at more than just bladder cancer and non-cancer controls - there are other cancers within this cohort. We create a filter to ensure we're only looking at the comparison of interest.

Sonrai Discovery filter interface showing selection of bladder cancer and non-cancer controls from mixed cancer cohort

Creating a filter in Sonrai Discovery to isolate bladder cancer and non-cancer controls from the wider cohort - the storyboard updates instantly.

t-SNE: Setting Parameters

With the filtered data linked, we move to the t-SNE application to visualise this complex high-dimensional data as a scatter plot, with each point representing a patient. No additional filters are necessary - they carry forward automatically from the previous step.

We first calculate PCA to determine the number of principal components that account for variability between the disease status groups. This is a peer-reviewed approach to reduce dimensions prior to t-SNE, speeding up convergence and improving result quality.

PCA calculation in Sonrai Discovery showing principal components accounting for variability between bladder cancer and non-cancer control disease status groups

PCA calculation showing the number of principal components accounting for variability between disease status groups. Click to read the original van der Maaten & Hinton t-SNE paper.

Interactive t-SNE Graph

Once satisfied with the parameters, clicking Apply generates the interactive t-SNE graph. Modifying parameters is easy, with real-time updates to both the explained variance and KL divergence scores.

Interactive t-SNE scatter plot in Sonrai Discovery showing two distinct clusters for bladder cancer and non-cancer controls from miRNA sequencing data

Interactive t-SNE graph showing two clearly distinct clusters - one for bladder cancer, one for non-cancer controls. This separation gives confidence that the data is well-suited for training a robust classifier.

t-SNE parameter exploration in Sonrai Discovery showing changes in explained variance and KL divergence scores with 3D visualisation option

Parameter exploration with real-time updates to explained variance and KL divergence. A 3D visualisation option is also available for more complex datasets.

XGBoost Classification

The t-SNE analysis clearly reveals two distinct clusters, giving us confidence the data is well-suited for training a classifier. We move to the XGBoost model, setting disease status as our target variable and excluding unique IDs, which don't contribute to the analysis. The train/test split defaults to 60/40.

XGBoost model configuration and training in Sonrai Discovery. The blue testing line reaches 0 after only 13 rounds - parameters are adjusted and the model retrained accordingly.

After retraining with adjusted parameters, we examine the confusion matrix to evaluate the model's performance on both training and testing rounds.

XGBoost retrained model with adjusted parameters in Sonrai Discovery for bladder cancer miRNA classifier

Confusion Matrix: Sensitivity and Specificity

The model demonstrates exceptional sensitivity and specificity - both at a perfect score of 1. You would be correct to be sceptical of a model that performed so well. However, in this particular case, these are genuine outcomes from the dataset, corroborated by an associated peer-reviewed publication.

Confusion matrix from XGBoost bladder cancer classifier in Sonrai Discovery showing sensitivity and specificity of 1 on both training and testing sets

Feature Importance

Knowing the model performs well, we now want to see which features it is using. XGBoost provides feature importance plots, making this a grey-box rather than a black-box ML tool. This is critically important - it enables individual verification of the identified markers, transforming the model output into actionable biomarker candidates for further validation.

XGBoost feature importance plot in Sonrai Discovery showing v5582 as top miRNA feature for bladder cancer classification

Boxplot: Individually Verify the Features

We configure our boxplot with disease status on the X-axis and the top XGBoost feature on the Y-axis. For instance, selecting v5582 - the top feature - instantly generates a boxplot with statistics comparing bladder cancer and non-cancer controls.

Boxplot configuration in Sonrai Discovery showing feature v5582 distribution between bladder cancer and non-cancer control groups

Result

The boxplot makes it immediately obvious why v5582 emerged as the top feature in the XGBoost model - bladder cancer participants exhibit notably lower expression compared to non-cancer controls, with a clear, statistically significant separation between the two groups.

Boxplot result in Sonrai Discovery showing significantly lower v5582 miRNA expression in bladder cancer patients compared to non-cancer controls confirming top biomarker feature

v5582 shows significantly lower expression in bladder cancer participants compared to non-cancer controls - confirming its role as the model's top discriminating feature and a candidate biomarker for further validation.

As demonstrated, crafting a personalised AI biomarker discovery workflow, fine-tuning parameters, and training machine learning models are seamless tasks in Sonrai Discovery. This adaptable workflow concept extends to a wide array of applications across precision medicine data modalities. Once candidate biomarkers are identified, the next step is a rigorous biomarker validation strategy to ensure findings translate to real-world impact.

Translational Analytics

See how Sonrai applies this end-to-end AI biomarker discovery approach within a governed, reproducible research environment to generate decision-ready evidence.

Learn more

Related Reading

Talk to our team

Sonrai acts as an extension of your team, combining scientific expertise, transparent and reproducible analysis, and a governed research environment to generate decision-ready evidence that can be revisited, reproduced, and built upon throughout the life of a programme.

Book a demo