AI Biomarker Discovery Workflow
13 min read · Updated July 2024
Data type
miRNA sequencing
Method
t-SNE + XGBoost + Boxplot
Platform
Sonrai Discovery
Key Takeaways
- This end-to-end AI biomarker discovery workflow takes miRNA sequencing data from raw upload through t-SNE exploration, XGBoost classification, and boxplot validation - all in a no-code environment
- Applied to bladder cancer vs non-cancer controls, the XGBoost model achieved sensitivity and specificity of 1 - validated by an associated peer-reviewed publication
- The workflow is fully configurable for different data types and disease contexts - demonstrating how translational analytics can move from raw data to decision-ready evidence without specialist bioinformatics expertise
Jump to section
End-to-End AI Biomarker Discovery
In the following use case, we'll showcase an illustrative workflow employing Sonrai Discovery for comprehensive AI biomarker discovery. This particular segment is dedicated to the analysis phase - but please contact us for a complete overview encompassing data management, secure collaboration, and reporting findings.
Our objective is to establish a classifier for identifying bladder cancer disease status using MicroRNA data. This workflow demonstrates how the right analytical approach - applied to the right data - can generate robust, publication-supported biomarker findings without requiring specialist bioinformatics expertise.
Full walkthrough of the AI biomarker discovery workflow in Sonrai Discovery.
Build Your Storyboard
We will develop a tailored workflow for this use case, configuring custom parameters including inclusion/exclusion criteria and fine-tuning parameters for various applications. After uploading, curating, and merging our data into an AI-ready format, our first step is constructing a storyboard:
- Choose pie charts and histograms to visualise the data and select the correct subcohort
- Select a t-SNE to visualise complex data as a scatterplot
- Train an XGBoost model and use a boxplot to validate the features identified
These applications can function independently but are also adaptable for use in sequence, with data seamlessly transferred between them.
Building the storyboard in Sonrai Discovery - selecting visualisations and configuring the cohort composition view.
Filter for Bladder Cancer and Non-Cancer Controls
As we can see in the initial view, we're looking at more than just bladder cancer and non-cancer controls - there are other cancers within this cohort. We create a filter to ensure we're only looking at the comparison of interest.
Creating a filter in Sonrai Discovery to isolate bladder cancer and non-cancer controls from the wider cohort - the storyboard updates instantly.
t-SNE: Setting Parameters
With the filtered data linked, we move to the t-SNE application to visualise this complex high-dimensional data as a scatter plot, with each point representing a patient. No additional filters are necessary - they carry forward automatically from the previous step.
We first calculate PCA to determine the number of principal components that account for variability between the disease status groups. This is a peer-reviewed approach to reduce dimensions prior to t-SNE, speeding up convergence and improving result quality.
PCA calculation showing the number of principal components accounting for variability between disease status groups. Click to read the original van der Maaten & Hinton t-SNE paper.
Interactive t-SNE Graph
Once satisfied with the parameters, clicking Apply generates the interactive t-SNE graph. Modifying parameters is easy, with real-time updates to both the explained variance and KL divergence scores.
Interactive t-SNE graph showing two clearly distinct clusters - one for bladder cancer, one for non-cancer controls. This separation gives confidence that the data is well-suited for training a robust classifier.
Parameter exploration with real-time updates to explained variance and KL divergence. A 3D visualisation option is also available for more complex datasets.
XGBoost Classification
The t-SNE analysis clearly reveals two distinct clusters, giving us confidence the data is well-suited for training a classifier. We move to the XGBoost model, setting disease status as our target variable and excluding unique IDs, which don't contribute to the analysis. The train/test split defaults to 60/40.
XGBoost model configuration and training in Sonrai Discovery. The blue testing line reaches 0 after only 13 rounds - parameters are adjusted and the model retrained accordingly.
After retraining with adjusted parameters, we examine the confusion matrix to evaluate the model's performance on both training and testing rounds.
Confusion Matrix: Sensitivity and Specificity
The model demonstrates exceptional sensitivity and specificity - both at a perfect score of 1. You would be correct to be sceptical of a model that performed so well. However, in this particular case, these are genuine outcomes from the dataset, corroborated by an associated peer-reviewed publication.
Feature Importance
Knowing the model performs well, we now want to see which features it is using. XGBoost provides feature importance plots, making this a grey-box rather than a black-box ML tool. This is critically important - it enables individual verification of the identified markers, transforming the model output into actionable biomarker candidates for further validation.
Boxplot: Individually Verify the Features
We configure our boxplot with disease status on the X-axis and the top XGBoost feature on the Y-axis. For instance, selecting v5582 - the top feature - instantly generates a boxplot with statistics comparing bladder cancer and non-cancer controls.
Result
The boxplot makes it immediately obvious why v5582 emerged as the top feature in the XGBoost model - bladder cancer participants exhibit notably lower expression compared to non-cancer controls, with a clear, statistically significant separation between the two groups.
v5582 shows significantly lower expression in bladder cancer participants compared to non-cancer controls - confirming its role as the model's top discriminating feature and a candidate biomarker for further validation.
As demonstrated, crafting a personalised AI biomarker discovery workflow, fine-tuning parameters, and training machine learning models are seamless tasks in Sonrai Discovery. This adaptable workflow concept extends to a wide array of applications across precision medicine data modalities. Once candidate biomarkers are identified, the next step is a rigorous biomarker validation strategy to ensure findings translate to real-world impact.
Translational Analytics
See how Sonrai applies this end-to-end AI biomarker discovery approach within a governed, reproducible research environment to generate decision-ready evidence.
Related Reading
Case Study
Predictive Biomarker Discovery
Case Study
GenoMe Biomarker Discovery
Article
A Guide to Biomarker Validation
Talk to our team
Sonrai acts as an extension of your team, combining scientific expertise, transparent and reproducible analysis, and a governed research environment to generate decision-ready evidence that can be revisited, reproduced, and built upon throughout the life of a programme.
Book a demo