Coral supports cancer researchers and computational biologists in interactively creating, refining, comparing, and characterizing patient cohorts from large multi-omic cancer datasets. The tool tracks the full cohort evolution as a provenance graph, enabling analysts to iteratively stratify patients based on metadata attributes and genomic markers, compare cohort attributes for significant differences, assess prevalences, and drill down to individual samples.
The input consists of a large cancer genomics database (over 128,000 samples from AACR Project GENIE, TCGA, and CCLE) containing mutation data, mRNA expression, DNA copy number, and clinical/demographic metadata per sample. Upon loading a dataset, the system automatically creates a root cohort containing all items and displays it as the initial node in the Cohort Evolution View.
The Cohort Evolution View (upper panel) displays all cohorts and operations as a directed graph showing how cohorts were derived from one another. The analyst selects one or more cohorts (each assigned a unique colour) to load them into the Action View (lower panel), which shows the Input Area with attribute distributions for context. This provides an immediate overview of the selected cohorts' composition.
The analyst enters an iterative cycle of cohort operations. They inspect attribute distributions (View operation) to understand how values are distributed across cohorts. When they identify relevant subgroups, they apply Filter (select items matching specific attribute values) or Split (divide a cohort into multiple sub-cohorts by attribute categories). Each operation creates new cohort nodes in the evolution graph connected by edges to their parents. The analyst can also apply the Compare operation (statistical tests for pairwise differences between cohorts), assess Prevalence (proportion of items with a characteristic within a reference population), or Inspect Items (drill down to individual samples in a Taggle table view). After each operation, the evolution graph grows, and the analyst assesses whether further stratification, comparison, or investigation is needed. This cycle repeats — creating, viewing, splitting, comparing, and refining cohorts — until the analyst has identified the relevant patient subgroups and their distinguishing characteristics.
The analyst synthesises findings from the cohort analysis: identified patient stratifications, confirmed or discovered biomarker associations (e.g., KRASG12C mutation prevalence across demographic subgroups), statistically significant differences between cohorts, and prevalence estimates. The integrated session management allows the analyst to store, share, and reproduce their findings.