
Why Biomarker Studies Need More Than Feature Ranking
High-dimensional omics data can produce attractive signatures that do not survive a new cohort. An integrated study is needed to separate repeatable biological signal from batch effects, cohort imbalance, overfitting, and chance correlations.
Study design determines what a model can legitimately answer. Wet-lab quality determines whether the intended signal reaches the data. Bioinformatics turns raw measurements into comparable features. Machine learning can then help prioritize combinations and quantify their robustness—but it cannot repair an unsuitable cohort or replace independent evidence.
We therefore begin with the intended decision: early discovery, treatment-response research, molecular stratification, longitudinal risk exploration, or validation of an existing hypothesis. That decision guides group definitions, molecular layers, sample allocation, locked evaluation strategy, panel size, and follow-up assay.
Questions the study should answer
- Is the signal stronger than a simple baseline?
- Is it stable across resampling and preprocessing?
- Could technical or clinical variables explain it?
- Can the shortlist be measured in new samples?
- What evidence is still missing?
Concrete Research Scenarios
The following examples show how experiments, analysis, and validation can be connected. Cohort sizes and methods are illustrative and are refined after feasibility review.
Blood-Based Biomarker Discovery for Earlier Disease Research
Client question: Can a compact blood-based molecular panel distinguish the target condition from both healthy and clinically relevant benign controls?
Illustrative design: Use serum or plasma from a discovery cohort containing target cases, benign-disease controls, and healthy controls. Generate discovery proteomic, metabolomic, cfDNA, or transcriptomic profiles; retain a locked subset or independent cohort for assessment. Relevant covariates may include age, sex, collection site, fasting status, hemolysis, storage time, and processing batch.
Machine-learning role: Compare a conventional baseline with regularized and ensemble approaches; perform feature selection inside resampling; report stability, ROC and precision-recall performance, calibration where relevant, and the incremental value of a compact panel. Shortlisted proteins or molecular features can move to targeted measurement in retained samples.
What the client receives: Evidence-ranked candidates, a compact panel proposal, model and stability report, confounder analysis, targeted-validation plan or results, and a go/refine/stop recommendation.
Research basis: Ney et al. used 539 serum samples, diagnosis-specific base learners, feature elimination, and held-out assessment to develop a pancreatic cancer protein panel. Xing et al. demonstrated a staged discovery-verification-validation proteomics workflow across 1,002 participants. These studies support the workflow logic, not guaranteed performance in another cohort.
Pre-Treatment Biomarkers of Research Response
Client question: Which baseline molecular features distinguish responders from non-responders, and does multi-omics information improve on clinical variables alone?
Illustrative design: Collect pre-treatment biopsies and a clearly defined response endpoint. Depending on mechanism and sample availability, combine variant profiling, expression, immune-state, epigenetic, or pathology-derived features. Reserve evaluation data by patient, site, or study cohort so that related samples cannot cross partitions.
Machine-learning role: Compare clinical-only, single-omics, and integrated models under the same evaluation plan; review class imbalance, treatment arm, tissue composition, and site effects; identify stable features that add measurable value beyond the baseline.
Validation route: Test the locked signature in an external dataset or follow-up cohort, then use targeted sequencing or expression measurement to verify a manageable candidate set.
Research basis: Sammut et al. integrated clinical, digital pathology, genomic, and transcriptomic features from pre-treatment breast tumor biopsies and evaluated the response model in an external cohort, illustrating why the entire tumor ecosystem and an independent dataset matter.
Single-Cell Discovery of Cell-State Biomarkers
Client question: Is the apparent bulk-tissue signal driven by a specific immune, stromal, or disease-associated cell population?
Illustrative design: Profile carefully balanced tissue or peripheral-blood samples at single-cell resolution. Define the donor—not the cell—as the independent sample unit. Discover cell populations, state programs, and abundance shifts associated with the endpoint, while controlling donor and processing effects.
Machine-learning role: Prioritize reproducible cell-state features, compare donor-level signatures, and avoid inflated performance caused by randomly splitting cells from the same donor. Candidate populations can be verified with flow-based, targeted expression, or orthogonal tissue assays in additional samples.
Research basis: A published melanoma study used single-cell RNA sequencing for discovery and then flow-based analysis in additional samples to investigate S100A9-positive monocytes as a response-associated biomarker. The design demonstrates the value of moving from high-dimensional discovery to a practical orthogonal assay.
Multi-Omics Stratification and Risk Research
Client question: Do genomic, transcriptomic, epigenomic, protein, or metabolite layers define reproducible subgroups or improve a risk model beyond established variables?
Illustrative design: Harmonize sample identity and metadata across layers, define which samples have complete or partial data, and establish whether integration is early, intermediate, or late. Evaluate each molecular layer separately before testing the combined model.
Machine-learning role: Identify stable latent factors or cross-layer features, compare clinical-only and single-layer baselines with the integrated model, perform sensitivity analyses for missing layers and batch effects, and translate the result into an interpretable panel or subgroup definition.
Validation route: Reproduce the subgroup or risk association in a separate cohort and verify key features with a smaller targeted assay. The objective is not to maximize the number of omics layers but to show which layer changes the research decision.
Research basis: Hoadley et al. integrated aneuploidy, DNA methylation, mRNA, microRNA, and protein measurements across approximately 10,000 tumors and showed how individual and integrated molecular layers reveal both shared and tissue-of-origin patterns. The study supports evaluating each layer before interpreting an integrated subtype.
Technology and Service Options
The technology is selected around the biological question, sample type, expected signal, and downstream validation route—not around algorithm novelty.
| Service Technology | What It Contributes | Common Biomarker Use | Validation Consideration |
|---|---|---|---|
| RNA Sequencing | Gene, transcript, splice, and pathway-level expression features | Response signatures, molecular subgroups, disease-state programs | Targeted expression measurement or independent transcriptomic cohort |
| Single-Cell RNA Sequencing | Cell populations, cell states, and donor-level cellular composition | Rare-cell and immune-state biomarker discovery | Flow-based or targeted expression verification in additional donors |
| ATAC-Seq | Chromatin accessibility and regulatory-state features | Regulatory biomarkers and mechanism-linked signatures | Targeted regulatory-region or expression follow-up |
| Whole-Exome Sequencing | Coding variants, mutational patterns, and copy-number-related features | Variant-informed stratification and response studies | Orthogonal variant confirmation and independent cohort assessment |
| Multi-Omics Services | Integrated evidence across molecular layers | Cross-layer subtyping, response, and risk research | Show incremental value over clinical and single-layer baselines |
| Bioinformatics Services | Data review, processing, harmonization, modeling, and reproducible reporting | Data-to-Insight and hybrid projects | Claims depend on metadata, quality, and available validation data |
Proteomics, metabolomics, targeted measurement, spatial profiling, and other project-specific assays may also be incorporated after feasibility review.
Project Entry Modes
| Entry Mode | Starting Material | Integrated Scope |
|---|---|---|
| Sample-to-Insight | Biospecimens plus research question and metadata | Study-design support, omics experiments, QC, bioinformatics, machine-learning-assisted discovery, and validation planning |
| Data-to-Insight | Raw or processed omics data and metadata | Data audit, harmonization, confounder review, feature engineering, model comparison, interpretation, and validation strategy |
| Hybrid Study | Existing data plus samples for a missing layer or follow-up cohort | Targeted new experiments integrated with client data, followed by discovery and research validation |
Integrated Biomarker Discovery Workflow
One coordinated workflow connects the intended research decision to cohort design, wet-lab execution, modeling, and validation.

Step 1 — Define the decision: Specify the comparison, endpoint, intended sample unit, candidate-panel constraints, and evidence needed for advancement.
Step 2 — Design the cohort: Review inclusion criteria, balance, covariates, batches, paired or longitudinal structure, and options for held-out or independent assessment.
Step 3 — Generate fit-for-purpose omics data: Select and execute the wet-lab assay or assay combination, with quality thresholds linked to downstream modeling needs.
Step 4 — Build analysis-ready features: Process raw data, annotate features, review missingness and technical variation, and lock outcome-independent preprocessing.
Step 5 — Compare baselines and models: Evaluate conventional statistics and justified machine-learning approaches under the same nested or held-out plan.
Step 6 — Test robustness and interpret biology: Quantify feature stability, review confounders, assess calibration where relevant, and connect candidates to pathways or cell context.
Step 7 — Validate the research finding: Evaluate the locked signature in independent data and/or verify candidates with a targeted orthogonal assay.
Step 8 — Deliver the decision package: Report the candidate panel, supporting evidence, limitations, reproducible outputs, and recommended next experiments.
Machine Learning with Research Safeguards
Machine learning is an assistive research layer, not an automatic answer. Depending on the endpoint and sample structure, methods may include regularized linear models, tree-based models, kernel approaches, ensemble learning, or justified integration methods. Every complex model should be compared with a simpler, interpretable baseline.
- Feature processing inside the resampling loop
- Nested cross-validation for tuning and internal estimation
- Grouped or site-aware splitting when samples are related
- Training-only handling of class imbalance
- Feature-selection frequency and stability reporting
- Confounder and sensitivity analyses
- Calibration assessment when probabilities matter
- Locked independent-cohort evaluation when feasible
A model is not advanced on one metric alone.
Discrimination, error profile, calibration, stability, biological interpretation, and practical assay feasibility are reviewed together.
Study Inputs and Sample Considerations
There is no universal minimum cohort size. Feasibility depends on endpoint prevalence, heterogeneity, effect size, feature dimensionality, cohort structure, missingness, and the validation claim.
- Sample type, preservation, input quantity, and anticipated quality
- Endpoint definition, group balance, time points, and paired measurements
- Age, sex, treatment, site, batch, and other relevant covariates
- Raw-data availability and metadata completeness
- Discovery, tuning, held-out, and independent-cohort options
- Candidate-panel size and downstream assay constraints
- Data-transfer, privacy, and reproducibility requirements
Deliverables
- Experimental QC and omics data reports
- Analysis-ready matrices and metadata review
- Locked splitting and validation plan
- Baseline and machine-learning comparisons
- Cross-validation and held-out summaries
- Discrimination, error, and calibration outputs
- Feature stability and selection-frequency report
- Evidence-tiered candidate biomarker panel
- Biological annotation and pathway context
- Targeted-validation results or recommendations
- Reproducible outputs and methods-ready report
- Limitations and recommended next experiments
References
- Walsh I, Fishman D, Garcia-Gasulla D, et al. DOME: recommendations for supervised machine learning validation in biology. Nature Methods. 2021.
- Diaz-Uriarte R, Gómez de Lope E, Giugno R, et al. Ten quick tips for biomarker discovery and validation analyses using machine learning. PLOS Computational Biology. 2022.
- Ney A, Nené NR, Sedlak E, et al. Identification of a serum proteomic biomarker panel using diagnosis specific ensemble learning and symptoms for early pancreatic cancer detection. PLOS Computational Biology. 2024.
- Xing X, et al. Proteomics-driven noninvasive screening of circulating serum protein panels for the early diagnosis of hepatocellular carcinoma. Nature Communications. 2023.
- Sammut SJ, Crispin-Ortuzar M, Chin SF, et al. Multi-omic machine learning predictor of breast cancer therapy response. Nature. 2022.
- Rad Pour S, Pico de Coaña Y, Martinez Demorentin X, et al. Predicting anti-PD-1 responders in malignant melanoma from the frequency of S100A9-positive monocytes in the blood. Journal for ImmunoTherapy of Cancer. 2021.
- Hoadley KA, Yau C, Hinoue T, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell. 2018.
For Research Use Only. Not for use in diagnostic or clinical procedures.
Example Decision-Oriented Output
A typical report combines performance, stability, calibration, and biological interpretation rather than presenting one accuracy number in isolation.

Illustrative composite output: nested assessment distributions, held-out ROC and precision-recall curves, calibration, feature-selection frequency, and an evidence-tiered compact panel. Final metrics and plots depend on the study endpoint and design.
Machine Learning Biomarker Discovery FAQs
1. Can the project start from biospecimens?
Yes. Sample-to-Insight projects may include study-design support, omics experiments, quality control, bioinformatics, machine-learning-assisted discovery, and validation planning. The assay scope is selected after sample and endpoint review.
2. Can you analyze data generated elsewhere?
Yes. We first review raw or processed files, metadata, quality information, cohort structure, and compatibility with the intended claim. Missing metadata or major batch differences may narrow the feasible analysis.
3. Which omics layer should we choose?
The choice depends on mechanism, sample accessibility, expected abundance, relationship to the endpoint, budget, and validation route. We prioritize the smallest defensible design rather than adding layers without a decision purpose.
4. Do you always use deep learning?
No. Regularized or tree-based models are often more appropriate for limited omics cohorts and easier to interpret. Method choice follows data size, endpoint, structure, and validation needs.
5. How do you reduce overfitting?
The plan may use nested cross-validation, grouped splits, locked held-out data, preprocessing inside resampling, stability analysis, and independent assessment. The design is agreed before model tuning.
6. Can selected biomarkers be validated experimentally?
Potentially. Follow-up may use targeted sequencing, targeted expression, flow-based measurement, or another orthogonal assay in retained or independent research samples, subject to feasibility.
7. Can performance be guaranteed?
No. Biomarker discovery is uncertain. The service reduces avoidable bias, quantifies robustness, and helps determine whether a candidate should advance, be refined, or be discontinued.
Published Case Study
Independent Research Highlight
Serum Proteomic Biomarker Panel Using Diagnosis-Specific Ensemble Learning
This independent publication is presented as a study-design example and is not a CD Genomics customer project.
Research question
Could a compact serum protein signature distinguish pancreatic cancer from healthy individuals and clinically relevant benign conditions among people with concerning symptoms?
Study design
The researchers analyzed 539 serum samples using an oncology protein panel plus additional markers. Sixteen specialized base learners were combined in a stacked ensemble. Feature elimination and cross-validation were used during development, while a held-out set was reserved for assessment.
Reported result
In the held-out validation set, the ensemble achieved an area under the ROC curve of 0.95 (95% confidence interval 0.91–0.99) and sensitivity of 0.86 (95% confidence interval 0.68–1.00) at 90% specificity.

Transferable lesson
The same performance should not be expected in another cohort. The useful lesson is the connected workflow: clinically relevant controls, broad protein measurement, feature reduction inside model development, complementary learners, held-out evaluation, and a compact panel suitable for further verification.
Reference
- Ney A, Nené NR, Sedlak E, et al. Identification of a serum proteomic biomarker panel using diagnosis specific ensemble learning and symptoms for early pancreatic cancer detection. PLOS Computational Biology. 2024.
Selected Publications
These customer-related publications illustrate omics datasets and research questions relevant to biomarker studies. Inclusion does not imply that each paper used the complete solution described here.
- Iparraguirre L, Alberro A, Iñiguez SG, et al. Blood RNA-Seq profiling reveals a set of circular RNAs differentially expressed in frail individuals. Immunity & Ageing. 2023.
- Van Goubergen J, Peřina M, Handle F, et al. Targeting the CLK2/SRSF9 splicing axis in prostate cancer leads to decreased ARV7 expression. Molecular Oncology. 2025.
- Joruiz SM, Von Muhlinen N, Horikawa I, et al. Distinct functions of wild-type and R273H mutant Δ133p53α differentially regulate glioblastoma aggressiveness and therapy-induced senescence. Cell Death & Disease. 2024.
