AI-Assisted Molecular Subtyping Service

A heatmap can show groups in a dataset. It does not tell you whether those groups are stable, biologically meaningful, or likely to appear in another cohort.

We help you build a research subtype system that can stand up to those questions. Our team can start with your existing data, generate new omics data from your samples, or combine both approaches.

  • Discover molecular groups without forcing a preferred answer
  • Test stability, batch sensitivity, and sample ambiguity
  • Explain each subtype through genes, pathways, cells, and tissue regions
  • Develop a rule for assigning new research samples
  • Test whether the subtype system reproduces in an independent cohort
Sample Submission Guidelines

Molecular subtyping service connecting omics experiments, stable subtype discovery, biological interpretation, and independent cohort replication

What You Receive

  • Recommended subtype number with supporting evidence
  • Subtype assignment for each research sample
  • Confidence and ambiguity flags
  • Subtype-defining features and pathways
  • Batch and sensitivity assessment
  • External replication and next-study recommendations
Table of Contents

    Molecular subtype evidence map covering group stability, batch effects, molecular programs, cell composition, spatial context, and external replication

    A useful subtype system should remain clear when the data, method, or cohort changes.

    Why Molecular Subtyping Needs More Than Clustering

    The goal is not to produce the most colorful heatmap. The goal is to find research groups that remain useful after careful testing.

    Different algorithms can divide the same samples in different ways. Missing data, batch effects, uneven group size, tissue composition, and a few unusual samples can also change the result.

    We therefore treat the first clustering result as a candidate model, not a final answer. We repeat the analysis under reasonable changes, check whether the same samples stay together, and look for technical variables that may explain the groups.

    We then connect the stable groups to molecular programs, cell states, tissue regions, and measured research phenotypes. If the study needs a stronger data layer, we can generate bulk, single-cell, spatial, or multi-omics data in our laboratories.

    Questions we help you answer

    • How many subtypes are supported by the data?
    • Do the same samples remain grouped after resampling?
    • Are batch, site, purity, or cell composition driving the result?
    • Which features and pathways define each subtype?
    • Can a new sample be assigned with a confidence score?
    • Does the pattern reproduce in another cohort?

    Molecular Subtyping for Specific Research Decisions

    We select the data layers and validation plan according to the biological question, not according to one preferred algorithm.

    1

    Find Stable Tumor Subtypes From Bulk or Multi-Omics Data

    Your question: Does this cohort contain repeatable molecular groups, and what biology separates them?

    Data and experiments: The study may use RNA expression alone or combine variants, copy-number changes, DNA methylation, chromatin accessibility, proteins, or metabolites. We can work with existing matrices or generate the missing data from biospecimens.

    How AI helps: AI-assisted methods compare patterns across one or more molecular layers. We test several reasonable group numbers and methods instead of accepting the first result.

    How we test the result: We use repeated subsampling, consensus measures, separation checks, feature sensitivity, and batch-aware review. We also examine whether one small group is caused by low-quality or unusual samples.

    What we deliver: You receive the recommended subtype number, sample assignments, stability evidence, defining features, pathway summaries, and important limitations.

    Published evidence: Hoadley et al. integrated several molecular platforms across more than 10,000 tumors and showed that tissue and cell-of-origin patterns strongly shaped molecular classification. This illustrates why biological context must be considered when interpreting a cluster.

    2

    Resolve Immune and Microenvironment Subtypes

    Your question: Are the groups driven by tumor cells, immune cells, stromal cells, or their spatial organization?

    Data and experiments: Bulk RNA can provide a cohort-wide view. Single-cell RNA, cell-surface protein, immune-receptor, or chromatin data can separate cell states. Spatial profiling can show where those cells and programs occur inside tissue.

    How AI helps: AI-assisted analysis connects cell populations, cell states, pathways, and tissue neighborhoods. It can compare whether a bulk subtype remains visible after accounting for changing cell proportions.

    How we test the result: We compare group structure across samples and data types, check donor and batch effects, and confirm important patterns with an independent dataset or a targeted assay when possible.

    What we deliver: You receive immune or microenvironment group assignments, cell and pathway profiles, spatial maps when included, and markers for follow-up research.

    Published evidence: Sivakumar et al. used matched single-cell multi-omics from pancreatic tumors and blood, together with two public datasets, to identify myeloid-enriched and adaptive-enriched microenvironments. Wang et al. used spatial transcriptomics in 92 triple-negative breast cancer cases and identified nine spatial archetypes that were examined in external datasets.

    3

    Identify Drug-Response or Experimental Phenotype Subtypes

    Your question: Do molecularly distinct groups show different response patterns, resistance mechanisms, or experimental phenotypes?

    Data and experiments: We can combine molecular profiles with measured dose-response values, perturbation results, resistant and sensitive models, longitudinal samples, or other quantitative phenotypes.

    How AI helps: We can first discover groups without using the response label, then compare response across the groups. When appropriate, semi-supervised methods can use a research phenotype to guide the analysis while keeping the molecular structure visible.

    How we test the result: We check whether the association remains after accounting for lineage, batch, and model source. Important findings should be repeated in additional models, samples, or a focused experiment.

    What we deliver: You receive subtype assignments, response distributions, subtype-specific molecular programs, candidate markers, and a validation plan.

    Published evidence: Lehmann et al. combined genomic, transcriptomic, epigenetic, protein, and screening data to study subtype-specific features and research vulnerabilities in triple-negative breast cancer. The study also showed that tumors can contain mixed or continuous subtype signals rather than clean boundaries.

    4

    Lock a Subtype Rule and Test It in Another Cohort

    Your question: Can the subtype system be applied to new samples, and does it reproduce outside the discovery set?

    Data and experiments: We define the discovery cohort, an independent test cohort, required features, sample eligibility rules, and acceptable platform differences. If a test cohort is not available, we can identify a suitable public dataset or plan a later validation study.

    How AI helps: AI-assisted feature selection and classification can convert the discovery groups into a fixed assignment rule. Each new sample can receive a subtype score and an ambiguity flag rather than a forced label.

    How we test the result: We measure assignment rate, confidence, group proportions, molecular concordance, and pathway agreement. We also test the effect of platform, batch, site, and sample composition.

    What we deliver: You receive a locked feature set or centroid rule, new-sample assignments, external replication results, and a clear statement of where the subtype system does and does not transfer.

    Published evidence: Langerud et al. analyzed 1,093 colorectal tumor samples from 692 patients. Multiregional sampling showed frequent subtype heterogeneity, while a lower-heterogeneity gene set supported a more congruent subtype framework that was also illustrated in external tumor series.

    Add the Molecular Layer That Clarifies the Subtypes

    The best technology is the one that separates biological heterogeneity from technical noise and changes the next research decision.

    Service TechnologyQuestion It Can AnswerCommon Role in Subtyping
    RNA SequencingWhich expression programs separate the samples?Cohort-wide subtype discovery, pathway analysis, and assignment signatures
    Whole-Exome SequencingWhich coding variants and copy-number patterns differ by group?Genomic context, subtype-defining alterations, and model comparison
    ATAC-SeqWhich regulatory programs and accessible regions differ by group?Regulatory subtype interpretation and upstream driver analysis
    Single-Cell RNA SequencingWhich cell types or cell states create the bulk signal?Rare states, mixed subtypes, lineage programs, and within-sample heterogeneity
    Spatial Multi-Omics Sequencing ServicesWhere do subtype programs occur inside tissue?Spatial archetypes, cell neighborhoods, boundaries, and resistant niches
    Multi-Omics ServicesDo several molecular layers support the same subtype?Integrated subtyping, cross-layer agreement, and mechanism refinement
    Bioinformatics ServicesCan existing data be processed and compared consistently?Quality review, harmonization, public-cohort analysis, and reproducible reporting

    Spatial profiling is especially useful when samples share a bulk subtype but differ in tumor regions, immune neighborhoods, or stromal organization. The spatial design is selected according to tissue type, preservation method, resolution need, and research question.

    Start With Samples, Existing Data, or Both

    Project TypeWhat You ProvideHow We Support the Study
    Sample-to-InsightBiospecimens, study groups, metadata, and the research decisionWe design and perform the omics experiments, process the data, develop the subtype system, and plan replication.
    Data-to-InsightRaw or processed omics matrices, sample metadata, and optional public datasetsWe review data quality, discover and test subtypes, interpret the biology, and develop an assignment approach.
    Hybrid StudyExisting data plus samples for a missing molecular or validation layerWe identify the evidence gap, generate focused new data, and update the subtype model.

    How We Decide Whether a Subtype Is Credible

    No single score proves that a subtype is real. We review several forms of evidence together.

    • Resampling stability: Do the same samples remain together when the cohort is repeatedly resampled?
    • Method sensitivity: Does the main structure remain visible with other reasonable algorithms, distances, or feature sets?
    • Group separation: Are groups distinct enough to interpret without hiding ambiguous samples?
    • Technical sensitivity: Are batch, site, platform, sample quality, or processing date linked to the groups?
    • Biological agreement: Do genes, pathways, proteins, cells, or tissue regions tell a consistent story?
    • Phenotype association: Are measured research phenotypes distributed differently across groups?
    • Assignment confidence: Can new samples be labeled without forcing borderline cases?
    • External replication: Are the groups and their biology recovered in an independent cohort?

    We report conflicting results as well as supporting results. If the data support a continuous gradient rather than separate groups, we will show that instead of forcing discrete subtypes.

    From Heterogeneous Samples to a Reusable Research Subtype System

    One connected workflow keeps data generation, subtype discovery, quality testing, interpretation, assignment, and replication focused on the same question.

    Molecular subtyping workflow from research question and omics study design through subtype discovery, stability testing, biological interpretation, new-sample assignment, and external replication

    Step 1 - Define the research decision: We clarify the cohort, biological question, expected sources of heterogeneity, and how the subtype result will be used.

    Step 2 - Review data and study design: We examine sample balance, metadata, batch structure, missing values, available molecular layers, and possible validation cohorts.

    Step 3 - Generate or process omics data: When samples are included, we perform the agreed experiments. All projects include data-specific quality control and processing.

    Step 4 - Discover candidate subtype structures: We compare reasonable feature sets, integration strategies, group numbers, and clustering approaches.

    Step 5 - Test stability and confounding: We use resampling, sensitivity analysis, group-separation checks, and technical-variable review.

    Step 6 - Explain the biology: We identify defining features, pathways, cell states, spatial patterns, and research phenotype associations.

    Step 7 - Build the assignment rule: We select a practical feature set or subtype centroid and report confidence or ambiguity for each sample.

    Step 8 - Test external replication: We apply the locked approach to an independent cohort and report agreement, differences, and limits.

    What We Need to Plan the Study

    Early metadata review helps prevent a technical variable from being mistaken for a biological subtype.

    • Research question and how the subtype result will guide the next study
    • Sample type, tissue source, preservation method, and collection site
    • Raw or processed single-omics or multi-omics data
    • Experimental phenotype, treatment, time point, and outcome variables when available
    • Batch, platform, processing date, study site, and sample-quality information
    • Available biospecimens for new bulk, single-cell, or spatial experiments
    • Candidate public or independent cohorts for replication

    Deliverables for Discovery, Review, and Reuse

    • Study design and data-quality review
    • Quality-controlled omics data when experiments are included
    • Recommended subtype number and selection evidence
    • Consensus and resampling stability results
    • Subtype assignment for every eligible sample
    • Confidence, ambiguity, and outlier flags
    • Subtype-defining features and candidate markers
    • Pathway, cell-state, and spatial interpretation as applicable
    • Batch, site, platform, and feature sensitivity results
    • Locked feature set, centroid, or assignment model
    • Independent cohort replication report
    • Reproducible tables, figures, methods, limitations, and follow-up plan

    Move From a Pattern in One Dataset to a Testable Study System

    An analysis-only provider can work with the data that already exist. That may leave the most important question unanswered when the cohort lacks the molecular layer, cell resolution, or tissue context needed to explain the groups.

    CD Genomics combines wet-lab omics, bioinformatics, single-cell analysis, spatial profiling, and AI-assisted pattern review. We can therefore identify the missing evidence, generate it, and connect it to the same subtype question.

    The final result is designed for use in the next research study: a transparent subtype system, a rule for new samples, an external replication assessment, and a focused plan for further testing.

    Research boundary

    Molecular subtypes depend on cohort composition, sample quality, assay design, and biological context. We do not guarantee that separate groups will be found or that a research subtype will transfer to every population, tissue, model, or platform.

    References

    1. Monti S, Tamayo P, Mesirov J, Golub T. Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Machine Learning. 2003.
    2. Hoadley KA, Yau C, Hinoue T, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell. 2018.
    3. Lehmann BD, Colaprico A, Silva TC, et al. Multi-omics analysis identifies therapeutic vulnerabilities in triple-negative breast cancer subtypes. Nature Communications. 2021.
    4. Langerud J, Eilertsen IA, Moosavi SH, et al. Multiregional transcriptomics identifies congruent consensus subtypes with prognostic value beyond tumor heterogeneity of colorectal cancer. Nature Communications. 2024.
    5. Wang X, Venet D, Lifrange F, et al. Spatial transcriptomics reveals substantial heterogeneity in triple-negative breast cancer with potential clinical implications. Nature Communications. 2024.
    6. Sivakumar S, Jainarayanan A, Arbe-Barnes E, et al. Distinct immune cell infiltration patterns in pancreatic ductal adenocarcinoma exhibit divergent immune cell selection and immunosuppressive mechanisms. Nature Communications. 2025.

    Example Molecular Subtyping Report

    The report is built to show both the proposed subtype system and the evidence that supports or limits it.

    Example molecular subtyping report with a consensus matrix, stability comparison, subtype feature heatmap, sample confidence, and external cohort replication

    The report can include a consensus matrix, stability curves, sample assignments, confidence or ambiguity flags, subtype feature heatmaps, pathway profiles, technical-variable checks, and independent cohort results. The exact panels depend on the project.

    AI-Assisted Molecular Subtyping FAQs

    1. How many samples are needed for molecular subtyping?

    There is no single minimum for every study. The answer depends on the expected number of groups, group balance, data type, missing values, and the need for a separate test cohort. We review these factors before recommending a design.

    2. Can the project start with existing data?

    Yes. We can review raw or processed matrices, metadata, and public cohorts. If an important molecular layer is missing, we can design a focused experiment using available biospecimens.

    3. Can you combine several omics layers?

    Yes. We can compare subtypes found in each layer, build an integrated model, or use one layer to explain groups found in another. The choice depends on sample overlap, data quality, and the research question.

    4. What does AI do in the study?

    AI-assisted methods help find complex patterns, compare molecular layers, select informative features, and assign new samples. Our scientists choose the study design, test stability and confounding, interpret the biology, and review the final result.

    5. How do you prevent batch effects from becoming subtypes?

    We review batch and site structure before modeling, use suitable normalization or harmonization, and test the association between proposed subtypes and technical variables. We also repeat the analysis under reasonable processing choices.

    6. Can spatial multi-omics be included?

    Yes. Spatial data can show whether a bulk subtype contains different tumor regions, immune neighborhoods, or stromal patterns. It can also define spatial archetypes that are not visible after averaging the whole tissue.

    7. Can you assign a new sample after subtype discovery?

    When the discovery data support it, we can lock a feature set, centroid, or classifier and apply it to new samples. Borderline samples receive a confidence or ambiguity flag rather than a forced label.

    8. What if the data support a gradient instead of clear groups?

    We report the gradient. Some biological systems contain continuous states or mixed programs. Forcing them into discrete groups can hide important information.

    9. Does external replication guarantee that the subtype system will work everywhere?

    No. Replication strengthens the evidence within the tested populations, assays, and sample types. Transfer to a different setting should be tested separately.

    Published Case Study

    Independent Research Highlight

    Testing a Colorectal Cancer Subtype System Against Within-Tumor Heterogeneity

    This publication is an independent research example. It is not a CD Genomics customer project.

    Background

    A subtype assigned from one tissue region may not represent the rest of the tumor. The researchers asked whether multiregional transcriptomics could identify classifications that were less sensitive to within-tumor heterogeneity.

    Methods

    Langerud et al. analyzed 1,093 colorectal tumor samples from 692 patients. The study included multiregional samples from 98 primary tumors and 35 matched primary-metastasis sets. Figure 1 maps subtype agreement and disagreement across the sampled regions.

    Results

    The study found frequent heterogeneity in consensus molecular subtype assignments. The researchers then focused on lower-heterogeneity, cancer-cell-intrinsic expression signals. Figure 4 compares these classifications across primary tumors and metastases. Figure 5 presents the proposed four-group congruent subtype framework and its research associations.

    Why It Matters

    The study shows why a subtype system should be tested against sampling location, mixed signals, and biological change. It also demonstrates that improved agreement within primary tumors does not remove every limitation: matched metastases could still show subtype switching.

    Conclusion

    A credible subtype project needs more than discovery clustering. Sampling design, stability, biological interpretation, new-sample assignment, external testing, and honest reporting of discordance all matter. Results from this publication do not predict the outcome of another research project.

    Open-access note: The article is licensed under the Creative Commons Attribution 4.0 International License.

    Reference

    1. Langerud J, Eilertsen IA, Moosavi SH, et al. Multiregional transcriptomics identifies congruent consensus subtypes with prognostic value beyond tumor heterogeneity of colorectal cancer. Nature Communications. 2024.

    Selected Publications

    These independent publications provide research foundations for stability testing, multi-omics subtyping, subtype-specific biology, spatial heterogeneity, and external replication. They are not presented as CD Genomics customer projects.

    1. Monti S, Tamayo P, Mesirov J, Golub T. Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Machine Learning. 2003.
    2. Hoadley KA, Yau C, Hinoue T, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell. 2018.
    3. Lehmann BD, Colaprico A, Silva TC, et al. Multi-omics analysis identifies therapeutic vulnerabilities in triple-negative breast cancer subtypes. Nature Communications. 2021.
    4. Langerud J, Eilertsen IA, Moosavi SH, et al. Multiregional transcriptomics identifies congruent consensus subtypes with prognostic value beyond tumor heterogeneity of colorectal cancer. Nature Communications. 2024.
    5. Wang X, Venet D, Lifrange F, et al. Spatial transcriptomics reveals substantial heterogeneity in triple-negative breast cancer with potential clinical implications. Nature Communications. 2024.
    6. Sivakumar S, Jainarayanan A, Arbe-Barnes E, et al. Distinct immune cell infiltration patterns in pancreatic ductal adenocarcinoma exhibit divergent immune cell selection and immunosuppressive mechanisms. Nature Communications. 2025.

    For Research Use Only. Not for use in diagnostic or clinical procedures.

    Nur für Forschungszwecke, nicht zur klinischen Diagnose, Behandlung oder individuellen Gesundheitsbewertung bestimmt.
    Verwandte Dienstleistungen
    Anfrage für ein Angebot
    ! Nur für Forschungszwecke, nicht zur klinischen Diagnose, Behandlung oder individuellen Gesundheitsbewertung bestimmt.