Proteomics and genomics data both appear in drug target discovery workflows, and both are described as providing "evidence" for target relevance. But they provide fundamentally different kinds of evidence, with different strengths, different failure modes, and different implications for how much confidence you should place in a target hypothesis derived primarily from one versus the other.
What Genomics Data Tells You
Genomic data in target discovery context primarily means GWAS data and rare variant analyses: statistical associations between genetic variants and disease phenotypes in human populations. The key property of this data is causal directionality. Genetic variants are inherited at conception, before disease onset, and are not substantially affected by disease processes, treatment history, or environmental exposures. When a genetic variant that disrupts a specific gene's function is associated with disease risk, the most parsimonious interpretation is that the gene's function is causally relevant to the disease.
This causal inference is the primary reason genetics has become the highest-weighted evidence tier in most target prioritization frameworks. The empirical validation supports it: drugs against targets with genetic evidence for human disease involvement have roughly twice the clinical success rate of drugs against targets without such evidence.
The limitations of genomics data are also worth being explicit about. GWAS identifies statistical associations between variants and phenotypes, not between genes and phenotypes. A GWAS hit in a region of high linkage disequilibrium implicates dozens to hundreds of variants that co-inherit in the population; determining which specific gene or variant is functionally responsible requires additional experimental work. Fine-mapping analyses using statistical methods and functional annotation can narrow the credible set, but ambiguity is common.
What Proteomics Data Tells You
Proteomics data -- typically from mass spectrometry experiments -- measures the abundance of proteins in a specific biological sample. In a disease context, proteomics can reveal which proteins are more or less abundant in disease tissue compared to normal tissue, which proteins are post-translationally modified in disease-specific ways, and which proteins are present in cellular fractions that suggest altered localization or complex formation.
Proteomics evidence is more directly relevant to drug targeting than genomics in one important sense: drugs act on proteins, not directly on genes. A protein that is upregulated in disease tissue and is physically accessible to a drug molecule is a more tractable target than a gene whose variants are disease-associated but whose protein product is ubiquitous, intracellular, and structurally difficult to target.
The limitation is causal ambiguity. A protein that is elevated in disease tissue might be a driver of the disease process, a consequence of the disease process, a bystander that changes in abundance for reasons unrelated to the disease mechanism, or a compensatory response to the disease. All four scenarios produce the same proteomics signature -- increased protein abundance in disease -- but have completely different implications for whether targeting that protein would be beneficial.
When They Agree and When They Conflict
The most confident target hypothesis comes from convergent evidence: a target with a genetic association to the disease, a disease-relevant change in protein abundance, and a relevant loss-of-function phenotype. Each evidence type is independently confirmatory, and their agreement substantially reduces the probability that any single evidence type is reflecting an artifact or a biologically irrelevant association.
Conflicts between genomics and proteomics data are actually more informative than agreement, because they point to specific biological questions that need to be resolved. A target with strong genetic association but no protein expression change in disease tissue might reflect the genetic variant operating through a mechanism other than abundance regulation -- perhaps through altered protein function, changed interaction partners, or cell-type-specific expression changes not captured by a bulk tissue proteomics experiment. A target with a strong proteomics signal but no genetic association might be a genuine driver in a genetic architecture that GWAS is underpowered to detect, or it might be a consequence of disease biology that is not itself causal.
The mRNA-Protein Correlation Problem
When researchers use transcriptomics data as a proxy for protein evidence -- which is common, because RNA-seq is cheaper and easier than mass spectrometry proteomics -- they implicitly assume a reasonable correlation between mRNA and protein levels for the genes of interest. For some gene families this assumption holds reasonably well. For others it does not.
Published studies of mRNA-protein correlation in mammalian cells typically find Pearson correlations in the range of 0.4 to 0.6 across the proteome, with substantial variance across individual genes. Proteins with high rates of translational regulation, post-translational modification, or regulated degradation can show minimal correlation between their transcript abundance and their protein abundance. Using transcriptomics as an uncritical proxy for proteomics data for these gene families produces incorrect target assessments.
Practical Integration Guidance
The right approach is to treat genomics and proteomics as complementary and largely orthogonal evidence types, each contributing a distinct dimension to a composite target assessment. Genomics provides the causal prior; proteomics provides evidence about protein-level disease relevance and tractability. Neither substitutes for the other, and a target assessment that relies exclusively on one to the exclusion of the other is missing a significant fraction of the available evidence.
When integrating the two, being explicit about what each data type is measuring and what its limitations are for a specific disease context is more valuable than applying a fixed integration formula. The right weights for genomics versus proteomics evidence depend on the quality of the specific datasets available, the disease context, and what specific biological questions the evidence is being used to answer.
The Role of Transcriptomics as a Bridge
Transcriptomics data -- RNA-seq -- occupies a middle position between genomics and proteomics in most evidence frameworks. It is more cost-effective than proteomics and provides broader coverage of the transcriptome than proteomics currently achieves for the proteome. As a result, it is often used as a proxy for protein evidence when proteomics data is not available for the relevant tissue or cell type.
The validity of that proxy depends on how well mRNA levels correlate with protein levels for the specific genes of interest. For genes where the correlation is high -- typically housekeeping genes and many structural proteins -- transcriptomics is a reasonable substitute. For genes with complex post-transcriptional regulation, high protein turnover rates, or significant translational control, the correlation can be weak enough that transcriptomics provides almost no information about protein abundance. Knowing which category a target gene falls into before treating transcriptomics as a protein proxy is worth the modest investment in checking the available literature on mRNA-protein correlation for that specific gene family.