Confidence scores in target prioritization are useful for communicating uncertainty and differentiating candidates, but they are frequently misunderstood in ways that lead to poor decisions. Understanding what a confidence score actually means -- and equally importantly, what it does not mean -- is necessary for using it appropriately.
What a Confidence Score Is
A target confidence score in the Assaygrove context is a composite measure that summarizes the cumulative evidence linking a target gene or protein to a disease indication, taking into account the quality, breadth, and consistency of that evidence across multiple data types. It is a summary statistic, not a direct measurement of any physical quantity.
The score is computed from evidence contributions across three primary tiers: genetic evidence (GWAS associations, rare variant burden, eQTL data), functional evidence (knockout phenotypes, protein interaction data, pathway enrichment), and literature evidence (publication frequency, recency-weighted citation networks, association scoring between target and disease entities). Each tier contributes a weighted component to the composite score, with weights reflecting the tier's general reliability for causal inference about target-disease relationships.
A score of 0.85 on a 0-1 scale means: given the available evidence in the indexed databases, this target has strong, multi-source support for relevance to the specified disease indication. It does not mean there is an 85% probability that the target is causal in the disease, an 85% probability that a drug against this target will succeed in clinical development, or any other specific probabilistic claim that would require epidemiological calibration data that does not currently exist.
What the Score Does Not Capture
Several important dimensions of target quality are not directly captured in a composite evidence score. Druggability -- whether the target protein has a structural binding site amenable to small molecule or biologic modulation -- is not a function of the biological evidence for disease relevance. A target can have outstanding genetic and functional evidence and be nearly undruggable in practice because the relevant protein domain lacks a well-defined binding pocket, the relevant protein interaction is a large flat surface, or the protein is membrane-associated in a way that limits conventional drug modality access.
Selectivity is similarly outside the score's scope. Inhibiting a highly connected hub protein that is deeply involved in disease biology might also disrupt functions in normal physiology in ways that produce unacceptable toxicity. The biological evidence for disease relevance does not speak to the likelihood of achieving the necessary therapeutic window.
Competitive landscape and intellectual property position are also not evidence-based and not captured in the score. A target with a 0.92 confidence score that is the subject of active programs at five major pharmaceutical companies is a different strategic opportunity than a target with a 0.78 confidence score that has not been pharmaceutically pursued.
How to Read the Evidence Decomposition
The composite score number is less informative than the decomposition of that score into its contributing components. A target that scores 0.82 because of a very strong genetic evidence contribution (0.91 in the genetic tier) and moderate functional evidence (0.74) is a different target from one that scores 0.82 because of moderate evidence across all three tiers (0.79, 0.81, 0.84).
The first profile -- high genetic, moderate functional -- points to a target with strong causal support whose functional biology in the specific disease context has not been deeply characterized. The risk profile is specific: the causal link is well-supported, but the mechanistic details that would guide assay development and medicinal chemistry are less clear. The right follow-up is functional characterization work.
The second profile -- consistent moderate evidence across all tiers -- has a different risk profile: no single strong anchor, but consistent supporting evidence that the target is genuinely disease-relevant. The right follow-up might be identifying which existing functional tools (CRISPR reagents, antibodies, chemical probes) are available to move this target into cell-based characterization.
Confidence Intervals and Evidence Depth
A point estimate confidence score without an associated uncertainty measure is incomplete. Two targets might both score 0.80, but one might have that score based on ten independent evidence sources including high-quality functional data, while the other has it based on three evidence sources of moderate quality. The score is the same; the confidence in the score is very different.
The appropriate way to represent this is through confidence intervals or through explicit display of the underlying evidence source count and quality tier. A target with a narrow confidence interval (score 0.80, CI 0.73-0.87) based on deep evidence is fundamentally different from a target with a wide confidence interval (score 0.80, CI 0.54-0.96) based on sparse evidence, even though both show the same point estimate.
Using Scores for Relative Ranking, Not Absolute Decisions
The most appropriate use of composite confidence scores is for relative ranking within a target list -- distinguishing the top candidates from the weaker ones -- rather than as absolute thresholds that determine yes/no selection decisions. A cutoff like "only pursue targets above 0.75" is operationally tempting but epistemically unjustified unless the score has been explicitly calibrated against a reference set of known outcomes, which for most scoring systems has not been done in a rigorous way.
The working model should be: use the score to identify the top tier of candidates, use the evidence decomposition to understand the specific strengths and weaknesses of each candidate's evidence profile, and use domain expertise to assess which evidence gaps are most critical to fill before committing bench resources. The score narrows the field; the evidence decomposition guides the program design.
Score Calibration and Ground Truth
A natural question about confidence scores is whether they are calibrated -- whether a score of 0.80 corresponds to a higher rate of wet-lab validation success than a score of 0.60, and whether the relationship is approximately linear. In principle, a well-designed scoring system should have this property. In practice, rigorous calibration is difficult because the ground-truth labels (did this target validate in a relevant model?) are sparse, often proprietary, and subject to publication bias that systematically underreports negative results.
The honest answer is that most target confidence scoring systems have not been calibrated against a sufficiently large, unbiased dataset of wet-lab outcomes to support strong claims about their absolute calibration. What can be said with more confidence is that relative ranking -- the top decile of scores is more likely to contain genuinely valid targets than the bottom decile -- is typically supported by the evidence. Using scores for relative ranking rather than as absolute probability estimates is both more defensible and more practically useful for the typical target selection decision.