Most research teams come to target selection with a shortlist that feels earned. Months of literature review, expert judgment, maybe a database pull or two. The list looks defensible. But the process that produced it has several structural failure modes that rarely get examined, and examining them changes how you approach prioritization.
The Literature Recency Trap
When researchers build target shortlists from literature, they implicitly weight recent, high-citation papers more than older or obscure work. That is rational in the sense that recent papers often reflect more refined experimental methods and larger sample sizes. But it introduces a systematic bias: targets that have been studied extensively look better than targets that have not. The data you are working with is a sample of what the community has chosen to measure, not a sample of biological reality.
This creates a feedback loop. Heavily-studied targets generate more data. More data increases confidence scores. Higher confidence scores lead to more wet-lab programs. More programs generate more publications. The same twenty gene families dominate drug pipelines not because they are the only valid targets, but because they are the ones the field has measured most intensively over decades. Targets in less-studied biological domains carry genuine uncertainty, but that uncertainty is not the same as lack of relevance.
Siloed Data, Siloed Decisions
The second failure mode is data siloing. Genomics data lives in one analysis pipeline. Proteomics in another. Literature evidence in a separate review process. The analyst integrating these sources is doing cognitive assembly under time pressure, and the integration is imprecise in ways that are difficult to notice from the inside.
When the GWAS signal and the proteomics abundance disagree, what happens? In practice, the analyst makes a judgment call based on their strongest prior. If they are a geneticist, they weight the GWAS. If they are a protein biochemist, they trust the mass-spec. The disagreement itself -- which is often the most informative signal in the dataset, pointing to a specific biological question that needs experimental resolution -- gets resolved rather than examined. The information content in the conflict gets discarded.
A scoring approach that maintains separate evidence tiers and surfaces disagreements explicitly is more useful than one that produces a single integrated number. The point is not to eliminate the need for judgment but to make it clear where judgment is being applied and on what basis.
The Validation Asymmetry
Target shortlists are almost never retrospectively examined for the quality of the selection process. Once a program moves forward, the question of whether the target was correctly ranked gets folded into the larger question of whether the program succeeded. If it fails, the failure gets attributed to the drug, the assay, the dose, the patient selection -- rarely to the target ranking itself.
This makes it nearly impossible to learn from target selection mistakes in any systematic way. The signal does not propagate back to the process that generated the list. A research organization that ran fifty target selection processes and tracked which selection criteria predicted wet-lab success would accumulate genuinely valuable knowledge. Almost none do this, because the data required is scattered across programs, and the programs are rarely analyzed at the level of the selection decision.
The Expert Agreement Problem
Expert consensus within a research group is sometimes treated as evidence that a target is well-supported. It is not. Expert consensus reflects shared training and shared reading of the same literature, which means it inherits all the same biases the literature introduces. A group of experienced computational biologists who all read the same set of high-impact papers on a disease area will tend to converge on the same target candidates regardless of the actual distribution of biological evidence.
Expert review is most valuable for assessing experimental quality -- recognizing weak controls, identifying assay artifacts, understanding which research groups' results are typically reproducible -- rather than for integrating large amounts of evidence across many data types. Using it primarily for the latter introduces precisely the cognitive shortcuts that systematic scoring is meant to replace.
What a More Systematic Approach Looks Like
Fixing this requires treating target ranking as an explicit multi-evidence synthesis problem, not an expert review exercise. That means building scoring that decomposes genetic, functional, and literature evidence separately, with explicit handling of data quality and missing modalities. It means taking data conflicts seriously as informative signals rather than resolving them by fiat. It means applying consistent analysis across the full candidate set rather than a curated shortlist that was itself selected by the process being audited.
None of that replaces domain expertise. The scoring output is most useful when it is reviewed by researchers who can recognize when a high-scoring target is in a biological domain where the evidence is generically poor-quality, or when a low-scoring target has a specific recent finding that the scoring model has not yet incorporated. But the scoring creates a starting point that is more consistent and more complete than expert review alone, and it creates a record that can be re-examined when programs succeed or fail.
The Starting List Problem
There is one more failure mode worth naming explicitly: the starting list that goes into the ranking process is itself often the product of the same informal expert review that systematic scoring is meant to improve. If the candidate set for a target ranking is defined by "proteins that people in this field generally think are interesting," then the ranking will be filtering within a biased prior rather than evaluating a representative sample of potential candidates. Well-studied proteins will dominate the input, and the scoring will differentiate among them, but the genuinely novel candidates that were never part of the informal discussion will never appear at all.
Broadening the initial candidate set -- starting from the full set of proteins with any genetic or functional evidence for involvement in the disease, rather than from an expert-curated shortlist -- is one of the highest-leverage changes a research team can make to its target selection process. It is also one of the least common, because it requires computational infrastructure to handle the larger candidate set and because the output may include unfamiliar proteins that trigger skepticism. Overcoming that skepticism, when the evidence genuinely supports an unfamiliar target, is the hardest part of building a better target selection process. It requires trusting the evidence over the intuition, and that trust has to be earned by demonstrating, over time, that the evidence-based approach produces better outcomes.