Executive summary
Hiring outcomes hinge on three questions most HR leaders have never had a metric to answer. Which of my current tools actually work? Why do referrals and internal candidates consistently outperform cold applicants regardless of process? And how do I upgrade the talent pool I'm evaluating in the first place? The Diagnostic Ratio makes all three answerable — and the answers point in the same direction.
The rest of this report develops these three points with data. The Diagnostic Ratio audit in Table 1 quantifies which tools work and which don't. The 2x2 strategic map shows where most corporate hiring processes currently live. Table 2 uses Bayes' Theorem to show how dramatically sourcing channel changes the downstream math. The deeper-dive sections explain the mechanism. The implementation guide at the end translates the findings into a 90-day action plan.
The strategic map — where your hiring process lives
The two dimensions that determine hiring quality are the DR of the tools you apply and how much you already know about the candidate. These are independent levers. Improving either one improves outcomes; improving both produces elite hiring.
Table 1 · Diagnostic ratios of common hiring tools
Numbers are best-estimate midpoints from meta-analytic validity research, primarily Sackett, Zhang, Berry & Lievens (2022). Hover section headers for strategic framing; hover DR values for interpretation.
| Tool | r | R² | Sensitivity | Specificity | DR | Impact statement |
|---|---|---|---|---|---|---|
| Foundation — what every other tool measures candidates against | ||||||
| Performance-based Job Description (PBJD) with Key Performance Objectives | ~0.50 | ~25% | 77% | 85% | 5.25 | Job analysis done as requirements engineering. Defines 6–8 KPOs the role must produce. Validated by Sackett (2022) at work-sample validity; operationalizes Gallup Q1 at hiring; legally defensible per the Littler Validation; provides the reference standard for 2-Step screening and the Win-Win career-move close. |
| Basic job analysis (formal task inventory — PAQ, Fleishman, incumbent interviews) | ~0.45 | ~20% | 74% | 84% | 4.63 | The conventional I/O approach — formal task-inventory analysis producing a KSAO catalog. Strong diagnostic foundation, legally defensible. Does not produce the outcome specification that drives 2-Step screening or the Win-Win close; the PBJD row above is the upgrade path. |
| Traditional job description (template-based) | ~0.05 | <1% | 50% | 50% | 1.00 | The corporate default. No diagnostic signal because the document doesn't specify performance outcomes. Every downstream tool measures candidates against a non-standard. |
| Sourcing channel — sets the prior before any tool is applied | ||||||
| Internal promotion (observed performance) | ~0.55 | ~30% | 85% | 90% | 8.50 | The highest-DR event in the entire hiring process. You have watched them do the work. Bypasses traditional screening. Most companies fill externally first, bypassing their best diagnostic signal. |
| Referral from trusted former coworker | ~0.45 | ~20% | 75% | 88% | 6.25 | Acquaintance networks carry observed-performance information no stranger-directed process can replicate. Typically bypasses the traditional screening stack and goes straight to high-DR evaluation. |
| Boomerang (returning former employee) | ~0.42 | ~18% | 70% | 85% | 4.67 | Known quantity with verified prior performance. Rivals structured interviews in diagnostic power. Usually overlooked because of "left us once" bias. |
| Recruiter-sourced passive candidate | ~0.20 | ~4% | 55% | 70% | 1.83 | Modest lift over cold applicants. Selection bias helps (recruiters target stronger profiles) but observation of actual performance is shallow. |
| Cold inbound applicant (job board) | ~0.05 | <1% | 50% | 50% | 1.00 | Zero prior information. This is the sourcing channel that most needs the 2-Step screening process — the 2-Step generates acquaintance-like evidence from strangers by capturing a comparable accomplishment at application time. |
| Low-DR screening tools — the tools to stop using | ||||||
| Keyword / resume screen | 0.10 | 1% | 40% | 55% | 0.89 | Worse than a coin flip. Rejects more top performers than average performers. Actively destroys pipeline quality while looking rigorous. |
| Skills-based screening (declared skills / resume tags) | 0.11 | 1% | 45% | 60% | 1.13 | Keyword screening rebranded. Opens the pool by dropping degree requirements — a real equity win — but adds almost no diagnostic power when skills are self-declared rather than measured. |
| Years of experience requirement | 0.09 | <1% | 75% | 30% | 1.07 | Explains less than 1% of who succeeds. Drives most corporate resume screening — a tool chosen for convenience, not evidence. |
| Degree / GPA requirement | 0.10 | 1% | 55% | 45% | 1.00 | Pure theater. One percent variance explained. Generates EEOC exposure with zero selection return. |
| Company pedigree ("must have FAANG") | 0.12 | 1.5% | 25% | 85% | 1.67 | Rejects 75% of real top performers to marginally enrich a small pool. Expensive Type II error generator. |
| Reference check (as typically conducted) | 0.13 | 2% | 45% | 55% | 1.00 | Near-zero diagnostic value. Social norms force positive references. Signal swamped by politeness. |
| Unstructured interview | 0.19 | 4% | 50% | 55% | 1.11 | Effectively noise. Interviewer confidence is inversely correlated with accuracy. The most-used evaluation method is among the weakest. |
| High-DR replacements — the tools to adopt | ||||||
| Personality test (Big Five) | 0.19 | 4% | 55% | 60% | 1.38 | Modest signal at best. Conscientiousness is the only trait with meaningful validity. Heavily oversold by vendors. |
| General Mental Ability (GMA) | 0.31 | 10% | 65% | 75% | 2.60 | Real signal, long overclaimed. Sackett's 2022 revision cut historical validity estimates in half. Still one of the better single tools, but 90% of variance remains unexplained. |
| Structured interview (rigorously executed) | 0.42 | 18% | 70% | 80% | 3.50 | Strong evaluation tool by modern standards. Most companies claim to use them; few actually do. The gap between "structured" and "claimed structured" is where validity is lost. |
| Validated skills assessment (work-sample platforms) | 0.45 | 20% | 73% | 83% | 4.29 | Platforms like TestGorilla, Codility, and Vervoe functionally deliver work samples. Near-gold-standard diagnostic power when rigorously implemented. |
| Performance-based Interview (probing Most Significant Accomplishments) | 0.48 | 23% | 75% | 85% | 5.00 | The gold standard single evaluation method. Probes Most Significant Accomplishments against KPOs. Statistical foundation underneath Performance-based Hiring's core interview methodology. |
| The 2-Step process and the full stack — upgrading the talent pool, then evaluating it | ||||||
| 2-Step screening (parallel Pass 1 + Pass 2) | ~0.55 | ~30% | 78% | 88% | 6.50 | Parallel OR gate at screening. Either traditional credentials (Pass 1) or performance evidence (Pass 2) advances the candidate. Rescues Hidden Gems. Upgrades the talent pool feeding into high-DR evaluation tools. |
| 2-Step screening + Performance-based Interview | ~0.60 | ~36% | 80% | 90% | 8.00 | Screening + evaluation working together. 2-Step screening surfaces Hidden Gems; Performance-based Interview confirms capability and fit. Diagnostic power no single tool reaches. |
| Full PBH stack (PBJD + referral sourcing + 2-Step + Performance-based Interview) | ~0.72 | ~52% | 90% | 95% | 18.00 | The diagnostic ceiling. PBJD foundation + high-prior sourcing + 2-Step screening + rigorous evaluation. What elite hiring looks like when all three stages of the funnel are tuned. |
Table 2 · How sourcing channel changes everything downstream
Using DR values from Table 1 as inputs, this table shows how the same evaluation tool applied to a different prior produces a different posterior probability. Bayes' Theorem in full view.
| Sourcing channel | Prior (before any tool) | After structured interview | After 2-Step + PBI |
|---|---|---|---|
| Cold applicant (job board) | ~20% | ~47% | ~67% |
| Passive candidate (recruiter-sourced) | ~30% | ~60% | ~77% |
| Boomerang (returning employee) | ~55% | ~81% | ~91% |
| Referral from trusted former coworker | ~65% | ~87% | ~94% |
| Internal promotion (observed performance) | ~75% | ~91% | ~96% |
An internal promotion candidate who passes a structured interview has a 91% probability of being a top performer. A cold applicant who passes the same interview has a 47% probability. Same evaluation tool. Same job. Same company. Dramatically different outcomes, driven almost entirely by the prior the sourcing channel establishes.
What the Diagnostic Ratio actually is
The Diagnostic Ratio (DR) — known in clinical diagnostics as the positive Likelihood Ratio (LR+) — measures how strongly a piece of evidence should update your belief. In hiring, the hypothesis is "this candidate is a top performer" and the evidence is "the candidate passed this tool."
A DR above 1 means passing is evidence in favor of the candidate being a top performer. A DR below 1 means passing is actually evidence against. A DR of 1 means the tool tells you nothing.
The three stages of the hiring funnel
The Diagnostic Ratio applies differently at different stages. Separating them clarifies what each tool does.
How the 2-Step screening process works
Most corporate hiring applies a serial filter stack at the screening stage: keyword screen, credential check, years-of-experience threshold, in sequence. Each stage compounds Type II errors — rejecting top performers who didn't match one narrow criterion. The cumulative effect rejects most real top performers before any interview happens.
The 2-Step screening process runs a second filter in parallel with traditional screening. Pass 1 is the traditional screen (skills, experience, credentials). Pass 2 is a short accomplishment write-up demonstrating work comparable to the role's KPOs. Candidates advance by passing either one. The four possible outcomes:
The parallel OR gate upgrades the talent pool feeding into the evaluation stage. A referral or internal candidate arrives with Pass 2 already satisfied because their performance is already observed. A cold applicant generates Pass 2 evidence via the accomplishment write-up. The net effect: strangers become eligible for the same high-DR evaluation treatment that acquaintances already receive.
On job analysis methodology — anticipating the I/O objection
An I/O psychologist reviewing this report may object that a Performance-based Job Description is not real job analysis. This objection gets the question backward. A PBJD is job analysis — done as requirements engineering rather than construct cataloging — and produces the outcomes job analysis exists to produce, often better than formal task-inventory approaches.
The PBJD methodology draws on O*NET task and work-activity taxonomies, satisfies Uniform Guidelines (1978) essential-functions requirements (as analyzed in the Littler Validation), and is operationalized through an 8-step wizard covering objectives, KSAO-to-outcome conversion, timeline, team, problem-solving, environment, and employee value proposition.
How to read the columns
Implementation — what to do with this
A 90-day sequence. Each step is independently valuable; the order makes them compound.
Important caveats
Sensitivity and specificity values are derived from reported validity coefficients using Taylor-Russell-style translations at an assumed top-25% performance threshold. Actual values depend on threshold, base rate, and implementation quality. Sourcing channel validity numbers are inferred estimates translated from performance-differential studies — the referral and internal-mobility advantage is supported by Castilla (2005), Burks, Cowgill, Hoffman & Housman (2015), and Baker, Gibbs & Holmstrom (1994).
The stacked-method numbers (2-Step DR = 6.5; 2-Step + Performance-based Interview DR = 8.0; full stack DR = 18.0) are theoretical combined estimates, not direct meta-analytic findings. Real implementations fall below these ceilings because component tools measure overlapping constructs. The PBJD numbers (r ~0.50, DR 5.25) are calibrated to Sackett et al. (2022) treatment of job analysis as a top-tier selection validity driver; legal defensibility is independently analyzed by the Littler Validation.
Primary sources: Sackett, Zhang, Berry & Lievens (2022) for selection method validity; O*NET / U.S. Department of Labor for occupational task taxonomies; Harter, Schmidt et al. (Gallup Q12 meta-analyses) for engagement validity; Goldstein (2013, Littler Mendelson) for legal compliance (the Littler Validation); Uniform Guidelines on Employee Selection Procedures (1978). Related: Schmidt & Hunter (1998) with modern revisions; McDaniel et al. (1994); Morgeson et al. (2007).