What you'll learn
- What Predictive Validity Actually Measures
- Where the Common Methods Actually Rank
- Why Combining Methods Beats Any Single Best Method
- What This Means for Redesigning a Hiring Process
Most hiring decisions still rest heavily on the unstructured interview — a free-flowing conversation with no consistent question set and no defined scoring standard — despite decades of meta-analytic research showing it's one of the weakest predictors of actual job performance among commonly used hiring methods, in a similar range to years of job experience, itself a surprisingly weak standalone signal. This guide covers what the research on predictive validity actually shows, where common assessment methods — structured interviews, work samples, cognitive ability tests — actually rank, why combining a small number of complementary methods produces meaningfully better predictions than relying on any single method alone, and how to use this evidence to prioritize the highest-leverage changes to a hiring process that's currently built around methods the research says are weaker than they feel.
What Predictive Validity Actually Measures
Quick answer
Predictive validity in hiring research refers to how well a given assessment method, taken during the hiring process, correlates with actual job performance measured later — typically supervisor ratings, objective output metrics, or training performance, collected some months or years after hire. It's expressed as a correlation coefficient, generally ranging from 0 (no relationship between the assessment and later performance) to 1 (a perfect relationship), and the most influential body of research on this topic comes from decades of meta-analytic work aggregating hundreds of individual studies across industries and roles, most notably ongoing work building on the foundational Schmidt and Hunter meta-analyses.
A useful reference point: a validity coefficient around 0.50 to 0.60 is considered quite strong for a hiring method, while anything below roughly 0.20 is weak enough that the method is barely better than chance at predicting who will actually perform well. Very few individual assessment methods, used alone, reach the top of that range — which is itself an important finding, since it means no single method, however well-designed, should be relied on exclusively for a consequential hiring decision.
This research matters directly to hiring practice because it provides an evidence-based way to prioritize where to invest limited time and process rigor. If a company is deciding whether to invest more heavily in structuring its interviews or in adding a work sample exercise, validity research gives a genuine, evidence-based basis for that decision, rather than defaulting to whatever hiring process feels intuitively rigorous or whatever a previous employer happened to use.
Where the Common Methods Actually Rank
Quick answer
Unstructured interviews — the free-flowing, unplanned conversational interview still used as the primary or sole evaluation method at a large share of companies — show validity coefficients in the range of roughly 0.20 to 0.38 across meta-analytic estimates, putting them near the bottom of commonly used methods and, notably, in a similar range to years of job experience, which is itself a weak standalone predictor despite being one of the most heavily weighted factors in typical resume screening. This is a genuinely uncomfortable finding for how most organizations actually hire, given how central the informal interview conversation remains to most hiring processes.
Structured interviews — where every candidate is asked the same predetermined, job-relevant questions and scored against a consistent rubric — show meaningfully higher validity, generally in the 0.40s to low 0.50s depending on the specific structuring approach, roughly double the predictive power of the unstructured version despite often taking a similar amount of interviewer and candidate time. The gap between structured and unstructured interviews is one of the most robust and consistently replicated findings in the entire body of hiring research, which makes the continued prevalence of unstructured interviewing across the industry somewhat surprising given how long this evidence has existed.
Work sample tests — having candidates actually perform a representative sample of the job's real tasks — consistently rank among the strongest individual predictors, generally in the 0.30s to 0.40s and sometimes higher depending on how closely the sample mirrors actual job content, with the added practical benefit of typically producing a better candidate experience than more abstract assessment formats, since candidates can directly see the assessment's relevance to the actual job. General cognitive ability tests also rank highly, generally in the 0.40s to 0.50s, and notably show some of the most consistent validity across a very wide range of job types and complexity levels, though they raise their own separate considerations around adverse impact that need to be actively managed alongside the validity benefit.
Unstructured interviews — still the most common hiring method by far — have one of the weakest predictive validity records of any method organizations regularly use, roughly on par with years of job experience, which itself is a surprisingly weak predictor on its own.
Why Combining Methods Beats Any Single Best Method
Quick answer
The research finding with the most direct, practical implication for hiring process design is incremental validity: combining two or more assessment methods with different, complementary validity profiles produces meaningfully better overall predictive power than any single method alone, even a strong one — because each method captures somewhat different aspects of what predicts job performance, and the combination reduces the noise and blind spots inherent in relying on just one lens. A cognitive ability test alone, a structured interview alone, and a work sample alone are each individually solid; the same three combined outperform any of them individually by a meaningful margin.
The specific combination that shows some of the strongest documented incremental validity in the research is a general cognitive ability measure paired with a structured interview, since cognitive ability captures general problem-solving and learning capacity while a well-designed structured interview captures job-specific knowledge, judgment, and interpersonal factors that a cognitive test doesn't directly assess — the two methods are measuring meaningfully different things, which is exactly why combining them adds real predictive value rather than just redundantly confirming the same signal twice.
There are diminishing returns to adding more and more methods, and practical costs — candidate time, interviewer time, process complexity — that need to be weighed against the marginal validity gain from each additional assessment stage. Most of the incremental validity benefit is captured with two to three well-chosen, complementary methods; a hiring process with six or seven separate assessment stages is unlikely to meaningfully outpredict a well-designed three-stage process, while imposing significantly more cost and candidate drop-off risk along the way.
What This Means for Redesigning a Hiring Process
Quick answer
If a hiring process currently relies primarily on unstructured interviews, the single highest-leverage change available, based on the weight of the research, is structuring those interviews — consistent questions, a defined scoring rubric, calibrated interviewers — since this captures a substantial validity improvement without adding an entirely new assessment stage or additional candidate time. This is also usually the lowest-cost change to implement relative to its expected impact, since it doesn't require new tooling or additional interview rounds, just a more disciplined approach to the interviews already happening.
Adding a well-designed work sample exercise, closely mirroring actual job tasks, is generally the next highest-leverage addition for roles where a genuine sample of the work can be constructed in a reasonable amount of candidate time — this is often more straightforward for technical, analytical, and craft-based roles than for roles with a longer time horizon between action and visible outcome, where an authentic short sample is harder to construct convincingly.
Cognitive ability testing carries real, documented adverse impact risk across some demographic groups, which needs to be weighed directly against its strong validity rather than treated as either an automatic inclusion or an automatic exclusion. Organizations considering it should review current legal guidance and their own internal adverse impact data specifically for this assessment type, and in many cases, combining it with methods that show more limited adverse impact — like a well-structured interview or a job-relevant work sample — can help offset the overall selection process's impact while retaining much of the incremental validity benefit that motivated its inclusion in the first place.
Related reading
Frequently asked questions
Common questions about candidate assessment and how InCruiter helps teams solve them.
InCruiter Editorial Team
AI Hiring Research · Interview Intelligence · Enterprise Talent Strategy
The InCruiter editorial team covers AI-driven hiring, interview intelligence, and modern talent acquisition strategy. Our guides draw on platform data from 2,000+ hiring teams, conversations with talent leaders, and published research in industrial-organizational psychology.



