Behavior Hiring
Evidence review

The predictive validity of behavioral assessments in hiring

An evidence review of how well common hiring methods predict job performance, using the 2022 Sackett et al. revised estimates, and where behavioral hiring fits.

Short answer

Structured interviews have the highest average validity of any method in the 2022 Sackett et al. re-analysis (.42). Unstructured interviews (.19) and years of experience (.07) are among the weakest predictors, yet remain the most common basis for hiring.

Key findings
  • Structured interviews have the highest average validity of any method in the 2022 Sackett et al. re-analysis (.42).
  • Unstructured interviews (.19) and years of experience (.07) are among the weakest predictors, yet remain the most common basis for hiring.
  • GoodJob reports validity of .65–.75 for its full PATH behavioral hiring process from action research since 1985. This comes from a different methodology and is not directly comparable to meta-analytic averages.

For most of the last 25 years, hiring teams were told that cognitive ability tests were the single best predictor of job performance. In 2022, a re-analysis of the underlying research overturned that conclusion. The method that came out on top was one many organizations already use, but rarely use well: the structured interview.

What question does this review answer?

When a hiring team chooses how to evaluate candidates, which methods actually predict who will perform well in the job? And where does a behavioral hiring process fit in that evidence?

What is predictive validity?

Predictive validity is the correlation between a selection method's results and later job performance. It runs from 0 (no relationship) to 1 (perfect prediction). In personnel selection, anything above .30 is generally considered useful, and anything above .40 is strong. A validity of .42 does not mean 42% accuracy; it describes how strongly scores and performance move together.

What does the current evidence show?

For decades, the reference point was Schmidt and Hunter's 1998 summary, which put cognitive ability tests at the top. In 2022, Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens showed that many earlier estimates had been inflated by over-correcting for range restriction. Their revised figures lowered most validities by .10 to .20 and changed the ranking.

Revised mean operational validity, Sackett et al. (2022)
Selection methodValidity (r)
Structured interviews0.42
Job knowledge tests0.40
Empirically keyed biodata0.38
Work sample tests0.33
Cognitive ability tests0.31
Integrity tests0.31
Assessment centers0.29
Conscientiousness measures0.19
Unstructured interviews0.19
Years of job experience0.07

Two caveats matter. First, these are averages. Structured interviews, for example, have an 80% credibility interval running roughly from .18 to .66, which means design quality makes a large difference. Second, the methods are not mutually exclusive. The best systems combine several.

What pattern explains the ranking?

The methods near the top share three properties. They are tied to the specific job, they are scored the same way for every candidate, and they look at evidence of what the person actually does. The methods near the bottom rely on general impressions (unstructured interviews), general traits (broad personality measures) or proxies (years of experience).

Behavioral hiring is designed around those three properties: a role-specific behavioral profile, a consistent assessment, and a structured interview that verifies the results with evidence.

Where does the PATH process fit?

GoodJob's PATH behavioral hiring process combines a work-context behavioral self-assessment with a RoleDNA profile of the role, structured behavioral interviews in which interviewers rank answers, and onboarding and coaching playbooks. PATH was developed in response to 1985 Harvard University research into the service value chain, and GoodJob reports predictive validity of .65 to .75 from its action research since then.

How to read this figure

The .65–.75 figure is reported by GoodJob for the full multi-step process, from its own action research. It has not been through the same independent meta-analytic process as the Sackett et al. estimates, which average many independent studies of single methods. It should be read as a practitioner result for a combined system, not as a direct comparison with the table above. We will publish an independent methodology summary as it becomes available.

What should hiring teams do with this?

  • Replace unstructured interviews with structured, role-specific ones. This is the single largest improvement most teams can make.
  • Stop treating years of experience as a performance predictor.
  • Use personality tools for development, not selection.
  • Combine methods: a role profile, a behavioral assessment and a structured interview cover more ground together than any one alone.

Method

This is a narrative evidence review. It draws on the published Sackett et al. (2022) revised estimates and the practitioner data reported by GoodJob for the PATH process. It is not a new meta-analysis.

Sources

  1. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
  2. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
  3. GoodJob. PATH behavioral hiring process and action research summary. https://goodjob.io

How to cite this page

Behavior Hiring. (2026). The predictive validity of behavioral assessments in hiring. https://behaviorhiring.com/research/predictive-validity-behavioral-assessments/

Monthly Research Brief

One email a month. New research, a practical tool, and what changed in the field.

No spam. Unsubscribe any time.