# Calculating Adverse Impact in Hiring: A Practical Guide

*Published 2026-09-11*

> Calculating Adverse Impact. Learn how to calculate adverse impact in hiring with selection rates, the 4/5ths rule, worked examples, and practical steps

Source: https://www.talentpronto.ai/blog-posts/calculating-adverse-impact

---

You're reviewing an ATS export after a busy hiring cycle. The requisition closed, the dashboard shows an acceptable overall hire rate, and yet one demographic group appears to have advanced through the process far less often than another. The immediate temptation is to label the result a compliance failure, or to dismiss it as noise. Neither response is reliable.

**Calculating adverse impact** starts with a simpler question: did a neutral hiring step select groups at materially different rates? The four-fifths rule helps surface that pattern, but it isn't a verdict. A defensible review combines reproducible math, stage-by-stage funnel analysis, appropriate statistical testing, validation evidence, and a clear record of what the employer examined and changed.

## Table of Contents
- [What Adverse Impact Means in Practice](#what-adverse-impact-means-in-practice)
  - [The two rates that drive the review](#the-two-rates-that-drive-the-review)
- [The Four-Step Calculation Workflow](#the-four-step-calculation-workflow)
  - [Step 1, prepare the export](#step-1-prepare-the-export)
  - [Step 2, calculate selection rates](#step-2-calculate-selection-rates)
  - [Step 3, calculate the impact ratio](#step-3-calculate-the-impact-ratio)
  - [Step 4, test before escalating](#step-4-test-before-escalating)
- [Choosing the Right Statistical Test](#choosing-the-right-statistical-test)
  - [Match the test to the evidence](#match-the-test-to-the-evidence)
  - [A practical decision sequence](#a-practical-decision-sequence)
- [Reading the 4/5ths Rule Without Misreading It](#reading-the-45ths-rule-without-misreading-it)
  - [Small samples can move sharply](#small-samples-can-move-sharply)
  - [The benchmark may not tell the whole story](#the-benchmark-may-not-tell-the-whole-story)
  - [Aggregate results can conceal a stage problem](#aggregate-results-can-conceal-a-stage-problem)
- [Checking Each Funnel Stage Separately](#checking-each-funnel-stage-separately)
  - [Find the gate with the largest swing](#find-the-gate-with-the-largest-swing)
- [AI Screening, Knockout Questions, and Hidden Bias](#ai-screening-knockout-questions-and-hidden-bias)
  - [Evidence should follow the model](#evidence-should-follow-the-model)
  - [Pair mitigation with an artifact](#pair-mitigation-with-an-artifact)
- [Documenting Results and Taking Corrective Action](#documenting-results-and-taking-corrective-action)
  - [The review record](#the-review-record)
  - [Correct the decision point](#correct-the-decision-point)

<a id="what-adverse-impact-means-in-practice"></a>
## What Adverse Impact Means in Practice

An HR analyst pulling a regional retail hiring export may see two groups entering the same requisition, completing the same application, and receiving different outcomes at the screening or interview stage. That difference doesn't establish intent. It does create a reason to examine the selection process.

**Adverse impact** is an outcome-based measure of whether a facially neutral employment practice disadvantages members of a race, sex, or ethnic group. It differs from disparate treatment, which concerns intentional differential treatment. A hiring manager might apply the same rubric to every applicant and still produce different selection rates across groups. For a helpful legal overview of the distinction between outcome-based disparity and intentional discrimination, see [Nick Norris, P.A.’s disparate impact resource](https://nicknorris.law/2026/03/08/what-is-disparate-impact-discrimination/).

<a id="the-two-rates-that-drive-the-review"></a>
### The two rates that drive the review

The first calculation is the **selection rate**:

**Selection rate = selected applicants ÷ total applicants**

You calculate that rate separately for each group and at each relevant hiring stage. “Selected” might mean advanced past application review, invited to an interview, or offered the job, depending on the decision being tested.

The second calculation is the **impact ratio**:

**Impact ratio = group selection rate ÷ highest group selection rate**

The group with the highest selection rate becomes the benchmark. Under the modern four-fifths rule formalized in the **1978 Uniform Guidelines on Employee Selection Procedures**, a selection rate below **80%** of the highest group's rate will generally be regarded by federal enforcement agencies as evidence of adverse impact. The guidelines were jointly issued by the EEOC, the Civil Service Commission, the Department of Labor, and the Department of Justice. The [EEOC's explanation of the Uniform Guidelines](https://www.eeoc.gov/laws/guidance/questions-and-answers-clarify-and-provide-common-interpretation-uniform-guidelines) provides the governing context.

> **Practical rule:** Treat a ratio below 0.80 as a prompt for investigation, not as proof that discrimination occurred.

The distinction matters in real hiring operations. A flagged ratio may reflect a questionable knockout question, inconsistent interviewer scoring, a narrow sourcing pool, or a small denominator. A ratio above 0.80 also doesn't guarantee that the process is fair. Teams need to show how they calculated the result, which stage generated it, what assumptions they used, and whether the selection method is job-related and supported by validation evidence.

<a id="the-four-step-calculation-workflow"></a>
## The Four-Step Calculation Workflow

A reliable analysis begins with a clean decision table, not a dashboard screenshot. Use one requisition and one defined period so another analyst can reproduce the result from the same ATS records.

For this example, Group A has **100 applicants**, **50 applicants advanced past application review**, and **10 hires**. Group B has **100 applicants**, **30 applicants advanced**, and **6 hires**. These figures are illustrative counts for showing the mechanics, not an empirical case study.

<a id="step-1-prepare-the-export"></a>
### Step 1, prepare the export

Export applicant records, demographic self-identification fields, disposition reasons, stage outcomes, requisition identifiers, and decision dates. Preserve rows with missing self-ID information in a separate category or an excluded-data log. Don't delete them, because removing records can change the denominator and obscure how complete the analysis is.

Clean duplicate applications, reconcile withdrawn candidates with the organization's defined selection rules, and confirm that “advanced,” “interviewed,” and “hired” mean the same thing across both groups. The [standard adverse impact workflow](https://www.prevuehr.com/resources/insights/adverse-impact-analysis-four-fifths-rule/) follows this same sequence, calculate selection rates, identify the highest rate, and compare other groups against it.

<a id="step-2-calculate-selection-rates"></a>
### Step 2, calculate selection rates

For application review, Group A's rate is **50 ÷ 100 = 0.50**. Group B's rate is **30 ÷ 100 = 0.30**.

For hiring, Group A's rate is **10 ÷ 100 = 0.10**. Group B's rate is **6 ÷ 100 = 0.06**.

<a id="step-3-calculate-the-impact-ratio"></a>
### Step 3, calculate the impact ratio

Group A is the benchmark at the hiring stage because **0.10** is higher than **0.06**. Group B's impact ratio is:

**0.06 ÷ 0.10 = 0.60**

That result is below **0.80**, so it flags potential adverse impact at the hiring stage. The same calculation at application review is **0.30 ÷ 0.50 = 0.60**, which also flags the stage.

| Step | Group A, Higher Rate | Group B, Lower Rate | Result |
|---|---:|---:|---|
| Applicants | 100 | 100 | Denominator established |
| Advanced past review | 50 | 30 | Selection rates: 0.50 and 0.30 |
| Hired | 10 | 6 | Selection rates: 0.10 and 0.06 |
| Impact ratio at review | 0.50 | 0.30 | 0.30 ÷ 0.50 = **0.60** |
| Impact ratio at hire | 0.10 | 0.06 | 0.06 ÷ 0.10 = **0.60**, flagged |

<a id="step-4-test-before-escalating"></a>
### Step 4, test before escalating

The four-fifths calculation is a screening test. Pair it with an inferential method suited to the cell counts and the question you're asking. Review whether the observed gap could plausibly arise from the available sample, then assess whether the selection procedure has job-related validation evidence.

Record the inputs, formulas, benchmark group, rounding method, missing-data treatment, and interpretation. A spreadsheet should show the cells that drive every ratio, not only a final red or green status.

<a id="choosing-the-right-statistical-test"></a>
## Choosing the Right Statistical Test

The four-fifths rule answers a practical screening question: **is one group's selection rate less than 80% of the highest group's rate?** Inferential tests answer different questions. They can help determine whether the relationship between group membership and selection is statistically distinguishable from random variation, or whether a scored assessment shows a meaningful difference in performance.

<a id="match-the-test-to-the-evidence"></a>
### Match the test to the evidence

**Fisher's exact test** is usually the safer choice when a contingency table has very small cell counts, such as fewer than **five hires in a subgroup**. It calculates an exact probability rather than relying on the large-sample approximation used by chi-square. The trade-off is that it can be conservative and doesn't tell you whether the disparity is practically important.

**Pearson's chi-square test** works well for larger applicant pools with adequate expected cell counts. It tests whether group membership and selection outcome are associated. It doesn't establish causation, job-relatedness, or the size of the disparity, so a statistically significant result should sit beside the impact ratio, not replace it.

**Standardized mean difference**, often expressed as Cohen's d, applies to continuous assessment scores or ranked candidate results. It compares average scores using a pooled standard deviation. The method is useful for AI-scored rankings, assessments, and structured rating outputs, but it requires meaningful score distributions and careful decisions about whether the score is comparable across groups.

| Test | Best For | Minimum Sample | What It Tells You |
|---|---|---|---|
| Four-fifths rule | Initial adverse impact screening | No universal reliable minimum | Whether a group's selection rate is below 0.80 of the benchmark |
| Fisher's exact test | Small cells and rare selections | Safer when a subgroup has fewer than 5 hires | Whether the observed association is statistically unusual |
| Pearson chi-square | Larger applicant pools and categorical outcomes | Requires adequate expected cell counts | Whether group membership and selection appear associated |
| Standardized mean difference | Scored or ranked assessments | Requires usable score distributions | The magnitude of score differences between groups |

<a id="a-practical-decision-sequence"></a>
### A practical decision sequence

Start with the impact ratio at each funnel stage. If it falls below **0.80**, confirm the data and benchmark, then use Fisher's exact test for sparse cells or chi-square for a sufficiently large table. Reserve standardized mean differences for scored or ranked assessments, particularly when an automated system produces a continuous ranking rather than a simple selected or rejected outcome.

Statistical testing shouldn't be isolated from the hiring design. Teams reviewing [HR data analytics practices](https://www.talentpronto.ai/blog-posts/data-analytics-in-human-resources) should connect each result to the actual decision rule, the role's requirements, and the validation evidence supporting that rule. No single test is universally superior. The right choice depends on the outcome type, sample structure, and decision you need to make.

<a id="reading-the-45ths-rule-without-misreading-it"></a>
## Reading the 4/5ths Rule Without Misreading It

The four-fifths rule is easy to calculate, which is why teams often misuse it. A ratio is a **screening heuristic**, not a legal finding that an employer discriminated. Federal enforcement agencies generally treat a result below **0.80** as evidence of adverse impact, while a result above that level isn't a guaranteed clean bill of health.

<a id="small-samples-can-move-sharply"></a>
### Small samples can move sharply

Suppose a subgroup has a small number of applicants and one or two selections change. Its ratio might move from **0.72 to 0.84** without any change to the underlying hiring design. That swing can be mathematically real and operationally unstable. Guidance on adverse impact analysis warns against relying on the rule mechanically with very small samples, and [Hawaii's selection-procedure guidance](https://labor.hawaii.gov/wp-content/uploads/2019/06/Element-Eight-Exhibit-B.pdf) emphasizes careful benchmarking and attention to small denominators.

<a id="the-benchmark-may-not-tell-the-whole-story"></a>
### The benchmark may not tell the whole story

The highest selection rate is the benchmark, even if that group is unusually strong for a particular role or location. In the worked example, Group A's hiring rate of **0.10** sets the comparison point and Group B's **0.06** produces a **0.60** ratio. That result identifies a disparity relative to Group A, but it doesn't explain whether the cause lies in sourcing, screening, interview scoring, or the job requirements.

<a id="aggregate-results-can-conceal-a-stage-problem"></a>
### Aggregate results can conceal a stage problem

A final ratio near **0.85** might appear acceptable while an interview-stage ratio falls below **0.80**. If later stages select disproportionately from the group that advanced earlier, the overall figure can hide the decision point that created the gap.

| Pitfall | Numeric Example | Correct Interpretation |
|---|---|---|
| Small sample | Ratio shifts from **0.72 to 0.84** after a small number of selections change | Recheck denominators and use an appropriate inferential test |
| Benchmark choice | Group A's **0.10** hiring rate is compared with Group B's **0.06** rate | Use the highest group as the benchmark, then investigate why it is highest |
| Stage shift | Overall ratio is near **0.85**, while an interview stage is below **0.80** | Review each gate instead of relying on the aggregate result |

The analytical record should state the benchmark group, applicant definitions, treatment of missing self-ID data, stage boundaries, sample limitations, and rounding rule. That documentation turns a ratio into an auditable finding rather than an unexplained status label.

<iframe width="100%" style="aspect-ratio: 16 / 9;" src="https://www.youtube.com/embed/9_cVzqW2DiI" frameborder="0" allow="autoplay; encrypted-media" allowfullscreen></iframe>

<a id="checking-each-funnel-stage-separately"></a>
## Checking Each Funnel Stage Separately

A hiring funnel can produce a tolerable overall ratio while one gate creates a serious disparity. Review the decisions in sequence: application review, phone screen, hiring-manager interview, and offer. At each gate, calculate the number selected from the group entering that stage, then compare the stage-specific rates.

Use the same two-group example with **100 applicants per group**. Suppose application review advances **50 Group A applicants** and **45 Group B applicants**. The rates are **0.50** and **0.45**, producing **0.45 ÷ 0.50 = 0.90**, which is above the screening threshold.

At the phone screen, assume **35 Group A applicants** and **30 Group B applicants** advance from the original applicant pools. The cumulative rates are **0.35** and **0.30**, so the cumulative ratio is **0.30 ÷ 0.35 = 0.86**. If you instead calculate the conditional pass-through rate among candidates who reached the phone screen, the denominator changes. That distinction must be labeled clearly.

![A marketing funnel illustration showing the five stages of the customer journey from awareness to retention.](https://www.talentpronto.ai/static/blog-img/calculating-adverse-impact-1.jpg)

<a id="find-the-gate-with-the-largest-swing"></a>
### Find the gate with the largest swing

Suppose the hiring-manager interview advances **20 Group A applicants** and **15 Group B applicants** from the original pools. The cumulative rates are **0.20** and **0.15**, yielding **0.75**, below **0.80**. The interview stage is now the most important place to inspect, even if final offers produce a less dramatic gap.

At the offer stage, assume **10 Group A applicants** and **8 Group B applicants** receive offers. The cumulative rates are **0.10** and **0.08**, producing **0.80**. That result sits at the threshold, but it doesn't erase the earlier interview-stage finding.

> **Audit habit:** Report both cumulative flow and conditional pass-through rates. They answer different questions and can point to different causes.

Use cumulative rates when the question is how many applicants from each group reach a stage from the original pool. Use conditional rates when the question is whether a particular gate treats candidates who reached it differently. Labeling the denominator prevents analysts from comparing unlike measures.

As candidates move toward the bottom of the funnel, subgroup counts shrink. Borderline ratios therefore need more caution, especially when only a small number of applicants reach interviews or offers. An OFCCP-style review may focus on the stage with the largest adverse swing, so preserve the stage definitions, disposition codes, interviewer identities, and decision timestamps that explain how candidates moved through the process.

<a id="ai-screening-knockout-questions-and-hidden-bias"></a>
## AI Screening, Knockout Questions, and Hidden Bias

Automated ranking can create a disparity before a recruiter reads a résumé. Resume parsers may interpret career histories differently, knockout questions may eliminate candidates based on proxy requirements, and predictive screeners may reward patterns embedded in historical hiring data.

Consider a knockout question requiring **three years of consecutive experience at one employer**. In an illustrative screening review, the question eliminates **62%** of one demographic group and **41%** of another, producing a screening impact ratio of **0.66**. The arithmetic is **38% retained ÷ 59% retained = 0.64**, so the stated elimination figures don't mathematically produce 0.66. That inconsistency is exactly why teams must preserve the underlying counts and not rely on a summarized vendor dashboard. Before interpreting the result, reconcile the retained counts, group denominators, and rounding method.

<a id="evidence-should-follow-the-model"></a>
### Evidence should follow the model

For every automated tool, retain a model and decision inventory that identifies:

- **Feature logic:** What inputs influence ranking, eligibility, or rejection?
- **Threshold values:** Which score or response triggers advancement or elimination?
- **Training data:** What datasets, labels, and historical outcomes shaped the model?
- **Validation evidence:** What job-related evidence supports the tool's use for this role?
- **Outcome monitoring:** Which group-level pass-through results are reviewed, and how often?

The employer remains responsible for understanding how a vendor's tool affects applicants. A contract, product brochure, or general assurance that a tool is fair isn't a substitute for role-specific validation and outcome data.

<a id="pair-mitigation-with-an-artifact"></a>
### Pair mitigation with an artifact

Test the tool before deployment using representative applicant data and protected-group pass-through rates. Keep the test dataset, calculation workbook, assumptions, and approval record. Monitor production outcomes by stage, and preserve periodic reports showing whether a ratio changes after the applicant mix or role requirements change.

If a threshold produces a flagged result, recalibrate it only after confirming that the underlying requirement remains job-related. A human review queue can help identify questionable eliminations, but reviewers need a documented rubric and an explanation of when they may override an automated result. [Guidance on reducing bias in AI screening](https://www.talentpronto.ai/blog-posts/how-ai-screening-reduces-bias-in-hiring) is useful for framing these controls.

A platform such as Talent Pronto can conduct conversational screening, apply role-specific questions, produce structured scorecards, and sync candidate statuses with ATS and HRIS systems. Those capabilities can support auditability, but the employer still needs to define valid criteria, review pass-through rates, and make accountable advancement and rejection decisions.

<a id="documenting-results-and-taking-corrective-action"></a>
## Documenting Results and Taking Corrective Action

A defensible adverse impact review should let another qualified analyst recreate the result without asking the original analyst to interpret undocumented choices. Create one record for each job group and analysis period, then attach the source export and calculation workbook.

<a id="the-review-record"></a>
### The review record

Capture:

- **Scope:** Analysis period, requisition or job group, location, and decision stages.
- **Population:** Applicant counts by group, including missing or unknown self-ID treatment.
- **Calculations:** Selection rates, benchmark group, impact ratios, and rounding method.
- **Testing:** Statistical test, p-value, effect size where applicable, and sample limitations.
- **Conclusion:** Whether the result was flagged for further review and why.
- **Evidence:** Job analysis, validation materials, rubrics, model documentation, and disposition definitions.

Use [practical reporting guidance for HR teams](https://www.talentpronto.ai/blog-posts/best-practices-for-reporting) to make the final report readable to HR, legal, recruiting operations, and business leaders. A report that contains only a ratio won't explain what happened.

<a id="correct-the-decision-point"></a>
### Correct the decision point

Start by validating the flagged stage. Check whether recruiters applied the stated criteria consistently, whether interviewers used the same rubric, and whether an automated tool behaved as configured. Remove or adjust a knockout question when it isn't necessary, retrain interviewers when scoring varies, and widen sourcing when the applicant pool is too narrow to support a stable comparison.

Then rerun the analysis in the next cycle and record the change with before-and-after ratios. If an OFCCP or EEOC inquiry arrives, establish an internal response process that can assemble the relevant records within **30 days**. Under Title VII recordkeeping guidance, retain the required records for at least **two years**. Confirm the applicable obligation with employment counsel, especially where another rule requires longer retention.

The practical standard is simple: document the signal, investigate the mechanism, validate the requirement, test less-discriminatory alternatives, and measure the result again. A flagged ratio should lead to disciplined action, not a rushed conclusion.

---

Talent Pronto helps employers standardize early-stage screening with conversational questions, role-specific criteria, structured scorecards, funnel analytics, and ATS or HRIS integrations. If you're building a stage-by-stage adverse impact review, visit [Talent Pronto](https://talentpronto.ai) to evaluate how its screening and reporting workflows could fit your hiring process.
