What you'll learn
- What Algorithmic Bias Actually Looks Like in Hiring Tools
- US Legal Exposure: NYC Local Law 144, the Illinois AI Video Interview Act, and EEOC Guidance
- How to Audit Your Current AI Hiring Tools
- Red Flags When Evaluating AI Hiring Vendors
- Practical Mitigation: Diverse Training Data, Human-in-the-Loop Checkpoints, and Structured Rubrics
- How to Build an Internal AI Governance Policy for Hiring
AI hiring tools were sold as the solution to human bias. Replace the subjective hiring manager gut check with objective algorithmic scoring, and your hiring process becomes fairer by default. That was the promise. The reality is more complicated. Algorithms trained on biased historical data reproduce that bias at scale, faster and more consistently than any individual human could. A resume screener that learned from ten years of hiring decisions at a company that historically promoted white men is not a neutral tool — it is a bias amplifier with an API. This is not a hypothetical problem. There are documented cases of resume screeners deprioritizing graduates of historically Black colleges, video interview AI penalizing non-native English accents, and coding assessment tools scoring women's submissions lower than men's for equivalent code. The regulatory response has arrived. NYC Local Law 144, the Illinois AI Video Interview Act, and EEOC guidance all create legal obligations that sit with the employer, not the vendor. This guide covers what algorithmic bias actually looks like in practice, what the law now requires, how to audit your tools, and what a defensible AI hiring strategy looks like in 2026.
What Algorithmic Bias Actually Looks Like in Hiring Tools
Quick answer
Algorithmic bias in hiring is the predictable output of any machine learning system trained on data that reflects historical discrimination. Resume screeners learn from past hiring decisions. If the engineers hired at a company over the past decade were predominantly male graduates of a small set of universities, the screener learns to weight those signals positively — not because they predict job performance, but because they correlated with the candidates who got hired. Amazon ran into this in 2018 and ultimately scrapped an ML resume tool that was systematically downgrading resumes from graduates of women's colleges and penalizing the word "women's" in candidate summaries. The tool did not intend to discriminate. It had learned to discriminate from a decade of predominantly male hiring data.
Video interview AI introduces a different and more pernicious form of bias: penalizing candidates for characteristics that correlate with protected class status but have no relationship to job performance. Several documented cases involve AI scoring systems that rate non-native English accents lower, independent of the actual content of what the candidate said. Candidates from certain geographic regions received lower automated scores not because their answers were weaker, but because their speech patterns differed from the majority of the training data. Facial analysis systems used in video AI have been shown in independent research studies to have meaningfully higher error rates for candidates with darker skin tones. When these tools make preliminary screening decisions — pass, fail, advance, reject — that error rate is not an abstraction. It is a direct barrier to employment for specific candidate populations.
Cognitive and personality assessments carry their own bias risks. Tests normed on predominantly white, college-educated American samples may systematically underpredict job performance for candidates from different cultural backgrounds where the same behavioral tendencies are expressed differently. A test that flags candidates as "low conscientiousness" based on response patterns calibrated to one cultural context may be measuring cultural difference rather than actual performance risk. The mechanism is always the same: a model trained to distinguish between high and low performers using data from a particular population produces output that reflects the demographic structure of that population, not just the underlying signal it claims to measure. Understanding this is the prerequisite for doing anything useful about it.
US Legal Exposure: NYC Local Law 144, the Illinois AI Video Interview Act, and EEOC Guidance
Quick answer
New York City's Local Law 144 took effect in July 2023 and applies to any employer using "automated employment decision tools" — defined broadly to include AI or machine-learning tools that substantially assist in an employment decision — for roles in NYC. Covered employers must commission an independent bias audit from a qualified third-party auditor and publish the results publicly before deploying the tool. The audit must include disparate impact analysis by sex and race/ethnicity, with selection rates shown for each group relative to the most-selected group. Employers who deploy an AEDT without a current, published audit are in violation regardless of what the tool's vendor claims about their own internal testing. If your company fills any NYC-based roles using automated resume screening, video interview AI, or any other algorithmic scoring tool, LL144 applies.
The Illinois Artificial Intelligence Video Interview Act has been in effect since 2020 and requires employers using AI to analyze video interview submissions from Illinois candidates to: notify candidates before AI is used, explain how the AI evaluates responses, and obtain consent before recording. Recordings may not be shared with third parties except for narrow technical purposes, and must be destroyed within 30 days of a candidate's request. Maryland enacted a similar statute in 2023. Washington state and Colorado have pending or enacted guidance covering automated hiring tools. If your company hires remotely across the US — which means any candidate could be in any of these states — you need a compliance map showing how your tools handle each jurisdiction's requirements, not a general assurance from your vendor that they are "compliant with applicable law."
The EEOC's 2024 technical assistance document clarified the employer liability question that vendors had been muddying: the fact that a hiring outcome was produced by a vendor's algorithm does not relieve the employer of liability for Title VII or ADA violations. If your AI resume screener produces disparate impact — screening out qualified candidates at higher rates for protected characteristics — you can be held liable even if you did not build the tool. The EEOC's position is that employers have an affirmative obligation to evaluate the tools they deploy, test for adverse impact in their own candidate population, and maintain records sufficient to defend their hiring decisions. "The vendor said it was unbiased" is not a defense. That obligation sits with you, and the documentation to support it needs to exist before someone files a complaint.
Algorithmic bias in hiring is not a hypothetical — resume screeners that learned from historically biased hiring data, video AI that penalizes non-native accents, and assessments normed on narrow reference populations all reproduce discrimination at scale, and under EEOC guidance the employer who deployed the tool is liable for the outcomes regardless of who built it.
How to Audit Your Current AI Hiring Tools
Quick answer
An audit starts with an inventory you may not have. Document every point in your hiring process where an automated tool produces a score, rank, recommendation, or pass/fail that influences which candidates advance. That includes resume screeners, chatbot pre-qualifiers, async video interview scoring, cognitive assessments, personality tests, and any custom scoring models your ATS vendor has embedded in pipeline views. Many TA teams have more of these than they realize — recruiting technology has embedded AI scoring into features that do not advertise themselves as AI. A recruiter who sorts a pipeline by "match score" without knowing how that score is calculated is relying on an AI tool they have not evaluated.
Once the inventory exists, request bias audit documentation from every vendor on it. Ask for: training data composition, validation studies showing predictive validity for the specific roles you use the tool for, and disparate impact analysis by race, sex, and age for candidate populations that match your own. NYC Local Law 144 sets a concrete standard for what adequate third-party auditing looks like — if a vendor cannot produce documentation that meets that standard, treat it as a gap regardless of whether NYC law applies to your company. Vendors who resist this request or provide only vague assurances about internal testing are telling you something important about what their data actually shows. Independent third-party audits and vendor-run internal validation are not equivalent; a vendor has an inherent financial interest in what their audit finds.
The internal component is disparate impact analysis on your own candidate population. Pull selection rate data by demographic group at each stage where AI tools are used — what percentage of candidates in each group passed the resume screen, completed the async screen, received an advancing recommendation. If the selection rate for any group is less than 80% of the most-selected group, the four-fifths rule — one of the EEOC's standard measures of adverse impact — is triggered. Your ATS likely has the raw data, and a basic spreadsheet can produce the initial read. If a gap surfaces, determine whether the AI tool is the cause: disable or bypass the tool for a test cohort, run the same analysis, and compare. If the gap narrows, you have identified the source. That finding needs to go to legal and result in either tool modification or a compensating process change.
Red Flags When Evaluating AI Hiring Vendors
Quick answer
The single biggest red flag in an AI hiring vendor evaluation is a scoring model with no explanation of what it actually measures. If a vendor tells you their algorithm is proprietary and they cannot explain what features drive their scores, walk away. You are the party that will face liability if those scores produce disparate impact. "Proprietary" is not a defense. Any vendor who cannot explain their model's key inputs — what signals produce a higher score — is asking you to outsource your hiring decisions to a system you cannot explain to a judge, an auditor, or a rejected candidate who files a complaint. The proprietary protection covers the technical implementation; it should never cover the question of what the model is measuring and whether it predicts performance.
Be skeptical of vendors who claim their tools "eliminate human bias" without producing bias audit results. That claim has become a marketing default in the HR tech category, and it is frequently unsupported by evidence. What eliminates bias is rigorous training data methodology, ongoing disparate impact testing across diverse candidate populations, and independent external audit of the results — not the absence of human reviewers. A vendor who says "our AI is unbiased because it removes human judgment" has misunderstood the source of bias. A vendor who says "here is our current third-party bias audit, showing selection rates by demographic group, conducted by this named third party using this methodology" is the vendor worth taking seriously. Audit documentation should be current — ideally within the past 12 months — and the auditor should be independent, not a firm with a commercial relationship with the vendor.
Also watch for vendors who validate their tools on populations that do not match your candidate base. A cognitive assessment validated on recent college graduates may behave differently when applied to experienced senior candidates. A resume screening tool validated on US-format resumes may produce different outcomes on internationally formatted applications. A video interview AI tested primarily on American English speakers may systematically underperform on candidates with non-native accents — which is both a bias risk and a direct liability under Illinois law if you have candidates in that state. Ask every vendor for the demographic and professional composition of their validation population and compare it directly to your own hiring demographics. A significant mismatch between the two is a documented predictor of both bias outcomes and reduced predictive validity in your specific context.
Related reading
Practical Mitigation: Diverse Training Data, Human-in-the-Loop Checkpoints, and Structured Rubrics
Quick answer
If your company builds internal AI hiring tools or works closely with a vendor on a custom model, training data composition is the first mitigation lever. Models trained on data that is demographically representative of the population they will evaluate produce more equitable outcomes than models trained on narrow reference populations. For internal models, that means auditing the historical hiring data used for training: if the majority of "high performer" labels are attached to candidates who fit a narrow demographic profile, the model will learn to prefer that profile. Synthetic data augmentation, stratified sampling, and explicit reweighting of underrepresented groups in training data are documented technical methods for improving data diversity — any ML team building hiring tools should be using at least one.
Human-in-the-loop checkpoints are the most practical mitigation for TA teams using commercial AI tools where training data is not under their control. The principle is straightforward: AI scores should inform human judgment, not replace it. Any AI-generated output that could trigger an advance or reject decision should have a human review step before action is taken. In practice, that means a recruiter who sees an AI-low-scored candidate takes 60 seconds to confirm the score against the actual resume or video content before declining. This is not a paperwork step. It is the mechanism that catches systematic errors before they become systematic discrimination, and it is also the documentation of human review that satisfies the "human in the decision seat" standard that regulators are increasingly expecting from employers.
Structured interview rubrics operate as a mitigation layer at the human evaluation stage. When interviewers score candidates against defined behavioral competencies with written anchor descriptions, the scoring process produces a consistent, documented record that is more defensible than unstructured impressions. Structured rubrics also reduce the influence of irrelevant characteristics — accent, appearance, cultural communication style — on interviewer scoring, because the rubric directs attention toward specific behavioral evidence for each competency. The combination of structured rubrics at the human stage and documented AI scores at the automated stage produces a candidate evaluation record that can be audited, compared across demographic groups, and defended if a selection decision is challenged. Neither layer alone is sufficient; together they produce the documentation infrastructure that a defensible AI hiring operation requires.
The compliance baseline for US employers in 2026 requires three things: a current third-party bias audit from any vendor whose tools influence hiring decisions, a documented human review checkpoint before any AI-generated score triggers an advance or reject action, and a governance policy with a named owner who is accountable for ongoing disparate impact testing — not a vendor assurance, an internal process.
How to Build an Internal AI Governance Policy for Hiring
Quick answer
An AI governance policy for hiring does not need to be a 50-page compliance document. It needs to do four specific things: identify who owns the decision to deploy a new AI hiring tool, define the due diligence requirements before deployment, establish a testing cadence for tools already in production, and create a documented process for what happens when bias findings surface. The ownership question is typically the most contested. TA, legal, HR operations, and compliance all have legitimate stakes in AI hiring tool decisions, and in most organizations, no one currently owns the full lifecycle. Assigning explicit accountability — a named function that approves new tool deployments and is accountable for ongoing auditing — is the foundational step that makes everything else functional. Without it, governance exists on paper and nowhere else.
The pre-deployment due diligence requirements should be specific enough to be operational. A reasonable baseline for any new AI hiring tool: a current third-party bias audit from an independent auditor, documentation of training data composition and demographic coverage, validation study results for the role types it will be used for, a candidate disclosure plan for jurisdictions with notification requirements, and a data retention and deletion policy meeting the applicable state law requirements. Tools that cannot satisfy this checklist should not be deployed. That is not an unreachable bar — responsible vendors in the category can meet it. It is a bar that filters out vendors who have not done the work to know whether their tools cause harm, and that distinction is exactly the one that matters when a complaint is filed.
Ongoing testing should run at minimum annually, and more frequently if candidate volumes are high enough to produce statistically significant results faster. At each cycle, pull selection rate data by demographic group at every AI-influenced stage and compare against the four-fifths rule threshold. Document the findings. If a gap is detected, the governance policy should specify what happens next — vendor escalation, tool suspension pending investigation, or a process design change that adds a human review layer before the next deployment cycle. The policy is only useful if it has teeth. A governance document that records bias findings and takes no action is not governance — it is a documented trail showing that the organization knew about the problem and did nothing, which is a worse legal position than not having run the audit at all.
What Compliant AI Hiring Looks Like in 2026
Quick answer
A compliant AI hiring operation in 2026 is not one that has avoided AI tools. It is one that has deployed AI tools with documented bias audits, disclosed that use to candidates as required by applicable law, implemented human review checkpoints at consequential decision points, and maintained records sufficient to defend its selection decisions. The distinction between compliance and exposure in 2026 is mostly documentation and governance: two organizations can use the same AI resume screener, and one is exposed while the other is not, based entirely on whether they ran their own disparate impact analysis, obtained a current third-party audit, built a human review process for AI-generated scores, and can produce all of that documentation on request.
The regulatory standard is rising. NYC Local Law 144 set a benchmark for audit quality and public disclosure that other jurisdictions are watching. The pattern in state-level AI hiring regulation is clear: affirmative disclosure requirements, mandatory bias testing, and employer liability for vendor tool outcomes. Companies that have built their compliance infrastructure now — internal audit capability, vendor assessment standards, governance ownership — will be ahead of the curve when the next round of state legislation takes effect. Companies that are waiting for federal preemption to simplify the picture are waiting for something that has not arrived and may not come in a useful timeframe. The cost of building governance infrastructure before a complaint is filed is a fraction of the cost of rebuilding it during litigation.
InCruiter's approach to AI in hiring is built around the principle that AI should surface structured signals for human review, not generate autonomous hiring decisions. Every candidate evaluation in InCruiter produces a documented scorecard from a human interviewer, with AI analysis serving as a supplemental signal layer rather than a pass/fail determination. That architecture is compliant with the current and emerging regulatory landscape because the consequential decision — advance or decline — is made by a human reviewer against documented criteria, not by an algorithm whose inputs cannot be explained. For TA teams that need AI efficiency without AI liability, the design principle is simple: keep humans in the decision seat, keep AI in the signal-generation role, and document both with enough specificity that you can explain every decision to the person who received it.
Frequently asked questions
Common questions about recruiting strategy and how InCruiter helps teams solve them.
InCruiter Editorial Team
AI Hiring Research · Interview Intelligence · Enterprise Talent Strategy
The InCruiter editorial team covers AI-driven hiring, interview intelligence, and modern talent acquisition strategy. Our guides draw on platform data from 2,000+ hiring teams, conversations with talent leaders, and published research in industrial-organizational psychology.


