Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When an AI hiring tool produces a result that seems wrong or uneven across groups, first identify the exact decision, stage, model version, inputs, job criterion, and threshold behind it. Then compare the tool’s record with the application, measure outcomes at each stage, and investigate whether the result follows from a job-related criterion or a data, accessibility, or configuration problem. A disparity metric is a reason to investigate—not proof that a system is biased, fair, accurate, or legally compliant.

Start by tracing the disputed decision

“AI candidate triage” can mean screening applicants out, assigning scores, ranking candidates, placing them in categories, or recommending who should advance. Those outputs are not interchangeable. Before changing a threshold or accepting a vendor’s explanation, establish what the tool actually did in the case at issue.

  1. Record the decision and configuration

    Capture the role, hiring stage, decision date, tool and model version, relevant settings, cutoff or ranking rule, and the output that screened out or downgraded the candidate. Preserve the configuration while comparing cases where feasible; otherwise, a changed setting can obscure the cause.

  2. Reconstruct the candidate’s inputs

    Identify the data fields used, where they came from, how old they were, how they were transformed, and how missing values were handled. Compare the tool’s record with the actual application. A parsing error, stale profile, omitted qualification, or incorrect job requirement can appear in the output as a candidate deficiency.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  3. Identify what the output represents

    Determine whether the result is a score, rank, recommendation, classification, or a binary pass/fail decision, and which stage consumed it. A ranking can change who receives attention without formally rejecting anyone; a cutoff can directly screen people out. The distinction matters when measuring outcomes and deciding what to correct.

This tracing sequence is a practical diagnostic method, not a technical procedure prescribed by the statutes discussed below.

Check whether the criterion belongs in the hiring decision

Compare each scored feature and cutoff with the role’s written requirements. Ask whether the feature measures a capability genuinely needed for the work, whether a proxy may be standing in for a protected characteristic, and whether the same standard is applied consistently to applicants.

  • Is the criterion tied to an actual task or qualification, rather than a convenient but weak correlate?
  • Would an applicant with the relevant ability be disadvantaged by the way the tool measures it—for example, by a particular format, response style, or assessment condition?
  • Is the cutoff justified for this role and stage, or was it selected without checking what it does to applicants?
  • Are equivalent qualifications treated consistently, including when they are described in different language?

EEOC guidance says selection criteria with a significant discriminatory effect must be job-related and consistent with business necessity. The EEOC’s national-origin guidance identifies objective written criteria that are communicated and consistently applied as a promising practice. These principles support close review of hiring criteria; the applicable legal analysis depends on the facts, protected trait, employer, and current law.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Measure outcomes at each stage, not just the final hire

For each stage, compare how many people applied, advanced, received scores above the relevant threshold, or were assigned each classification. Break out the categories relevant to the law and use case, and examine intersections where the data allows. A final-hire comparison alone can hide where an outcome gap first appears.

Measure What it tells you Interpretation check
Selection rate The share of a relevant applicant or promotion-candidate group moved forward or assigned a classification. State the denominator, stage, and decision rule. A rate without those details is difficult to interpret.
Impact ratio The category’s selection rate divided by the most-selected category’s rate; for scores, the category’s scoring rate divided by the highest scoring category’s rate. It compares outcomes, but does not establish why they differ or whether the tool predicts job performance.
Scoring rate The share of people in a category whose score is above the sample median. Identify the sample used to calculate the median and the scoring stage being examined.

These definitions follow New York City Rules §§ 5-300 and 5-301. For covered audits, § 5-301 specifies calculations for sex, race/ethnicity, and intersectional categories. It also calls for reporting the number of assessed people in unknown categories. For tools that classify candidates into groups, calculations apply to each group as specified in the rule. For scoring tools, the rule specifies the sample’s median score, category scoring rates, and impact ratios.

An independent auditor may exclude a category comprising less than 2% of audit data from required impact-ratio calculations under § 5-301, but must disclose the justification, the number of applicants, and the rate for that category. Do not treat a small or unstable group comparison as conclusive. Record group sizes, unknown or missing demographic counts, category definitions, comparison population, and thresholds so a reader can judge what the numbers mean.

An impact ratio is one lens on outcomes. It does not establish that a tool uses valid job criteria, predicts performance, provides an accessible assessment, or complies with every applicable law. A favorable metric—or a completed audit—does not by itself prove the system is fair or accurate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find the mechanism behind a gap or error

If the results show an unexpected disparity or an incorrect decision, investigate the pathway that produced it rather than stopping at the aggregate. Compare results by input, criterion, threshold, role, and assessment step. Where feasible, test whether correcting a suspected input or criterion changes the affected applications, then check overall job-related performance as well as the suspected failure.

  • Check whether a data source is incomplete, stale, or differently populated across groups.
  • Check whether a parsing or transformation step drops equivalent qualifications or handles missing data inconsistently.
  • Check whether one criterion or cutoff accounts for the observed change in advancement or score outcomes.
  • Check whether results differ by job family or stage, rather than assuming one role’s findings apply to every use.

These are recommended debugging steps, not procedures expressly required by the NYC rules. If a change is made to a criterion, threshold, data pipeline, or model version, repeat the relevant checks before relying on the changed result.

Check disability access and accommodation routes

The EEOC and Department of Justice warn that employment software can screen out people with disabilities who could perform a job with or without reasonable accommodation. AI assessments can also raise concerns if they elicit disability or medical information. Review how the assessment works for different disabilities, whether candidates can request an accommodation or an alternative process, and whether the system asks for information that creates medical-inquiry concerns.

A human reviewer can catch a parsing or judgment error, but human review is not an automatic remedy for bias. Reviewers should apply consistent, job-related criteria, and the employer should monitor overrides and outcomes rather than assuming that a person in the loop has resolved the problem.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What NYC Local Law 144 requires for covered AEDT use

New York City’s requirements are specific to covered automated employment decision tools (AEDTs) and covered uses. NYC DCWP’s FAQ, dated June 29, 2023, says the law applies when the job is located at an NYC office at least part time, a fully remote job is associated with an NYC office, or the employment agency using the AEDT is located in NYC. The FAQ describes covered use as substantially helping assess or screen applicants at any point in hiring or promotion; scanning a resume bank or contacting someone who has not applied for a specific position is outside the requirement as described there. Check for later city guidance before relying on the FAQ.

For a covered use, NYC Administrative Code § 20-871 generally bars an employer or employment agency from using the AEDT for an employment decision unless the tool had a bias audit no more than one year before use and a summary of the most recent audit, including the distribution date of the tool to which it applies, was posted publicly before use. The law also requires notice to covered NYC-resident candidates at least 10 business days before use. The notice must state that an AEDT will be used and identify the job qualifications and characteristics it will assess. Candidates may request an alternative selection process or accommodation.

The law requires information about the data type, source, and retention to be available on the employer or agency website or, if not already available, within 30 days of a written request, subject to legal exceptions. The city FAQ says Local Law 144 requires an audit but does not itself require a particular action based on the audit’s results; other anti-discrimination laws still apply. These are procedural requirements, not proof that a tool is valid, accessible, or legally safe. Check current official text and obtain qualified employment-law advice for a specific situation; the code host notes its database may not immediately reflect the latest changes.

Document fixes and compare systems consistently

Keep a record of the question investigated, data and model version, metric definitions, findings, reviewer decisions, overrides, and remediation. When comparing tools or processes, use the same decision stage and job family where possible, and examine:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • output type and decision rule;
  • job-relatedness and validation evidence for the criteria;
  • outcome rates by group and intersection, with sample sizes and missing data;
  • data sources, completeness, and age;
  • accessibility and accommodation options;
  • human-review and override patterns; and
  • monitoring after a model, threshold, criterion, or data change, along with jurisdiction-specific audit and notice duties.

Where an independent audit is required, keep its scope and independence clear. An audit can help identify patterns, but it is only one part of a broader review of the hiring process.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.