No items found.

Bias in AI interviews: what you can test, what vendors leave out

By Jürgen Ulbrich

Bias in AI interviews is testable only when an employer compares the same job-relevant evidence under different conditions, reviews outcomes and drop-off by meaningful segments, and keeps a human accountable for consequential decisions. AI can make an interview more consistent than an exhausted first screener. It can also turn an unseen transcription error, a narrow prompt, or an inaccessible channel into a repeatable barrier.

The practical question is not whether a product uses AI. It is what the system changes in the hiring decision, what happens to an answer between speech and recommendation, and who has a worse experience in the real funnel. This overview of AI interviews and voice recruiting provides useful context; procurement needs an evidence plan on top of it.

What does fairness in an AI interview mean?

Bias is a systematic distortion: candidates with equally relevant evidence receive different transcripts, scores, recommendations, or process opportunities for reasons that are not necessary for the job. Fairness does not mean every group must have identical outcomes. It means that equivalent job-relevant evidence is treated equivalently, and that any difference can be explained, tied to the role, and challenged.

This is a process property, not a badge or a single model metric. A good aggregate score may hide poorer transcription for a particular accent, a question that rewards familiarity with a local office norm, or a higher abandonment rate in one channel. Looking only at the final rank makes the source of the difference invisible.

Likewise, one difference in a small sample is an investigation signal, not proof of unlawful discrimination. A clean-looking small sample is not proof that the system is fair. A responsible hiring team has to hold both ideas at once.

Where bias can enter an AI interview

Language-model assumptions can be mistaken for job evidence

A language model recognises patterns in language. Unless the interview guide and rubric are tightly tied to the role, it may treat familiarity with a particular wording, communication style, or career narrative as a sign of competence. The risk is not limited to the base model: it also sits in the definition of a strong answer and in the follow-up questions the system chooses to ask.

Speech recognition changes the material being scored

In a voice process, an answer is commonly converted to text before it is assessed. If the system transcribes an accent, dialect, multilingual answer, or ordinary phone audio differently, the later evaluator is working from altered evidence. That does not establish that every speech system disadvantages a group. It does mean that a voice interview workflow should be tested with the languages and speech patterns relevant to the actual candidate population.

Question design can reward an unstated norm

A question about motivating a team is not inherently unfair. It becomes risky when it is vague, assumes an unstated cultural convention, or rewards a preferred form of self-presentation more than the experience needed for the role. Follow-ups matter too: a system that probes one type of answer deeply and moves quickly past another is not gathering comparable evidence.

Scoring data can carry forward earlier preferences

When a scoring layer is calibrated on historic hiring, interview, or performance judgments, past human preferences may be present in the labels. Removing an explicitly protected field does not automatically remove all related signals. The same question applies to AI-assisted CV screening: a field can be absent while proxies remain in the information used for an outcome.

A channel can exclude candidates before scoring begins

A mandatory phone or video conversation assumes suitable technology, a workable environment, and the ability to use that channel at that moment. Candidates who cannot participate or decide to leave never appear in a score analysis. A fair process offers an equivalent alternative, a reachable human contact, and appropriate accommodations instead of treating access to one channel as evidence of suitability. A candidate portal with multiple process options helps only when that choice is real in practice.

What an employer can test: a four-signal field audit

The most useful information gain does not come from one generic fairness score. It comes from four separate signals that locate whether a problem sits in scoring, speech processing, process access, or the live decision path. Before testing, fix the role, questions, rubric, system versions, and observation period. If sensitive group data is used, collect it lawfully, minimally, and with appropriate privacy review.

  1. Review samples by segment: For the same role, take random completed cases and compare scores, movement to human review, and progression by relevant segment. Segments may include offered channel, interview language, or speech variant; protected characteristics should never be collected casually. Check that the job-relevant evidence and process stage are comparable before interpreting a difference.
  2. Run the same answer through different voices: Prepare one answer and have it spoken by several voices relevant to your candidate population, including different accents or dialects where appropriate. Send every recording through the complete workflow. If the transcript, follow-ups, score, or recommendation changes, you have a reproducible signal rather than an assumption.
  3. Isolate the point of failure: Submit the canonical written answer directly to the scoring layer as well as comparing it with the generated transcripts. If a score is stable for identical text but varies after audio processing, investigate speech recognition first. If identical text itself produces different outcomes, inspect the guide, prompt, and scoring logic.
  4. Track outcome distributions and drop-off separately: Do not compare averages alone. Review the distribution of scores, recommendations, and human overrides, then measure who leaves before starting, during a particular question, or after a channel change. A high drop-off rate is not automatic proof of bias, but it is a material clue that access or communication may be failing.

Do not collapse these four signals into a reassuring composite number. Balanced final scores can coexist with poor transcription; low abandonment can coexist with a faulty rubric. Each signal asks a different diagnostic question, and that is why the audit is more useful than a vendor slogan.

What evidence should a vendor provide?

Ask for evidence for the configuration you plan to deploy, not a generic product demonstration. In the AI interview and voice tool category, an important distinction is whether a system merely documents a conversation or also scores, filters, ranks, or recommends candidates.

  • Purpose and decision depth: a written account of whether the system asks questions, summarises, scores, prioritises, or recommends a hiring action, including every automatic threshold.
  • Version-bound validation: the model, prompt, and speech-recognition versions; test date; languages and speech patterns covered; sample description; metrics; and known limitations. An undated white paper is not enough.
  • Scoring logic: the role-specific guide and rubric, treatment of uncertainty, examples of reasoned scores, and an explanation of which inputs can influence a recommendation.
  • Paired and production-test results: same answer across voices, transcript comparisons, outcome distributions, and drop-off data. Ask to see failed tests, corrective actions, and retests as well as positive findings.
  • Operational control: audit logs, change records, monitoring after updates, a route for human override, and an equivalent route for people who cannot or do not wish to use the AI channel.
  • Governance material: data provenance and quality, technical documentation, risk management, privacy documentation, and a named owner for candidate complaints or incidents.

A claim that a model is unbiased is not falsifiable if it omits the groups, task, metric, test set, product version, and time window. A testable claim is narrower: for this version, this role, these languages, and this evaluation, these differences were observed, these limitations remain, and this is how they are monitored. That is not a universal clearance. It is evidence that can be repeated and scrutinised.

AGG and the EU AI Act answer different questions

For hiring in Germany, the General Equal Treatment Act, or AGG, covers selection criteria and hiring conditions. It prohibits disadvantage on the listed grounds within its scope and addresses apparently neutral practices that can cause indirect disadvantage; see the official text of sections 1, 2, 3, and 7 AGG. A vendor assessment does not transfer the employer’s responsibility. Employment-law review should cover the selection criteria, documentation, and response to concerning findings.

Under its intended purpose, a system used to evaluate candidates in recruitment or selection falls within the employment and access-to-self-employment area in Annex III of the EU AI Act. As of 20 August 2026, the consolidated text sets the application date for Chapter III, Sections 1 to 3 for Annex III high-risk systems at 2 December 2027; the regulation otherwise applies from 2 August 2026. See the consolidated EU AI Act, including Article 113 and Annex III. The precise classification depends on intended use and statutory exceptions, not on a vendor’s marketing label.

The later date is not a reason to delay the work. Risk management, data governance, technical documentation, logging, and human oversight are not credible last-minute additions. AGG compliance and AI Act readiness are also not substitutes for one another: a formal compliance pack does not prove a non-discriminatory selection process, and a strong field audit does not replace legal analysis. For US roles, this EU-and-Germany-focused checklist is not a complete legal framework; obtain advice for the applicable federal, state, and local rules.

The limit of this approach

No audit can prove fairness for every individual, future model version, role, or context. Small or skewed samples may miss rare failures, and even a careful segment analysis can overlook important individual circumstances. The four-signal audit is a way to find and narrow risks, not a certification and not legal advice.

The responsible conclusion is therefore not that bias has been eliminated. It is that the intended use is known, relevant risks were tested before launch, live signals are monitored, and the process will be paused or changed when evidence warrants it. That standard is more useful to candidates and hiring teams than an untestable neutrality promise.

Frequently asked questions about bias in AI interviews

Can an AI interview be fairer than a human first interview?

It can be, when candidates receive the same role-relevant questions, the assessment is explainable, and people review consequential recommendations. It is not automatically fair because software is involved. Guide quality, transcription, scoring, and access to alternatives all matter.

Is removing protected fields from an application enough?

No. Language, career history, availability, location signals, or technical conditions can act as proxies. Test the whole workflow and document why each assessment criterion is necessary for the specific role.

How often should the audit be repeated?

Test before launch, after a material change to the model, prompt, rubric, or speech component, and at defined operating intervals. Repeat it whenever patterns in scores, overrides, or abandonment look concerning. The vendor must make version changes visible enough for this to be possible.

Should every candidate have to complete the AI interview?

A mandatory single channel increases the risk that access is confused with suitability. Offer an equivalent alternative and explain clearly what the system does and who makes the decision. The alternative should not lead to slower or worse treatment.

What is the most important vendor question?

Do not start by asking whether the model is bias-free. Ask which identical-answer cases were tested across voices and languages, what differences appeared, which version the result covers, and how the vendor monitors changes after updates. A credible answer includes documents, limitations, and an escalation path.

Jürgen Ulbrich

CEO & Co-Founder of Sprad

Jürgen Ulbrich has more than a decade of experience in developing and leading high-performing teams and companies. As an expert in employee referral programs as well as feedback and performance processes, Jürgen has helped over 100 organizations optimize their talent acquisition and development strategies.

Free Templates &Downloads

Become part of the community in just 26 seconds and get free access to over 100 resources, templates, and guides.

No items found.

The People Powered HR Community is for HR professionals who put people at the center of their HR and recruiting work. Together, let’s turn our shared conviction into a movement that transforms the world of HR.

Similar Posts