To create an AI interview, turn a clear role profile into a structured interview guide, permitted follow-up questions and a reviewable scorecard. AI can suggest questions or conduct the initial conversation. People remain responsible for criteria, approval, edge cases and the decision about what happens next.
What do you need before creating an AI interview?
This guide covers two approaches: using a general language model to draft a guide for a human interviewer, or configuring a specialist platform for an asynchronous chat, voice or video interview. The quality logic is the same in both cases.
Do not start from the job advertisement alone. Prepare a concise role profile that separates essential requirements, learnable capabilities and mere preferences. Add real work situations in which the relevant skills become visible.
| Building block | Useful input | Weak input |
|---|---|---|
| Role outcome | What should the person deliver after onboarding? | A generic task list |
| Essential criterion | An observable requirement with a reason | “Top talent” or “culture fit” |
| Work sample | A real decision or typical situation | A puzzle unrelated to the job |
| Evidence | A concrete example, approach and result | A numeric self-rating without an example |
| Exclusion | A legally and professionally reviewed minimum condition | An assumption about personality |
The hub on AI interviews and voice recruiting helps place the guide, channel and workflow stage.
Step 1: define the purpose of the interview
Write one sentence that states which decision the interview prepares. A useful goal might be: “The initial interview clarifies shift availability, English for customer conversations and two concrete examples of de-escalation.” A weak goal is: “The AI finds the best candidates.”
The purpose limits questions and data collection. A short screen should not pretend to be a complete assessment. It gathers the information missing for the next human step.
Step 2: translate requirements into observable criteria
Phrase each criterion so multiple reviewers can classify the same answer similarly. “Strong communicator” is too open. “Explains the next step to an upset customer, sets a boundary and records the commitment” is observable.
Limit a short initial interview to a small number of criteria. When everything matters, the scorecard becomes arbitrary. Mark each criterion as essential, developmental or supplementary evidence.
Step 3: design core questions with the same starting point
Every candidate should receive the same core questions. This improves comparability and prevents spontaneous rapport from determining interview depth. Good questions ask for a concrete example or a reasoned choice.
- Experience: “Describe a situation in which …”
- Approach: “How would you handle the following situation …?”
- Outcome: “How did you know your approach had worked?”
- Reflection: “What would you do differently now?”
Avoid questions that reveal the desired answer. Questions about protected characteristics or private circumstances do not belong in an automated guide.
Where automated evaluation crosses into an automated decision is a legal question, not a product one — we cover it in can AI reject candidates automatically.
Step 4: define permitted follow-up questions
A good AI interview is neither a rigid form nor an unrestricted conversation. Define a small fixed set of permitted follow-up types: ask for an example, clarify the person’s own role, and request an outcome or learning. Limit repetitions and total duration.
Set stop and handoff rules as well. If someone reports a technical problem, does not understand a question or shares sensitive personal information, the system should not improvise. It needs an alternative channel and a human contact.
The same structure applies to any structured conversation, whether or not AI is involved. Our guide to building an interview guide covers how to derive criteria from a role profile.
Step 5: create an evidence-based scorecard
A scorecard separates the criterion, observed evidence and judgement. It should never produce only a number. Reviewers need a path back to the relevant transcript passage or response.
| Level | Description | Customer de-escalation example |
|---|---|---|
| No evidence | No usable response | No example or irrelevant answer |
| Partial evidence | Some relevant indicators | Calms the customer but gives no next step |
| Solid evidence | The criterion is supported clearly | Clarifies issue, explains next step and documents it |
| Strong evidence | Evidence covers approach, result and reflection | Adds a clear boundary, ownership and verified outcome |
These levels are not universal. Adapt them to each role and criterion. A total score must not hide an essential condition: proof of a required work authorisation cannot be offset by points elsewhere.
Step 6: use a controlled prompt
You can use a general language model to produce a first draft. Do not enter real applicant data until the use, contract and privacy basis have been reviewed. A useful prompt states purpose, role, criteria, output format and explicit prohibitions.
Prompt template:
You are an interview designer. Create a structured initial interview guide for [role]. Purpose: [decision]. Essential criteria: [list]. For each criterion, write one open core question and a small fixed set of neutral follow-ups. Create a scorecard with clearly named evidence levels, visible evidence anchors and “insufficient information” as a separate option. Do not ask about private or protected characteristics. Do not make a hiring decision. Output questions, follow-ups, scorecard and review notes separately.
Review every output professionally. A model can misread requirements, generate duplicate questions or propose criteria that are irrelevant to the role.
Step 7: configure channel and candidate experience
In a specialist platform, choose between a web form, chat, video, voice, messaging or phone. The best channel is the one your audience can use reliably. Browser video may fit office roles, while phone or messaging can be more accessible for non-desk candidates.
The invitation should explain purpose, expected duration, channel, evaluation, next step and contact person. Offer an alternative for technical, language, personal or accessibility barriers. The candidate portal illustrates how several routes can join one consistent process.
Step 8: test the interview before launch
Test at least the following cases: an ideal answer, a weak but honest answer, an unexpected career history, a technical failure and a response containing sensitive information. Use fictional test profiles, not real applicant data.
- Ask multiple reviewers to score the same answers independently.
- Compare score, reasoning and missing information.
- Test transcription, accents and domain terms.
- Test stopping, resuming and an alternative channel.
- Correct the guide and rubric before sending real invitations.
If reviewers repeatedly disagree, the likely cause is an unclear criterion or scale. Adding more AI does not fix that design problem.
What effective human review actually requires
Human oversight only works when the reviewer sees more than a recommendation. They need access to the relevant response, its criterion, the evidence anchor and any uncertainty in the output. They must also be able to correct the system and reject its proposed next step.
Keep three outcomes separate: the candidate did not provide evidence, the question was not understood, or the interview never collected the required information. They should not receive the same assessment. Record the reviewer’s own reasoning as well. Otherwise, nominal human control becomes a rubber stamp on an automated summary.
Review is not an opportunity to introduce new selection criteria after the interview. If relevant information is missing, carry it forward as an open question for the next human conversation. If a requirement changes permanently, revise and retest the role profile, guide and scorecard together.
What does a complete interview blueprint contain?
| Stage | Content | Owner |
|---|---|---|
| Invitation | Purpose, duration, data notice and alternative | Recruiting |
| Opening | Process and agreement to the channel | System and recruiting |
| Core questions | The same job-related questions | Interview owner |
| Follow-ups | Only approved clarification | Configuration |
| Scorecard | Criterion, evidence, judgement and uncertainty | Human reviewer |
| Handoff | Transcript, summary and open points into the ATS | Integration |
| Decision | Next step and documented reason | Responsible person |
How do you monitor quality after the pilot?
Passing a test run is not permanent approval. Transcription, model behavior, role requirements and candidate journeys can change. Assign clear owners for approval, error analysis and shutdown. Treat every change to the prompt, model, channel or evidence anchors as a new version.
Do not monitor completion alone. Review transcript corrections, irrelevant follow-ups, loops, abandonment, use of alternative routes, disagreement between system output and human review, and candidate feedback. Look for patterns across languages, devices and candidate groups without drawing premature conclusions from small samples.
Review anomalies with recruiting and the hiring team. A technical failure requires a different remedy from a vague criterion. Stop an affected version when an error could influence selection decisions. Return it to use only after the correction is documented and the relevant tests have passed again.
Which legal questions need review?
The EU AI Act places certain AI systems used for recruiting and selection in the high-risk category in Annex III. Whether a product falls within that category depends on function and use. Review provider and deployer roles, human oversight, documentation, data quality and error handling.
The GDPR applies to personal-data processing. Relevant areas include transparency, data minimisation, storage limitation and potentially Article 22 on solely automated decisions. For German operations, Section 87, paragraph 1, number 6 of the Works Constitution Act can be relevant to technical devices intended to monitor behavior or performance. The guide to AI recruiting and works councils explains the practical involvement using primary sources.
When is an AI voice interview useful?
A voice interview can help when spoken answers add relevant evidence and scheduling is the bottleneck. Sprad supports structured conversations through a portal, messaging or phone; the voice interview overview explains the flow. The boundary remains clear: software collects and organises evidence while people own the next step.
Frequently asked questions about creating AI interviews
Can ChatGPT create an interview guide?
Yes. It can draft questions, follow-ups and a scorecard. You must review job relevance, fairness, privacy and professional accuracy. Do not enter real applicant data without an approved basis.
How many questions should an AI interview contain?
The number follows the purpose. A short screen benefits from a few core questions rather than a comprehensive catalogue. Each question should serve one relevant criterion.
Can AI ask follow-up questions?
Yes, when the permitted types and limits are defined. Useful follow-ups request an example, the candidate’s role, an outcome or clarification. Unrestricted probing is inappropriate.
Should AI reject candidates automatically?
Automated decisions can create substantial legal and quality risks. Use human review, documented criteria and a correction route, especially for exclusions and uncertain evidence.
What does a good output look like?
A useful interview record connects each criterion to the relevant answer, judgement and remaining uncertainty. A general fit score without source evidence is not enough.
