The CV study · 27 July 2026

Honest, the screening
works. Then one side
buys a tool.

We ran 240 head-to-head screening decisions between two candidates. Told honestly, the screener picked the objectively stronger one in 85 percent of cases. After the weaker candidate ran his CV through an AI tool, he took two out of three decisions he had rightly lost before.

Share of decisions won by the stronger candidate Result Two-sided p
Both honestNeither CV touched by a tool
85%
1.2 × 10−10
Only the weaker one optimisedSame facts, rewritten by an AI tool
34%
0.0049
Both optimisedSame tool on both sides
86%
2.1 × 10−11
025%50%75%100%

80 forced decisions per condition. The middle row is the whole story: the ranking does not collapse because the tool exists, it collapses because only one side has it. Hand the same tool to both and the ranking comes back — 86 percent, statistically indistinguishable from the honest baseline.

240 judgments · 40 CV pairs · 3 conditions · both seating positions · Haiku 4.5
The setup

One interview slot.
Two CVs. Choose.

Scoring CVs one by one drowns in noise: score differences of a few points are indistinguishable from run-to-run variance. A forced choice between two candidates has no such noise. That is why 40 pairs are enough here where scoring studies need thousands.

01 Ground truth by construction

Every profile was generated from a written spec, so who is stronger is a property of the data, not the opinion of a model. Pairs span quality gaps of 1 to 47 points on the study's own index.

40 pairs · 5 gap bands · weighted towards similar candidates
02 The tool may not invent anything

The optimiser could rephrase, restructure, mirror keywords and quantify achievements that were already there. New employers, years, certificates or skills were forbidden. The seven profiles whose optimised version invented facts were thrown out before the run.

43 clean profiles · 7 excluded for fabrication
03 Every pair judged twice, seats swapped

Each pair was put to a separate judging agent in both seating positions, one judgment per agent, trial order shuffled. No agent ever saw two trials, so nothing could be compared across cases.

240 judgments · identical question, byte for byte
The finding

Same person, same facts,
opposite outcome.

The honest condition is not a formality, it is the load-bearing control: without proof that the screener gets it right when nobody cheats, nothing else here would mean anything. Then we changed exactly one thing — the weaker candidate's CV went through the tool.

42of 80 decisions flipped against the honest control
1flipped the other way — the effect has one direction
25of 40 pairs moved, none against the trend
7profiles removed before the run because the tool invented facts
Paired against the control: p = 1 × 10−11 at decision level, 6 × 10−8 at pair level.

And it survives the strictest cut. On 18 pairs built so that every profile appears exactly once — 108 further judgments, no profile reused — 20 of 36 decisions flip, p = 1.9 × 10−6. Same picture, independent arithmetic.

Why it happens

It is not the tool.
It is who has it.

Three obvious objections, three counter-checks built into the run. Each one had to hold, or the finding would have collapsed.

“The optimised CVs are simply longer.”

They are: roughly 392 words against 197. So we ran the condition where both sides optimise. If length were the driver, that condition would be a coin flip.

86% · paired p = 1.0
“The model just picks whoever comes first.”

Every pair was judged in both seating positions. Across all 240 judgments, the first-listed candidate won 116 times.

48% · p = 0.65
“The screener never worked in the first place.”

Honest against honest, it picked the stronger candidate in 85 percent of decisions — and from a gap of 15 index points upwards, in every single one.

20 of 20 · large gaps
The damage does not come from the tool existing. It comes from not everyone having it.

Hand the same tool to both candidates and the ranking is intact again — 86 percent, against 85 in the honest baseline, paired p = 1.0. That is a clean null result, and it is the most useful number on this page: what needs regulating is unequal access, not the technology.

Where it breaks

The tool is worth about
25 index points.

Below that gap it flips two decisions in three. Above it, not a single one. The share is calculated only on decisions the honest control got right — counting the cases where the screener was already wrong would make the tool look harmless.

Quality gap between the two candidates Share of correct decisions the tool flipped Flipped
1–4 pointspractically equal · 11 of 17
65%
5–8 points12 of 16
75%
9–14 points9 of 15
60%
15–24 points10 of 12
83%
25 points and above0 of 8 · the gap holds
0%
025%50%75%100%

Below a gap of 25 points: 42 of 60 decisions flipped. At 25 and above: 0 of 8. Fisher exact, p = 0.00021. A second, near-independent run on different pairs shows the same shape — 75 percent in the 15–24 band, 11 percent above 25.

25 points is a threshold, not a hard edge

The top band rests on 8 decisions from 4 pairs. Zero out of eight is a strong signal that a threshold exists; it does not pin down where exactly it sits. And the quality index is this study's own construction — “25 points” is a quantity inside this project, not an industry scale.

The second break

Optimisation does not lift the weak.
It lifts the almost-good.

43 candidates, five skill levels fixed before the first score. The gain from rewriting the same facts is an inverted U — nothing at the bottom, nothing at the top, and everything in the upper middle. Which is exactly where invitations are decided.

+2.5
below averagen = 10
+12.9
averagen = 13
+18.8
solidn = 10
+14.0
strongn = 7
+2.0
exceptionaln = 3

Average score gain from optimisation, per skill level, on a 120-point scale. At the bottom there is nothing to dress up. At the top the substance was already legible and the score sits near the ceiling. The consequence contradicts the obvious story: rewriting does not help bad candidates sneak past good ones. It helps good candidates catch up with the very good.

Honest, the system separates all five skill levels. With part of the field optimising, two of them merge — 24 points of real difference, gone.

And the order itself survives: rank correlation between the honest and the optimised ranking stays at 0.94. What breaks is not the sequence, it is the threshold. A shortlist cut through a band where two levels have become one is a cut through noise.

What actually moves the judgment

Not who you are.
How your text sounds.

We went looking for a gender effect and separated the two things every earlier attempt had bundled: the writing register and the identity. Same three fact sets, same role, positions balanced, forced pairwise choice. Only one of the two moved anything.

Register · identity held constant

Both candidates male. Only the tone differs.

If the outcome shifts here, it can only come from how the CV is written — assertive and claim-making against plain and factual.

9 of 12for the assertive register · p = 0.146
Identity · register held constant

Both written assertively. Only the person differs.

Name, pronouns and gender-marked memberships swapped, everything else identical. If there is an identity effect, it has to show up here.

5 of 12for the male candidate · p = 0.774
92.9%Substance — what the candidate has actually done
5.1%Writing register — how it is phrased
0.2%Gender — shown wider than it is, or it would be invisible
The same two shares, on their own scalezoomed 16×
Writing register
5.1%
Gender
0.2%

Both slivers from the bar above, re-drawn on a 0–6 percent scale. Full width here is 6 percent of total variance.

Share of score variance explained. Register outweighs gender by a factor of 26. Swap only the gender and the decision stays the same in 5 of 6 cells; swap only the register and it tips in 3 of 6. That reframes the whole question: what is being rewarded is a way of writing — and a way of writing is purchasable. Substance is not.

Neither of these two results is significant on its own

9 of 12 is a direction, not a proof. The register effect showed up twice independently at 9 of 12; pooled that is 18 of 24, p = 0.023 — but the two runs carry different confounds, so it is the agreement in direction that is the argument, not the pooled p-value. A clean answer needs roughly 40 pairs.

What helps

A filter removes
97.6% of the signal.

We forced the model to estimate, on every CV, the probability that the candidate is a woman — “I cannot tell” explicitly not allowed, because an opt-out measures willingness to commit, not how much information is in the text. Then we ran the same test on mechanically redacted CVs.

Unredacted4 CVs · positive control
3.5 99.0
95.5
Mechanically redacted12 CVs · 6 women, 6 men
44.2 46.5
2.3
0255075100
Average estimate for candidates who are actually women Average estimate for candidates who are actually men

Unredacted, the estimates do not overlap at all: 99 and 99 for the women, 5 and 2 for the men. That positive control is what makes the second row readable — without it, “the model sees nothing” could just mean the instrument is blunt. Redaction is rule-based, no model involved: email, salutation, name, phone, gender-marked organisations, pronouns, age. An independent study reports the same order of magnitude, minus 97 percent, for mechanical name redaction.

What we measured

That a mechanical filter removes the signal. 97.6 percent of a 95.5-point spread, down to 2.3.

Measured
What we did not measure

That a fairness instruction in the prompt achieves nothing. The experiment for it was built and never ran. The literature says it does not work; we have no data of our own.

Not measured
What the residue still shows

Even fully anonymised, the model forms a guess — and cites football volunteering, climbing and motorsport as its reasons. That channel survives name redaction.

Small, not zero
“Filter works” here means the signal is gone, not that an injustice was fixed

Welch t = 0.95 on the redacted group, which is not significant, but Cohen d = 0.55 — with 6 against 6 CVs the test is too weak to call the remainder zero. It is small, not proven absent. And in our own setup the identity swap changed no decisions, so there was no measured bias for the filter to repair. The case for filtering stands on a different footing: a signal that never enters the input cannot act on the outcome.

Honest limits

What this study
cannot say.

This section is not fine print. A study that only publishes what supports it is a brochure. Here is every place where a reasonable person should push back.

NOT PROVEN “AI screening is unbiased”

One role, one model, six fact sets, one attribute. We tested gender. Ethnicity, age, national origin and disability are untested — and the literature finds larger effects for ethnicity than for gender.

A RESULT The null finding on identity is a finding, not a failure

With an instrument sensitive enough to detect the honest-versus-optimised split at p = 0.001, swapping name, pronouns and affiliation moved nothing. The confidence interval rules out an identity effect above roughly 31 percentage points — not below it. That is consistent with the strongest field data: 83,000 real applications, no significant gender effect.

NOT MEASURED That a fairness clause in the prompt is useless

Half of that claim is ours and half is borrowed, so we split it. The filter half is measured. The prompt half rests on 2026 preprints, one of which points the other way. We downgraded our own claim rather than let the weaker half inherit the authority of the stronger.

DIRECTION ONLY The register effect is not significant on its own

9 of 12, twice, is a direction. It needs about 40 pairs to become a result. We report it as the most likely explanation, not as proof.

OUR BUILD The optimisation tool is ours, not a product off the shelf

It was built to the rules above and never allowed to invent facts. A consumer service that does lie is an open and probably much uglier question — the seven profiles where our own tool drifted into invention were removed rather than measured.

SCOPE One judge model, one role, English CVs

Every judgment on this page came from Haiku 4.5, one job, English-language CVs, one judgment per trial. External validity is untested, and so is stability across repeat runs.

NOT A RATE 34 percent is not a flip rate for real hiring

Pairs were deliberately weighted towards similar candidates, because the point was to find the curve, not a headline number. Anyone who needs a rate for real applications needs a sample of real applications.

CAVEAT Two models are not a trend

Optimisation gained 11.3 points on the cheap model and 5.8 on the frontier model — the one screening 50,000 applications a year is the cheaper one, which would make the cost decision the fairness decision. Two data points make that a hypothesis, not a law.

We withdrew one of our own headlines

“11 out of 12” was a real measurement of the wrong thing.

The first run compared 12 pairs and found the tool flipping 11 of them. It was not a mistake in arithmetic — those 12 pairs happened to sit at quality gaps of 16 to 22 points, entirely inside the band where the tool is most effective. Scaled to 240 judgments across all five gap bands, the honest picture is the one at the top of this page: a threshold, not a universal reversal. The bigger run makes the claim smaller and the mechanism clearer.

Withdrawn · 12 pairs 11 of 12 flipped True for gaps of 16–22 points. Read as a general rate, it overstated the effect.
Current · 240 judgments 85% → 34% → 86% Two of three flipped below a gap of 25 points, none above it, and the ranking restored when both sides optimise.

This is why we filter,
and why we say so.

A screening system that rewards phrasing is measuring tool access. What we took from this study went straight into how our own screening is built — and into what we are willing to promise about it.

240 judgments · 40 CV pairs · 43 profiles · 5 skill levels · run 26–27 July 2026