The reference check is the new skills test
Key takeaways
- When a senior candidate’s work is AI-assisted, a clean portfolio only proves that they, or a tool, could produce a clean output, it can no longer attribute their judgment.
- The hard part of the hire is now verification: separating the person from the assistance in front of you. The interview and the deck can no longer do that on their own.
- The reference check, run properly, is where you recover the signal, where you reconstruct the decisions the candidate made under ambiguity, before the polish went on.
The portfolio stopped proving whose judgment it was.
For most of hiring history, a senior candidate’s work product was a reliable proxy for the person. A strategy memo, a system design, a board deck, a piece of analysis, these carried the fingerprints of the mind that made them. You could read a document and infer how someone thinks. That inference is what a portfolio was for.
AI assistance has loosened the link between the output and the author. A strong first draft of almost anything is now available to anyone who can write a good prompt, and the polish that used to signal seniority is the cheapest part to buy. A candidate can present work that’s genuinely good and genuinely theirs, or good and mostly the tool’s, and the artefact often looks the same either way. The signal the portfolio used to carry has thinned.
AI-assisted work is often better, and a senior person who uses these tools well is doing exactly what you want them to do on the job. The problem is narrower and harder: the work in front of you no longer tells you which of the two you’re looking at. The verification step, deciding whether the judgment on display belongs to the candidate, used to be nearly free. It’s become the expensive part of the hire.
The interview and the take-home have the same blind spot.
The instinct is to fix this in the interview. Give a harder case, a live exercise, a take-home with a tight clock. Each of these has a version of the same weakness. A take-home can be run through the same tools the candidate would use to pad a portfolio, so it tests access, not authorship. A live case tests how someone performs under a stopwatch in a room, which is a real skill but a narrow one, and not the skill most senior roles are actually paid for.
What these formats measure well is capability at a moment. What they measure poorly is the trail of decisions a person made over months on real work, with incomplete information, competing stakeholders and no clean answer. That trail is where senior judgment actually lives, and it’s exactly the part a polished artefact compresses out of view. You see the conclusion. You don’t see the 7 forks the candidate navigated to reach it, or the 2 they got wrong first.
So the interview confirms that the person is articulate and quick, which the portfolio already suggested. It rarely tells you whether the judgment behind the portfolio was theirs. To recover that, you have to talk to the people who watched them decide.
A candidate can rehearse the story of a good decision. What they can’t rehearse is the version their former manager tells about the same week, that’s where you find out who actually made the call.
Rajesh Pandian · Chief of Staff, Tech and Strategy, Recruise
The reference check has to become a structured probe.
Most reference checks are worthless for this, because they were designed to confirm facts a candidate already gave you. Dates, title, “would you rehire.” The referee is chosen by the candidate, primed to be positive, and asked questions with obvious right answers. That format can’t surface who owned a decision, because it never asks about a decision.
The reference check that recovers judgment does two things differently. It goes partly off the candidate’s list, a back-channel to a former peer or skip-level who saw the work up close but wasn’t hand-picked to praise it. And it replaces the confirmation script with scenario probes: specific, retrospective questions about a hard call the candidate was part of, asked of someone who was in the room. The aim is to reconstruct the same episode from two sides and see whether they match.
Done well, this is the assessment stage. You’re using people who watched the candidate decide over time to do what an artefact can no longer do, attribute the judgment to a person. When a former manager describes a messy call the candidate led, names the constraint they were under, and can tell you what the candidate got wrong before they got it right, you’re hearing something no deck and no take-home can fake.
| Dimension | Confirmation reference (the default) | Judgment probe (what to run) |
|---|---|---|
| Purpose | Verify facts the candidate supplied | Attribute the judgment behind the work to the person |
| Who you call | Referees the candidate chose | Candidate’s list plus a back-channel peer or skip-level who saw the work |
| What you ask | Dates, title, “would you rehire” | A specific hard call: what the constraint was, what the candidate decided, what they got wrong first |
| What good looks like | Consistent, positive, unsurprising | Specific, textured, willing to name a real limitation |
| What a bad signal looks like | Rarely surfaces one | Vague ownership, credit that keeps sliding to “the team,” a story that doesn’t match the candidate’s |
| What it actually tests | That the resume is accurate | Whether the decisions in the portfolio were the candidate’s |
What to ask the people who watched them decide.
The questions that work share a shape: they name a real episode and ask the referee to reconstruct it. Skip “how strong were they analytically.” Ask about a specific decision the candidate owned, and let the detail either arrive or fail to.
Useful probes for a senior AI-augmented role include: walk me through a call this person made where the data was thin and they had to commit anyway, what did they choose, and what did it cost them if they were wrong. Tell me about a piece of work they presented that looked finished; how much of the thinking underneath was theirs, and how would you know. When they used AI tools or automation, where did their own judgment sit, what did they override the tool on. And the one that separates a real reference from a courtesy: what did this person get wrong in the time you worked together, and how did they handle being wrong.
A referee who was actually there answers these with texture: a constraint, a date, a disagreement, a correction. A referee who wasn’t, or a story that was authored rather than lived, gets vague exactly here. The tell is a smooth answer with no friction in it, where every hard edge has been sanded off. When ownership keeps dissolving into “the team did,” or the referee can describe the outcome but not a single fork on the way to it, you’re hearing about work the candidate was near, not work they led.
The failure mode is trusting a portfolio that reads too clean.
The mistake this pattern guards against is specific. A senior req comes in, a candidate presents a portfolio that’s articulate, well-structured and impressive, the interview goes smoothly, and the team reads the polish as proof of the person. Six months in, the hire can’t reproduce under pressure the judgment the portfolio implied, because the portfolio was assisted in ways the process never checked. The cost of that lands at the senior level, where a single wrong hire reshapes a team’s direction.
The counter-intuitive part is that a portfolio which reads too clean should now raise a question. Real senior work carries the residue of the constraints it was made under: the compromise, the thing that didn’t quite work, the decision made with worse information than anyone wanted. Output that has none of that texture is either exceptional or synthetic, and from the artefact alone you can’t tell which. The reference probe is how you tell.
None of this means treating AI-assisted work as a red flag. A candidate who uses these tools fluently is showing you a skill the role needs. The discipline is to stop reading the artefact as evidence of the author, and to move the weight of the decision onto the stage that can still separate them. For senior hires, that stage is a reference check run as an assessment, and the teams that get this right have already made the shift.
Frequently Asked Questions
Why is a portfolio less reliable when the work is AI-assisted?
Because the artefact no longer tells you whose judgment produced it. AI tools make a strong, polished draft available to anyone with a good prompt, so the qualities that used to signal seniority, structure, fluency, a clean conclusion, are the cheapest part to acquire. A candidate can present work that’s genuinely theirs, or mostly the tool’s, and the document often looks the same either way. The output can still be excellent; it just stops being evidence of the author’s decisions.
Can’t a harder interview or take-home solve this instead?
Only partly. A take-home can be run through the same tools a candidate would use to pad a portfolio, so it tests access rather than authorship. A live case tests performance under a stopwatch, which is a narrow skill and rarely the one a senior role is paid for. Both formats measure capability at a single moment; neither reconstructs the trail of decisions a person made over months on real work, which is where senior judgment actually lives. That trail is best recovered from people who watched the candidate decide.
What makes a reference check able to surface judgment?
Two changes to how it’s run. First, go partly off the candidate’s chosen list, a back-channel to a former peer or skip-level who saw the work up close but wasn’t hand-picked to praise it. Second, replace the confirmation script with scenario probes: specific retrospective questions about a hard call the candidate was part of, asked of someone who was in the room. The aim is to reconstruct the same episode from two sides and check whether they match, which attributes the judgment to a person rather than to an artefact.
What should we actually ask a referee for a senior AI-augmented hire?
Ask about specific episodes, not ratings. Walk me through a call this person made when the data was thin. When they presented work that looked finished, how much of the thinking underneath was theirs, and how would you know. When they used AI tools, where did their own judgment sit and what did they override. And what did they get wrong, and how did they handle it. A referee who was there answers with texture: a constraint, a date, a correction. A smooth answer with no friction in it, or ownership that keeps sliding to “the team,” is the signal that the candidate was near the work rather than leading it.
One hiring pattern worth knowing, every ten days.
The Mandate Desk is our read on the senior GCC talent market — one signal that moved, the read behind it, and one thing worth doing. Written from live placement data.
More from Recruise Insights.
AI in the Workplace Most GCC AI pilots stall at the same place. Here's why.
We've mapped GenAI rollouts across the GCCs we work with. The failure mode is consistent, and it isn't the technology.
AI in the Workplace GenAI and AI/ML are two different hires. Most JDs miss it.
Three distinct profiles are being confused for one. The mis-hire rate is predictable. The fix is upstream of recruitment.
AI in the Workplace The roles AI is quietly creating inside GCCs
Not the ones in the headlines. The senior seats that appear when a function moves from doing the work to governing it.
Have a senior seat to fill?
Tell us the mandate — the role, the level, the market. We’ll come back with what the market is really doing on it, and how we’d run the search.