An interview loop is a measuring instrument, and most of them are badly calibrated in the same three ways. Here is what we changed across 58 candidates in Tokyo, in the order the changes should be made — step 2 accounts for most of the eleven-day result.
We ran three loop designs over roughly a year of hiring English-speaking engineers into Tokyo teams: a five-stage loop with an unpaid take-home, a four-stage loop with a live coding exercise, and a three-stage loop built around a single paired working session.
Hire quality, judged at six months, was indistinguishable between the five-stage and three-stage designs. Elapsed time was not: six weeks against eleven days. And the five-stage loop lost four candidates to competing offers in the gap between stages three and four.
That is the whole finding. What follows is how to get there without breaking the parts that were working.
Step 1 — Name the two things the loop is testing
Before touching the format, write down what you are actually measuring. Two things, not six. Every additional dimension adds a stage, and every stage costs days.
For most engineering roles on an existing product, the two that matter are: can this person make progress in an unfamiliar codebase, and does this person change their mind when shown evidence. Almost everything else — specific framework knowledge, algorithmic recall, tooling familiarity — either follows from the first or is learnable in weeks.
This step feels like paperwork and is the reason the rest works. When a panellist proposes adding a system design round, the question becomes concrete: which of the two are we not currently measuring? Usually the answer is neither, and the round does not get added.
Step 2 — Compress to three stages and publish the shape
Three stages: a short screen (thirty minutes, with the hiring manager, not a recruiter), one substantial working session (ninety minutes), and a decision conversation (forty-five minutes, covering the role, the team and the candidate’s questions).
The reason this is the single highest-leverage change is scheduling, not interview length. Each additional stage adds five to eight elapsed days once you account for coordinating calendars across a panel — and it adds them at exactly the moment competing offers arrive.
Then publish the shape in the job posting: how many stages, what each contains, and the target elapsed time. This costs nothing, materially improves response rates to outreach, and has a useful internal effect — once the process is public, adding a stage requires a decision rather than a drift.
Step 3 — Replace the take-home with a paired working session
The unpaid multi-hour take-home is the most defended and least defensible part of a typical loop. It filters on available time, not on ability, and it removes exactly the people worth hiring: those with caring responsibilities and those already employed on demanding teams.
The replacement is a ninety-minute paired session on a small, real problem from your codebase — a genuine defect works best. The candidate drives; an engineer sits with them and answers questions as a colleague would.
What you observe is not the solution. It is the process: how someone approaches unfamiliar code, what they check first, whether they read error messages or guess, whether they ask when stuck or spiral, and how they respond when told their hypothesis is wrong. None of that is recoverable from a submitted repository, which is why take-home reviews so often end in a shrug.
One rule makes or breaks it: the session must be collaborative, not observational. An engineer sitting in silence taking notes turns a working session into a performance, and you are back to measuring composure under artificial pressure.
Step 4 — Score independently before anyone discusses
The most common defect in an otherwise sound loop is the debrief that starts with a conversation. Whoever speaks first, or most confidently, sets the frame, and the rest of the panel adjusts toward it without noticing.
The fix is procedural and takes no extra time: every interviewer submits a written score and a one-paragraph justification before the debrief opens. Submissions are visible only once all are in.
Then track how often people change their score after discussion. A small amount of movement is healthy — that is what debriefs are for. If half the panel moves, you are not measuring the candidate; you are measuring the room. In our first design that number was above forty per cent, which meant a substantial part of our hiring signal was whichever senior engineer happened to speak first.
Step 5 — Design the panel for mixed language fluency
In Tokyo teams the panel is frequently mixed: some interviewers fully comfortable conducting a technical session in English, others much stronger in Japanese. Left undesigned, one failure mode follows reliably — a highly capable engineer contributes little because the session runs in their weaker language, and their assessment quietly carries less weight than it deserves.
Two practical fixes. Let panellists submit written questions in advance for a fluent colleague to pose, so their technical judgement enters the loop regardless of who is speaking. And allow written scoring in whichever language the panellist prefers, translating only at the summary stage.
The thing to avoid is assessing a candidate on language performance when the role does not require it. Decide in advance what level the job genuinely needs — and if the answer is that the team works in English, then testing anything beyond that is measuring something you have already decided does not matter.
Want candidates who reach the loop already vetted?
We screen English-speaking engineers for Tokyo teams on the two things that matter — progress in unfamiliar code, and response to evidence — so your three stages do less work. Shortlist in 7 days.
Start now — see vetted engineersStep 6 — Decide within 48 hours, or change the process
The debrief happens within two working days of the final stage, and a decision is made in that meeting. Not “let’s see who else is in the pipeline”.
Comparing a candidate against hypothetical future candidates is the most expensive habit in hiring, because the comparison is always favourable to the imaginary person. The question at the debrief is binary: would this person be a clear improvement to the team? If yes, make the offer. If no, decline and say so quickly.
If you genuinely cannot decide within 48 hours, the problem is upstream. Either the loop did not test the right things, or the role is not defined well enough to evaluate anyone against. Both are fixable; waiting is not a fix.
Step 7 — Track two numbers and recalibrate quarterly
Loops drift. A stage gets added after one bad hire, a round gets extended after an awkward session, and eighteen months later you are back at five stages without anyone having decided to be.
| Number | What it measures | Threshold |
|---|---|---|
| Elapsed days, first contact to offer | Competitive exposure, not just efficiency | Above 15 days — investigate |
| Share of interviewers changing score after debrief | Whether you measure the candidate or the room | Above 25 % — score before discussing |
| Number of stages | Drift | Above 3 — justify each one in writing |
| Offers declined on process experience | Damage already done | Any — ask the candidate why |
Elapsed time deserves emphasis because it is misread as an efficiency metric. In a market where strong candidates are in three processes at once, it is a quality metric: a candidate who waits six weeks is not waiting, they are being interviewed elsewhere.
The three mistakes we made along the way
Adding a stage after a bad hire. The instinct is that more filtering would have caught it. Reviewing the case honestly, the signal had been present in the existing loop and had been overridden in the debrief. We added a stage that cost every subsequent candidate a week and fixed nothing.
Letting the recruiter run the first stage. A thirty-minute screen with the hiring manager is worth more than an hour with anyone else, because only the hiring manager knows which imperfection matters. When we moved this back, our pass-through rate to the working session fell and our offer rate rose.
Treating the paired session as an exam. Our first version had an engineer observing silently and taking notes. Candidates froze, and we were back to measuring composure. Making it explicitly collaborative — questions answered as a colleague would answer them — changed the quality of the signal completely.
The same loop travels reasonably well but not perfectly. Colleagues at HireDeveloper.sg find the paired session works equally well in Singapore but that three stages is sometimes one too few for regulated employers, while the team at HireDeveloper.ae reports that in Dubai the binding constraint is usually the debrief scheduling rather than the number of stages. If you are still deciding what the team will build, our guides on how to build an edtech platform in Japan and how to build an e-commerce platform cover the scoping work that determines what you should be testing for in the first place.
Frequently asked questions
How many interview stages should a technical loop have in Tokyo?
Three, and the case for a fourth is almost always weaker than it appears in the meeting where someone proposes it. Each additional stage adds roughly five to eight days of elapsed time once scheduling across a panel’s calendars is included, and it adds them at precisely the point where competing offers are landing. In our comparison across 58 candidates, the five-stage loop and the three-stage loop produced indistinguishable hire quality judged at six months, while the five-stage loop lost four candidates to competing offers in the gap between stages three and four. The three stages that earn their place are a short screen run by the hiring manager, one substantial working session, and a decision conversation covering the role and the candidate’s own questions.
Are take-home exercises worth it for engineering hires in Japan?
Rarely, and least of all for experienced candidates. An unpaid multi-hour take-home is a filter on available time rather than on ability, and it systematically removes people with caring responsibilities and those already employed on demanding teams — precisely the population most worth hiring. A ninety-minute paired working session on a real, small problem from your own codebase produces more signal because you observe the process rather than the artefact: how someone reads unfamiliar code, what they check first, whether they read error messages or guess, and how they respond when told their hypothesis is wrong. None of that is recoverable from a submitted repository, which is why take-home reviews so often end inconclusively. One condition matters: the session must be collaborative rather than observational, or it becomes a performance test.
How do you interview when the panel has mixed English fluency?
Design for it explicitly instead of hoping it works out. The reliable failure mode is that a highly capable interviewer contributes little because the session runs in their weaker language, and their assessment quietly carries less weight than it should — which means you are discarding technical judgement for reasons that have nothing to do with technical judgement. Two practical fixes work well: let panellists submit written questions in advance for a fluent colleague to pose, and allow written scoring in whichever language the panellist prefers, translating only at the summary stage. What must not happen is a candidate being assessed on language performance when the role does not require it. Decide beforehand what level the job genuinely needs, and test only that.
What are the signs that an interview loop is broken?
Two numbers reveal it, and both are cheap to track. The first is elapsed days from first contact to offer, which in a competitive market is a quality measure rather than merely an efficiency one: a candidate who waits six weeks is not waiting, they are being interviewed by three other companies during those weeks. Above roughly fifteen days, look for scheduling gaps rather than long interviews, because that is almost always where the time goes. The second is the proportion of interviewers who change their score after group discussion, which should stay low; a small amount of movement is what debriefs are for, but if half the panel moves you are measuring social dynamics rather than the candidate. In our first loop design that figure exceeded forty per cent, which meant our hiring signal was substantially determined by whoever spoke first.
Eleven days is achievable. Six weeks is a choice.
Send us your current loop. We will tell you which stage to remove first and supply pre-vetted English-speaking engineers to run the new one on.
Start now