How to share customer interviews across a language gap

A company spends four months on customer interviews in Brazil, forty of them, run properly by a local team. The findings arrive at head office as a nine slide deck. Six months later the product ships with a payment flow that the Brazilian team had flagged in interview eleven, in a sentence that never made it into any slide.

Nobody was careless. The deck was a good deck. What happened is that a body of research in Portuguese passed through a single translating human, and that human had to decide what mattered before anyone else got to look.

Why customer interviews lose the most in translation

Because the compression and the translation happen in the same step, performed by the same person, with no way for anyone downstream to check either one. A summary is already a set of judgements about what is important. When it is also the only version that exists in the reader’s language, those judgements become invisible and permanent. The reader cannot tell the difference between something a customer never said and something the summariser did not think worth carrying.

Separating the two steps is most of the fix, and it is cheaper than it sounds.

What a summary strips out

Three kinds of thing reliably fail to survive, and they are the three kinds you were paying for.

The first is the customer’s own wording. People describe their problems in vocabulary that does not match the vocabulary of the company selling to them, and that mismatch is often the finding. Once a summariser has translated boleto anxiety into concerns about payment friction, the specific thing is gone and only the category remains.

The second is hesitation and contradiction. Jakob Nielsen’s guidance on interviewing users is blunt about the limits of what people report: what users say and what they do differ, memory is unreliable, and people readily invent an opinion when asked. A transcript preserves the moment where someone says one thing and then quietly walks it back. A summary records the tidier version.

The third is frequency. A summary that says several participants mentioned delivery times is not checkable. A set of transcripts is. Seventeen of forty is a different fact from four of forty, and only one of them justifies changing a roadmap.

The chain of custody problem

It helps to look at how a finding actually travels.

A customer says something in Portuguese in minute thirty one. A researcher hears it and writes a note. The note goes into a synthesis document, in English, alongside notes from thirty nine other sessions. The synthesis becomes a slide. The slide becomes a bullet in a strategy memo. Somebody reads the memo and makes a decision.

Five hops. At every hop the material gets shorter, and at exactly one hop it also changes language. By the time it reaches the decision, there is no route back. Asking what the customer actually said is not a question anyone can answer without reopening the audio, and the audio is in a shared drive nobody outside the local team can use.

Fixing this does not require changing the research. It requires that one artefact in the chain be complete, and that it be reachable.

Making the transcript the artefact of record

The transcript is the natural candidate, because it is the last point at which nothing has been thrown away.

Transcribe in the language of the interview. This part matters. Transcribing into English directly merges recognition and translation into one operation and produces a document that cannot be checked by the person best placed to check it, namely the local researcher. Vomo publishes its Portuguese entry under its own name, transcrever áudio em texto, and the reason to use the Portuguese-language tooling rather than an English pipeline is that the local team can read, correct and sign off on the result.

Then translate the transcript as a second, separate document, and keep both. A machine translation of an accurate transcript is usually good enough for a head office reader to work with, and when a passage matters, the original is one click away and a native speaker can be asked.

Speaker labels are worth insisting on, since an interview transcript without them makes the researcher’s leading question indistinguishable from the customer’s answer. Timestamps earn their keep the first time someone disputes a quote.

What changes at head office

The visible change is that people stop asking whether the summary is reliable, because they can check.

The less obvious change is that more than one person can read the research. A synthesis is a bottleneck by design, and it caps the number of insights at whatever one researcher noticed in one pass. Forty searchable transcripts can be read by the product manager looking for pricing signals, by the designer looking for onboarding language, and by the person writing the localised copy who needs actual sentences rather than approved terminology.

They will find different things, because they are looking for different things. None of that is available from a deck.

There is also an archival argument that only becomes obvious later. Research done in year one is usually unreadable by year three, because the researcher has left and their notes were shorthand. Transcripts survive staff turnover in a way that synthesis documents do not.

Where this goes wrong in the other direction

Two failure modes come from overcorrecting, and both are common enough to be worth flagging before anyone builds a process around this.

The first is dumping raw transcripts on people and calling it access. Forty unedited transcripts is not research output, it is a filing cabinet, and a head office reader handed forty documents in a language they do not speak will read none of them. The synthesis still has to be written. What changes is that it stops being the only thing that exists.

The second is treating the transcript as authoritative when it has not been checked. Recognition output in Portuguese is generally good and it still gets names, regional terms and numbers wrong, and a machine translation of an uncorrected transcript compounds both errors while making them fluent. Somebody who speaks the language has to read the transcript before it is treated as a record. That is the ten minutes referred to below, and skipping it produces confident nonsense that is worse than a careful summary.

There is a middle position that works. The synthesis remains the document people read. The transcripts sit behind it, corrected and searchable, for the cases where somebody needs to verify a claim, look for something the synthesis was not looking for, or quote a customer accurately in a document that will be seen outside the company.

What it costs to keep both

Worth being concrete, because the objection is usually about effort rather than principle.

Transcription itself is now a background task rather than a line item. The real cost is the correction pass, which runs at roughly a quarter of the recording length for a careful reader working in their own language, less if the audio is clean and the interviewer avoided talking over people.

Translation of a corrected transcript costs close to nothing if machine translation is acceptable for the reading copy, which for internal research it usually is. Where a quote will be published or shown to a customer, that specific passage gets translated properly by a person, and the passages that need this treatment are typically a handful per study rather than the whole corpus.

Storage and access are the parts that get neglected. A transcript nobody can find is the same as no transcript, so the files need to sit somewhere the head office team can actually search, with the study name and date in the filename rather than a recording timestamp.

Consent, and what you promised

Recording an interview requires the participant’s agreement, given before the recording starts and in the language of the interview.

Two details are worth handling explicitly when material will cross borders. Tell participants who will have access, and be accurate about it, because a customer agreeing that a local researcher may record is not automatically agreeing that a transcript will circulate at a head office in another country. And say what happens to the recording afterwards, including when it is deleted.

Where a translated transcript will be shared more widely than the original, the safer default is to strip names and identifying details from the translated version and keep the identified original with the local team.

A version of this that is worth doing

For a small research programme, the whole change amounts to three files per interview instead of one line in a document: the audio, a corrected transcript in the original language, and a translation.

The added effort per interview is on the order of ten minutes, most of it correcting names. The thing it buys is that six months later, when somebody asks whether any customer actually said the payment flow was a problem, there is an answer, and it takes about a minute to find.