# Turning Vietnamese call-centre recordings into call summaries

> PhoWhisper transcribes the Vietnamese and WhisperX works out who said what. Summaries usually break where the two are joined.

Original: https://fdetimes.net/en/guides/vietnamese-call-center-speech-to-text-summaries/

Picture a call centre that records thousands of calls a day and almost nobody listens back. When the client asks an FDE to "do something with all these recordings", the model is not the hard part. The hard parts are low-quality 8kHz files, two people talking over each other, and regional accents the demo set never contained.

AWS describes speech-to-text as software that listens to audio and returns a word-for-word, editable transcript. It also lists call analytics, meaning insights drawn from customer conversations, among the technology's applications. This guide builds a small version of that problem on your laptop.

You will end up with a transcript, speaker labels and a summary, plus a quality score you calculated yourself.

## What will you build, and what do you need?

The pipeline has five steps: normalise the audio, transcribe it, add timestamps and separate the speakers, assemble a labelled conversation, then summarise. A final step measures where the whole chain goes wrong.

You need Python, the Hugging Face `transformers` library and a few call recordings. If you have no real data yet, role-play an agent and a customer with a friend and record it on a phone. A GPU helps but is not required for the smaller models.

Whisper comes in six model sizes that trade speed against accuracy, so the client's GPU budget will decide which one you use.

## Step 1: why do call-centre files need to go up to 16kHz?

Call-centre recordings are usually stored at 8kHz, while PhoWhisper needs input audio sampled at 16kHz. This is an input requirement of the model, so do not skip it just because the files seem to play normally. The first step is to resample every file to 16kHz and name the outputs clearly, for example `call_001_16k.wav` ("call 001, 16k").

The code below is a minimal version that assumes you have installed `librosa` and `soundfile`. You can use `torchaudio` or `ffmpeg` instead, as long as the output is 16kHz:

```python
# Simplified: resample a recording to 16kHz, mono
import librosa
import soundfile as sf

audio, sr = librosa.load("call_001.wav", sr=16000, mono=True)
sf.write("call_001_16k.wav", audio, 16000)
print(sr)  # should print 16000
```

**Check:** open a few converted files and listen to them, and print the sample rate to confirm it is 16000. A common mistake is resampling only part of a folder. Results then look good on some files and poor on others, and you end up blaming the model.

## Step 2: original Whisper or PhoWhisper?

OpenAI's repository presents Whisper as a general-purpose speech recognition model: it transcribes many languages, translates speech and identifies languages. The same repository states plainly that Whisper's quality varies widely by language. For Vietnamese, that is the reason to try a model fine-tuned for the language.

VinAI's PhoWhisper is multilingual Whisper fine-tuned on 844 hours of Vietnamese data covering many regional accents, and it comes in five versions. You call it through the `transformers` ASR pipeline:

```python
from transformers import pipeline

transcriber = pipeline("automatic-speech-recognition", model="vinai/PhoWhisper-small")
# Simplified: take the text part of the result
output = transcriber("call_001_16k.wav")["text"]
print(output)
```

If you want to compare against the original Whisper, there is one trap: the `turbo` model was not trained for translation. For Vietnamese call-centre audio, use it to transcribe and do not ask it to translate into English.

**Check:** read the transcript while listening to the audio. Watch proper names, order numbers and phone numbers. These matter most to the client and are also where errors are most likely.

## Step 3: who is speaking, and when?

A continuous block of text is not yet a call record. To write a summary such as "customer complained about late delivery, agent promised a refund", you need to know who said each line. The original Whisper only gives sentence-level timestamps, which are too coarse to cut accurately where two people talk over each other.

WhisperX fills this gap in three ways. It aligns with wav2vec2 to get accurate timestamps for every word, and it uses pyannote-audio to tell speakers apart and return speaker ID labels.

It also runs VAD (voice activity detection) as a preprocessing step to reduce hallucination, and it runs batched inference without increasing WER, reaching about 70 times real time with large-v2.

One point needs stating clearly: these are WhisperX features, and nothing guarantees that a PhoWhisper transcript will feed directly into WhisperX's wav2vec2 alignment step. Treat combining the two tools as a design you must test and verify yourself.

The idea is that one stream gives you each word with its timing, another gives you time spans with speaker labels, and your job is to join them by time.

Take the exact commands and parameters for alignment and diarization directly from the README of the `m-bain/whisperX` repository, matching the version you have installed.

The joining code you can write yourself, and it is the part most worth practising. The minimal version below uses a data structure you define yourself; it is not the WhisperX API. Each word is assigned to the speaker whose time span overlaps it most. (`overlap` means "overlap"; `assign_speakers` means "assign speaker".)

```python
# Simplified: words and turns are data you convert into this shape yourself
# words = [{"word": "hello", "start": 0.4, "end": 0.7}, ...]
# turns = [{"speaker": "SPEAKER_00", "start": 0.0, "end": 3.2}, ...]

def overlap(a_start, a_end, b_start, b_end):
    return max(0.0, min(a_end, b_end) - max(a_start, b_start))

def assign_speakers(words, turns):
    for w in words:
        best = max(turns, key=lambda t: overlap(w["start"], w["end"], t["start"], t["end"]))
        w["speaker"] = best["speaker"]
    return words
```

**Check:** diarization only gives you `SPEAKER_00` and `SPEAKER_01`. It does not know which one is the agent. You have to write your own relabelling rule, for instance that whoever says the scripted opening greeting is the agent. Test that rule on every call, because a single call in which the customer speaks first will swap the roles across the whole summary.

## Step 4: what should the summary contain to be useful?

Once you have a labelled conversation, pass it to a language model to summarise. Do not ask something generic like "summarise this call". Ask for exactly the fields a call-centre manager needs to read. The template below is only an illustration: reason for the call, whether the customer was satisfied at the end, what the agent committed to (with any deadline), and next steps. Replace them with the fields the client actually uses:

```text
# Example summary template
- Why the customer called:
- Was the customer satisfied at the end of the call:
- What the agent committed to (with deadline, if any):
- Next steps:
```

Splitting the summary into fields lets you compare thousands of calls instead of reading thousands of free-form paragraphs. It also exposes errors in earlier steps: if the "what the agent committed to" field contains the customer's words, step 3 almost certainly assigned the wrong speaker.

## Step 5: measure on the client's own data

Because Whisper's quality varies so much by language, nobody's published figures can replace your own measurement on the client's Vietnamese calls.

The usual metric is WER, the word error rate. Suppose a reference sentence has 20 words, and the model gets 3 words wrong and drops 1. The WER is 4 divided by 20, or 20%.

Calculate this for each call rather than only taking the average, because one call with a strong regional accent and a high WER can be hidden by ten calls in a standard accent.

A further caution about borrowing published numbers: an arXiv paper on a pipeline for assessing Vietnamese service quality from speech was withdrawn by its own authors because its results contained significant errors. Before putting anyone's figures on a slide for a client, check that the source still stands.

| Failure mode | Symptom | First fix |
|---|---|---|
| Part of the files not resampled | Unusually uneven quality between files | Confirm a sample rate of 16000 for every file |
| Hallucination in silent stretches | Sentences nobody said appear while the caller is on hold | Enable VAD before transcription |
| Agent and customer roles swapped | Summary says the customer "promised a refund" | Review the speaker relabelling rule |

## How does this skill show up with clients?

If you are the FDE on a project like this, spend the first days measuring before you choose a model. Ask for 20 to 30 calls that represent different regions, times of day and complaint types. Sit with a shift lead to type reference transcripts, then run the pipeline and produce a WER table broken down by group.

That table will settle which model size to use, whether a GPU is needed and which calls must go to a human reviewer.

When reading FDE or solutions engineer job descriptions, look for phrases such as "call analytics", "speech", "contact center" or "ASR". On your CV, do not just write "used Whisper". Say that you measured WER on Vietnamese data, found errors caused by 8kHz audio or swapped speaker roles, and explain how you fixed them.

Choosing a model takes one line of code. What the client pays for is knowing which calls the pipeline gets wrong, and why.

**Try this week:**

- Record 5 mock calls of 2-3 minutes with friends (each with a different regional accent if possible), resample them to 16kHz and run them through vinai/PhoWhisper-small.
- Type reference transcripts for those 5 calls by hand and calculate WER for each one; note which call had the most errors and why.
- Write a one-page README describing your pipeline, with the WER table and an example summary, and link it from your CV as a call analytics case study.

## Sources

- [openai/whisper (GitHub)](https://github.com/openai/whisper)

- [VinAIResearch/PhoWhisper (GitHub)](https://github.com/VinAIResearch/PhoWhisper)

- [m-bain/whisperX (GitHub)](https://github.com/m-bain/whisperX)

- [What is Speech to Text? (AWS)](https://aws.amazon.com/what-is/speech-to-text/)

- [Speech-based Multimodel Pipeline for Vietnamese Services Quality Assessment (arXiv)](https://arxiv.org/abs/2412.09829)
