Shinji Watanabe

Notes ·

The story behind cpWER

Long-form, multi-speaker conversational ASR is everywhere right now: meeting transcription, podcasts, multi-talker speech foundation models, and now open releases from big companies such as Microsoft’s VibeVoice-ASR and NVIDIA’s multi-talker Parakeet with Sortformer diarization. Many of the papers in this area report a number called cpWER, and every time I see it, I am a little happy. It comes from a paper that a lot of people worked very hard on, and that is now remembered mostly for that one metric:

S. Watanabe et al., “CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings,” Proc. CHiME 2020. paper

Much of that momentum traces back to 2019. Long-form, multi-talker recognition has a long history, from ICSI and AMI meetings onward, and our work is only one piece of it. But 2019 was the year several of us decided, more or less at once, that the segmented setting was no longer the problem worth solving, and CHiME-6 was one of the places that happened.

It was also the first (and, so far, last) CHiME challenge I led as first author. Here is the story, and a small request at the end. Affiliations are as they were in 2019 and 2020: I was at JHU then, and many of the people named here have moved since.

What cpWER is

CHiME-6 Track 2 asked a question nobody had asked in a challenge before: take a two-hour recording of a real dinner party with four people talking over each other in multiple rooms, and produce a transcript with who said what, with no segmentation given. Diarization and recognition together, on unsegmented audio, with a fully open-source baseline. It is a hard task today; in 2019 it felt almost unreasonable.

The problem is how to score it. Word error rate needs a reference and a hypothesis for the same utterance, but on unsegmented audio the system does not know where the utterances are, or who is speaking. cpWER, the concatenated minimum-permutation word error rate, is the simplest thing that works:

  1. Concatenate all utterances of each speaker, for both the reference and the hypothesis.
  2. Compute WER between the reference and every speaker permutation of the hypothesis (24 permutations for four speakers).
  3. Take the lowest one.

That is the whole definition. No time alignment, no segment matching. The reasoning was: CHiME-5 already had WER, and diarization already had DER, which measures time. So the ASR metric did not need a time axis at all. It only needed words and speaker attribution. Everything else was left to DER and JER.

Where it came from

After CHiME-5, Marc Delcroix told me that evaluating on oracle segments wasn’t realistic, and that we should do the whole thing. That comment is where Track 2 came from. The idea was in the air: by the end of 2019, Marc and Zhuo Chen were proposing a JSALT workshop on exactly this problem, unsegmented multi- talker recordings.

The baseline started in June 2019, when I wrote to Christoph Boeddeker and Reinhold Haeb-Umbach in Paderborn to ask whether guided source separation (GSS) could be added to a new Kaldi baseline. Their CHiME-5 system was already public; Christoph brought it in, and Aswin Shanmugam Subramanian put it into the Kaldi recipe. I wanted the baseline to be state of the art, and it was; multi-array GSS took about a week to run. Shota Horiguchi, visiting JHU from Hitachi with Yusuke Fujita, sat in our regular Hitachi-JHU meetings and helped with some of the early prototypes. They were also our line to the end-to-end diarization work at Hitachi; we even talked about putting EEND into the baseline, but we simply ran out of time.

Sanjeev Khudanpur offered the Kaldi team at JHU. I do not think my student Xuankai Chang could have done it alone. Vimal Manohar built much of the first Track 2 pipeline, including the first multi-speaker scoring script, before he left JHU that fall. Desh Raj built the speech activity detection, David Snyder the x-vector diarization (the heart of the diarization baseline), Ashish Arora ran the recognition side, and Bar Ben-Yair worked on the dereverberation side with WPE. Yenda Trmal and Dan Povey, the main Kaldi developers, made it reproducible and merged it into the Kaldi main repository, so it was a real recipe and not a one-off script. We met every Thursday evening in my lab that October.

By early November 2019, we were stuck on scoring. The NIST tool for this, asclite, was too complicated for us, and it silently drops regions when too many people overlap. Our other option was to concatenate everything per speaker and run edit distance over all permutations. I wasn’t confident, so I asked Naoyuki Kanda and Takuya Yoshioka at Microsoft. They came back with a clean recommendation, sclite over all 4! permutations, and one warning: a word given to the wrong speaker is counted twice, once as an insertion and once as a deletion, which is not how NIST RT defined speaker-attributed WER. Yusuke Fujita argued that the double count was the right thing. A human would never move a single word to another speaker; the mistake is really a diarization mistake and should cost more. I agreed. Yusuke also helped with the implementation. And we did not settle it by argument alone. We computed several flavors of WER on the outputs we already had, with different diarization results plugged in, and checked that cpWER moved the way we wanted it to before committing to it.

On November 15, I wrote to Jon Barker, Emmanuel Vincent, and Mike Mandel asking them to name the metric, and proposed cpWER, concatenated minimum- permutation WER. Mike countered with CWERMP and CWMP. Emmanuel had one condition, that the acronym end in WER so it was clearly a form of WER. Jon said it did what it said on the tin. That was it. They are strict about metrics, and I expected a fight; it took three days, and there was none.

The question I still get: if it is concatenated minimum-permutation WER, why is it cpWER and not cmpWER? What I told everyone at the time was that the concatenation is the important part, that minimum permutation is common in multi-speaker ASR anyway, and that cmpWER just looked ugly. The real reason is that, as a physics student, I worked on CP violation, and I wanted a CP in there. I never told anyone. Now you know.

The reference was the last piece. In February 2020, in the middle of the challenge, Maxim Korenevsky at STC pointed out that the reference segments contained long pauses that inflated everyone’s DER, and Christoph had already found a bug in how the reference RTTM was handled. Neville Ryant suggested forced alignment per participant. Because the CHiME-6 arrays were now synchronized, we could align once on the binaural audio and be done. Mike and Zhaoheng Ni did it at CUNY. That is what made a proper DER and JER possible on this data in the first place.

What survived

Looking back, cpWER is the piece that traveled furthest. Many of the techniques in that baseline are still in use: the unsegmented task became the template for later CHiME rounds, GSS is still a standard front end, and the Kaldi recipe is still there. Speech LLM systems use none of them, though. The metric is the part that crossed over: it is in much of the long-form ASR literature now, speech LLMs included, and it has grown a time-constrained descendant, tcpWER, from Christoph’s group. Systems get replaced every few years. Once people agree on a metric, it stays.

So this is the request. cpWER is a three-step definition, but behind it are twenty-one authors, listed alphabetically within each institute, several teams, and the work around the challenge: the JHU system, the GSS front-end, the Hitachi diarization, the Microsoft multi-speaker work. If you use cpWER, please cite the CHiME-6 paper, and please also look at the papers around it. When I sent the co-author email in April 2020, Naoyuki replied by sharing his first Microsoft paper, which turned out to be serialized output training. Almost everyone on that author list is doing remarkable things now, and I am proud to have been on it with them.

One last thing. The CHiME 2020 workshop was supposed to be in Barcelona on May 4, 2020, where we would have presented and argued all of this in person. Then COVID arrived, and it was held online. We never got to celebrate it in the same room. Maybe that is part of why I wanted to write this down.