Shinji Watanabe

Software

ESPnet

End-to-End Speech Processing Toolkit

An open-source toolkit for speech recognition, text-to-speech, speech enhancement, speech translation, and spoken language understanding. It provides reproducible recipes and a complete setup for speech foundation model research.

VERSA

Versatile Evaluation of Speech and Audio

A toolkit for evaluating speech and audio quality. It provides seamless access to over 90 evaluation and profiling metrics with 10x variants, assessing audio through multiple dimensions.

OWSM

Open Whisper-style Speech Models

Reproduces Whisper-style training using publicly available data and ESPnet. Data preparation scripts, training and inference code, pre-trained model weights, and training logs are all publicly released.