Software
ESPnet
End-to-End Speech Processing Toolkit
An open-source toolkit for speech recognition, text-to-speech, speech enhancement, speech translation, and spoken language understanding. It provides reproducible recipes and a complete setup for speech foundation model research.
VERSA
Versatile Evaluation of Speech and Audio
A toolkit for evaluating speech and audio quality. It provides seamless access to over 90 evaluation and profiling metrics with 10x variants, assessing audio through multiple dimensions.
OWSM
Open Whisper-style Speech Models
Reproduces Whisper-style training using publicly available data and ESPnet. Data preparation scripts, training and inference code, pre-trained model weights, and training logs are all publicly released.