NVIDIA is addressing the lack of AI language support for many languages with new tools for high-quality speech recognition and translation for 25 European languages, including Croatian, Estonian, and Maltese. The Granary dataset offers over a million hours of audio, while models like Canary-1b-v2 and Parakeet-tdt-0.6b-v3 provide accurate transcription and translation capabilities. The team behind Granary collaborated with researchers to develop the dataset using the NVIDIA NeMo Speech Data Processor toolkit, offering clean, ready-to-use data for developers to build inclusive speech technologies for European languages. The Canary and Parakeet models, optimized for accuracy and speed, enable developers to innovate in speech AI with permissive licensing and improved performance compared to larger models. NVIDIA NeMo software suite accelerates speech AI model development, with tools like NeMo Curator and the Speech Data Processor toolkit enhancing data quality and processing efficiency.

Read more at NVIDIA: NVIDIA Releases Open Dataset, Models for Multilingual Speech AI