Jessie A Ellis
Aug 15, 2025 09:01
NVIDIA introduces the Granary dataset and fashions designed to enhance speech recognition and translation throughout 25 European languages, addressing information shortage in AI language fashions.
NVIDIA has unveiled a brand new open dataset and fashions aimed toward advancing multilingual speech AI, addressing the restricted language help in present AI language fashions. The Granary dataset, alongside the NVIDIA Canary and Parakeet fashions, seeks to boost speech recognition and translation capabilities for 25 European languages, together with underrepresented ones akin to Croatian, Estonian, and Maltese, in line with NVIDIA’s weblog.
Granary Dataset: A New Useful resource for AI Builders
The Granary dataset is a complete assortment of multilingual speech datasets, encompassing roughly 1,000,000 hours of audio. This contains practically 650,000 hours devoted to speech recognition and over 350,000 hours for speech translation. The dataset is accessible on Hugging Face, offering a beneficial useful resource for builders to scale AI purposes globally, facilitating the creation of multilingual chatbots, customer support voice brokers, and real-time translation providers.
Developed in collaboration with Carnegie Mellon College and Fondazione Bruno Kessler, the Granary dataset makes use of NVIDIA’s NeMo Speech Knowledge Processor toolkit to rework unlabeled audio into structured, high-quality information. This modern processing pipeline permits for enhanced public speech information with out the necessity for intensive human annotation, making it a crucial useful resource for AI coaching within the European Union’s official languages, plus Russian and Ukrainian.
Introducing NVIDIA Canary and Parakeet Fashions
The NVIDIA Canary-1b-v2 and Parakeet-tdt-0.6b-v3 fashions, educated on the Granary dataset, provide highly effective instruments for transcription and translation. Canary-1b-v2, a billion-parameter mannequin, helps high-quality transcription of European languages and translation between English and 24 different languages. In the meantime, Parakeet-tdt-0.6b-v3, with 600 million parameters, is optimized for real-time or large-volume transcription duties.
Each fashions are designed to offer correct punctuation, capitalization, and word-level timestamps of their outputs. Canary-1b-v2 is especially notable for its effectivity, providing transcription and translation high quality corresponding to fashions thrice its measurement, whereas operating inference as much as ten instances sooner.
Advancing Speech AI Innovation
By sharing the methodology behind Granary and its related fashions, NVIDIA is empowering the worldwide speech AI developer group to adapt related information processing workflows to different automated speech recognition (ASR) or automated speech translation (AST) fashions, thereby accelerating innovation within the discipline. The fashions and dataset are publicly accessible beneath a permissive license, encouraging widespread use and adaptation.
The Granary dataset and NVIDIA’s new fashions characterize a big step ahead in addressing the challenges of knowledge shortage in speech AI, notably for languages which were traditionally underrepresented in AI language fashions. This initiative not solely broadens the scope of multilingual speech recognition and translation but additionally enhances the inclusivity and effectiveness of AI applied sciences globally.
The Granary dataset and fashions can be found for exploration on Hugging Face, and additional particulars could be accessed on NVIDIA’s weblog.
Picture supply: Shutterstock


