AI Models

Cohere Transcribe Arabic open-source model targets dialect challenges

Cohere has launched Cohere Transcribe Arabic, an open-source 2-billion-parameter ASR model designed for Arabic speech recognition. The company claims it is the most accurate open-source Arabic speech-to-text system, outperforming Whisper Large V3 and the standard Cohere Transcribe model in benchmarks. The model handles dialect variety, code-switching, bilingual conversations, and specialized vocabulary.

Neura News

Neura News

Neura Market Editorial

July 7, 20262 min read
Cohere Transcribe Arabic open-source model targets dialect challenges

Cohere releases open-source Arabic speech recognition model

Cohere has released a new open-source model called Cohere Transcribe Arabic, built specifically for Arabic speech recognition. The 2-billion-parameter ASR model is, according to Cohere, the most accurate open-source Arabic speech-to-text system currently available.

The model targets the unique challenges of transcribing Arabic speech. These include the wide variety of dialects across different regions, bilingual conversations where speakers switch between Arabic and English, code-switching within a single sentence, and specialized vocabulary used in fields like medicine, law, or technology.

Performance compared to existing models

Cohere says that Cohere Transcribe Arabic outscores Whisper Large V3, the standard Cohere Transcribe model, and other systems in benchmarks. Human ratings of Arabic transcripts on a scale of 1 to 5 showed that the new model scored higher in overall quality, dialect faithfulness, and code-switching accuracy compared to those alternatives.

Cohere provided an image illustrating these human ratings, showing Cohere Transcribe Arabic outperforming both Whisper Large V3 and the standard Cohere Transcribe model in all three categories.

Availability and licensing

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The model is released under the Apache 2.0 license, making it free to use, modify, and distribute. It is available for download on Hugging Face and can also be accessed through the Cohere API. More benchmarks, example transcriptions, and technical details are available on the Cohere blog.

Cohere is a Canadian AI company founded in 2019, known for its large language models and enterprise AI solutions. The company has focused on making its models available both as open-source releases and through cloud APIs, offering flexibility for researchers and businesses.

Significance for Arabic language technology

Arabic speech recognition has historically lagged behind English due to the language's complex morphology, numerous dialects, and lack of labeled training data. Open-source models like Cohere Transcribe Arabic could help accelerate development of Arabic-language applications in customer service, healthcare, media transcription, and education.

The release underscores a growing trend among AI companies to release specialized models for underrepresented languages, rather than relying on general-purpose models that may not capture linguistic nuances accurately.

Cohere has not disclosed the specific training data used for this model, but the company says it focused on representative samples of various Arabic dialects and code-switched speech.

Related on Neura Market:

More from Neura News

Industry

Monday.com Joins Tech Layoff Trend Citing AI as Factor

Monday.com announced it will lay off about 20% of its workforce, or over 600 employees, citing a restructuring tied to its AI-driven growth strategy. The Tel Aviv-based work management software company joins a growing list of major tech firms, including Amazon, Meta, and Microsoft, that have cited artificial intelligence as a factor in job cuts this year. A new Financial Times analysis shows U.S. tech companies have slashed nearly 140,000 jobs since January, with AI often cited as a reason.

Jul 26·12 min read
General

Open-weight AI mirrors Kubernetes ecosystem shift

Tobi Knaup, co-founder of Mesosphere, draws parallels between the rise of Kubernetes and the current trajectory of open-weight AI models. He argues that open-weight models are becoming a neutral substrate for innovation, attracting a global ecosystem of developers, startups, and enterprises. The piece warns against US restrictions on Chinese open-weight models, advocating instead for American leadership through open releases, procurement strategies, and standards.

Jul 25·7 min read
General

Open-weight AI mirrors Kubernetes rise, US warned on bans

The author, a Mesosphere co-founder, draws parallels between the rise of Kubernetes and the current open-weight AI ecosystem. He argues that open-weight models are becoming a neutral platform for innovation, and warns that US restrictions on Chinese open-weight models could isolate American developers from a global ecosystem. The piece urges the US to compete by releasing frontier models, using procurement to create demand, building the stack, and setting standards rather than imposing bans.

Jul 25·7 min read
Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read