NewsTradingSentimentCalendarCommunityBriefing
Tech

UCLA Researcher Secures Funding to Improve AI for African Languages

By Tech Desk · 2026-09-11 · 2 min read
A stylized globe with abstract network nodes connecting different continents
Illustration: Tradingbird

A new grant aims to close the data gap that leaves many African languages invisible to modern artificial intelligence systems.

Saadia Gabriel, a computer scientist at UCLA, has received $250,000 in compute credits to improve how artificial intelligence handles African languages. The funding comes from Amazon’s Build on Trainium program, which supports research using high-performance chips. She is one of 34 researchers selected from 30 universities globally to receive this support.

The project, reported by GN technics/ai (en-US), focuses on a critical weakness in current AI: the lack of sufficient training data for many languages. Most AI models are built on English-centric data, leaving speakers of other languages with tools that often fail to understand them. Gabriel and her doctoral student Sheriff Issaka plan to use the credits to develop models that better connect spoken and written African languages.

Data Scarcity Limits AI Accuracy

Artificial intelligence systems rely heavily on large volumes of high-quality data to learn language patterns. African languages, which account for nearly one-third of the world’s languages and are spoken by about 1.5 billion people, remain underrepresented in these datasets. This imbalance creates a significant barrier to creating inclusive technology. The researchers intend to address this by evaluating multimodal understanding and translation across five major linguistic groups.

These groups include Niger-Congo, Afroasiatic, Nilo-Saharan, Khoisan, and Khoe-Kwadi. Selecting these specific families provides broad geographic and linguistic coverage. By focusing on these diverse groups, the team aims to build datasets and models that work effectively for a much broader range of speakers. This approach seeks to reduce the bias that currently favors English and other well-resourced languages.

Expanding Resources With Native Speakers

Beyond the immediate grant, Gabriel’s team has launched a five-year effort to expand existing datasets. They are working directly with native speakers to collect more than 15,000 hours of validated speech data and 50 billion tokens of text across 40 African languages. This massive collection will serve as the foundation for training their new models. The data will support both speech recognition and text processing applications.

The trade-off in this approach is the time and effort required to validate such a vast amount of data. Unlike web-scraped text, speech data requires careful curation to ensure accuracy and cultural relevance. However, the team believes this investment is necessary to create AI systems that are truly representative of global linguistic diversity. Their recent work on this project was recognized with a Senior Area Chair Highlights Award at the Association for Computational Linguistics 2026 conference.

Broader Institutional Support at UCLA

Gabriel is part of a larger ecosystem of research supported by Amazon at UCLA. The Build on Trainium program has funded 14 Amazon Trainium Fellows at the UCLA Samueli School of Engineering. Additionally, 28 engineering doctoral students received fellowships for the 2026 winter quarter. These efforts are coordinated through the Science Hub for Humanity and Artificial Intelligence, a collaboration established in 2021.

As the director of the Misinformation, AI and Responsible Society Lab, Gabriel is well-positioned to lead this work. Her role also includes co-directing the Natural Language Processing Group at UCLA. The combination of dedicated funding, institutional backing, and a clear mission to improve data representation marks a significant step toward making AI more accessible to non-English speakers.

Based on reporting by GN technics/ai (en-US), compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories