South African startup Vambo AI has released MORENA, an open-weight 1.5-billion-parameter language model trained from scratch for 12 African languages plus English and French.
Languages covered: ChiShona, Kiswahili, Hausa, Yorùbá, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele and Nigerian Pidgin.
Co-founder Isheanesu Misi describes it as the largest open African AI model trained from scratch — a qualifier that distinguishes it from models fine-tuned on an existing base.
The claim, and what it rests on
Misi says MORENA outperforms a Google model eight times its size on African text modelling.
No benchmark is named, the Google model is not identified, and no results have been published. That is Misi’s claim rather than a demonstrated finding.
Because the weights are open, it is checkable — which is more than most performance claims in this field allow for, and worth someone doing.
Why tokenisation matters more than size
The more consequential argument concerns cost.
“It tokenises African languages more efficiently so it’s both cheaper to run in African contexts and cheaper to adapt,” Misi said, “which means startups in Nigeria, or researchers in Kenya, will spend less to make their own models and use them in real world settings.”
Tokenisation is how a model breaks text into units it can process. Models built primarily on English often fragment African languages into far more tokens than necessary, which raises the cost of every query and every training run. Fixing that at the tokeniser level compounds across everything built on top.
Vambo’s recommended use is fine-tuning rather than direct deployment — the model is positioned as a base for others to build on.
“It’s over three times the size of the second largest African language model, which gives other researchers a new point of reference,” Misi said. “It also proved that other African initiatives from scratch do not necessarily need billion dollar budgets.”
The company
Vambo AI was founded in April 2023 by Chido Dzinotyiwei and Isheanesu Misi, building AI that “understands the languages of people in the rest of the world.”
Individual users can write, search, translate and transcribe. Businesses and developers access an API and tooling to build across multiple geographies and localise solutions that have worked elsewhere.
The platform previously covered 11 languages including Arabic, Kiswahili, isiZulu and French.
What hasn’t been published
No licence terms, download location, training data provenance, compute budget or funding detail.
The licence matters most. “Open weight” covers a range of terms with materially different permissions, and developers deciding whether to build on MORENA need to know which applies.





