The Luxembourgish language will not be left behind in the AI revolution, thanks to a University of Luxembourg project that has just received Microsoft backing.
LuxVLD (Luxembourgish Vision-Language Dataset for Education and Digital Inclusion) lays the foundations for a Luxembourgish vision-language model capable of understanding, reasoning, and interacting in Luxembourgish across text and images.
It does this by creating high-quality, AI training data in Luxembourgish, such as text, audio and images paired with texts, ensuring that AI systems have structured, reliable material to learn Luxembourgish properly. When done correctly, it means the AI will be able to reason directly in Luxembourgish, reducing loss of nuance, awkward phrasing and cultural inaccuracies.
And it paves the way for the creation of Luxembourgish language chatbots, classroom and public sector communication tools, among other things.
“This ambition opens the door to a new generation of human-centered AI applications, developed for Luxembourg and in Luxembourgish,” said Professor Djamila Aouada, Deputy Director of SnT and head of the CVI2 research group.
Boosting AI adoption
Today, nearly half of all web content is in English and it is estimated that 90% of current training data for generative AI is English.
LuxVLD is led by the SnT (Interdisciplinary Centre for Security, Reliability and Trust) at the University of Luxembourg. Its solution for Luxembourgish is among 10 different projects being supported by Microsoft’s LINGUA Open Call, an initiative aimed at boosting underrepresented languages in artificial intelligence. Luxembourgish is spoken by an estimated 400,000 people and as such faces challenges of digital underrepresentation. Recent research from the Microsoft AI for Good Lab shows that countries with mainly underrepresented languages adopt AI 20% more slowly, even when connectivity and digital skills are comparable.
“AI adoption happens much faster when people can use their own language,” explained Marijke Schroos, General Manager of Microsoft Belux. “By supporting projects such as LuxVLD, we are helping ensure that Luxembourgish speakers can fully participate in the AI revolution without having to abandon their language. This is the kind of technological progress we want to see: innovation that truly includes everyone.”
LINGUA was a European call for projects supporting initiatives that create open, high-quality datasets for underrepresented languages, making it possible to train multilingual AI models and help communities develop language technologies that reflect their cultural and linguistic realities. Part of EU Digital Unlock, it is developed in coordination with the APERTUS, a fully open large language model (LLM) designed by EPFL and ETH Zurich, with the Council of Europe. The selected projects cover 16 languages ranging from Icelandic and Basque to Maltese, and Romani, receive financial support for dataset creation, as well as technical guidance through Microsoft’s collaboration with the APERTUS research consortium.