July 6, 2026

UvA and SURF AI model works in six languages

PhD candidates Dheeraj Varghese and Mohammad Derakhshani, at the VIS Lab of the University of Amsterdam’s Informatics Institute, have developed NeoBabel: a text-to-image AI model that generates images directly from prompts in six languages, including Dutch, Chinese, Hindi, French and Persian. Both the VIS Lab and SURF, the Dutch collaborative organisation for IT in education and research that supplied the computing power behind the model, are based at Amsterdam Science Park.

Words lost in translation

Most AI image generators are built on English-language datasets. A prompt in another language is typically translated into English in the background first, a step that can distort meaning. Varghese points to a Dutch prompt describing a “beer” (bear) at a table: two leading Big Tech models, Salesforce’s BLIP3o and DeepSeek’s Janus Pro 7B, both produced an image of a drink instead of an animal, mistaking the Dutch word for its English homophone. NeoBabel, working directly in Dutch, generated the correct scene.

An Amsterdam Science Park collaboration, scaled up in Europe

Training a model across six languages required far more image-label pairs than were publicly available. The VIS Lab’s own algorithms expanded the dataset from 40 million to 124 million pairs, a task that demanded more computing capacity than the Netherlands’ national supercomputer Snellius could offer. SURF, whose Amsterdam office sits just across the park from the Informatics Institute, arranged access to LUMI, one of Europe’s most powerful supercomputers, located in Finland and operated through the European EuroHPC initiative. Over several months, the team worked with up to 1,000 GPUs at a time, supported by LUMI’s user support team as they adapted to unfamiliar AMD-based hardware.

Smaller, and open to everyone

Despite being four times smaller than the Big Tech models it was tested against, NeoBabel performed on par in English and more accurately in the five other languages. All code and data behind the model have now been released as open source, allowing other researchers to build on the work. Varghese’s next step is to apply what the team learned to so-called world models, AI systems that simulate how environments evolve over time.

Source: SURF

Related news

How can we help you?

If you are interested in Amsterdam Science Park, exploring opportunities, or simply have a question, feel free to get in touch. We’ll be happy to help.

For business inquiries contact

Petra Baarendse

Let's connect