AI · Analysis
The small language problem is a commercial problem, not a technical one.
Finnish, Norwegian, Danish and Icelandic model quality has improved sharply. The economics of maintaining those models have not improved at all.

Independent coverage
Published 9 September 2026
7 min read
Evidence: Analysis
Ask a researcher in Reykjavik or Oulu why frontier models handle their language poorly and you will not get a story about tokenisation, at least not first. You will get a story about maintenance budgets.
The technical gap has narrowed. Open weight base models fine-tuned on well-curated national corpora now perform acceptably on summarisation, classification and question answering in every Nordic language, Icelandic included. The remaining weakness sits in long-form generation and in domain-specific vocabulary, particularly clinical and legal.
Why the gap persists anyway
A model is not a artefact you finish. It is a commitment. Every base model generation resets the work, and a national language model with three full-time maintainers competes against a global model with a team of hundreds and a commercial reason to keep going.
Iceland is the clearest case. The national language technology programme produced high quality resources, and those resources now sit inside several commercial models. The public programme gets neither the revenue nor the compute that the improvement generates.
The pattern that is working
The approach that has held up best is unglamorous. Rather than training a national model end to end, teams in Norway and Finland have concentrated on evaluation sets, clean domain corpora and adapter layers that can be re-applied to whatever open weight base is current.
This treats the base model as a commodity that will be replaced and the language data as the durable asset. It is the correct reading of where value actually sits.
What it means for buyers
For a Nordic hospital, bank or municipality evaluating an AI vendor, the useful question is not which model the vendor uses. It is whether the vendor can demonstrate performance on a locally built evaluation set in the working language, and whether the vendor commits to re-running that evaluation when the underlying model changes.
Almost no procurement document currently asks for the second thing. The ones that do get materially better answers.
Sources
Related reading