AI Models

Norway Uses 2 PB of Huawei Flash Storage for National LLM

Norway's National Library is developing a sovereign large language model for the Norwegian language, using 2 petabytes of Huawei OceanStor Dorado all-flash storage in its AI training pipeline. The project addresses the lack of commercial Norwegian LLMs and highlights challenges in moving data from archival storage to high-performance AI systems.

Neura News

Neura News

Neura Market Editorial

May 25, 20264 min read
Norway Uses 2 PB of Huawei Flash Storage for National LLM

Norway's National Library is building a large language model that understands the Norwegian language and is using 2 PB of Huawei OceanStor Dorado flash storage to handle its AI training data pipeline.

Marius Husnes, Head of IT Platform at the library, discussed the project at Huawei's ID Forum 2026 in Paris. He explained that no commercial LLM provider was developing a local Norwegian language model. He argued that any country with its own language that lacks a sovereign LLM trained in that language would be at a disadvantage. A globally trained, English-speaking LLM would not know about that country's history, news, and culture described in the local language.

The Library's Role as AI Builder

Norway's Ministry of Culture tasked the National Library with building a sovereign AI because the library holds the largest digital collection of Norwegian books, newspapers, web pages, and other materials. Like many state libraries, it receives copies of every published book and broadcast content. Its legal deposit mandate extends beyond books, making it duty-bound to collect and preserve all of Norway's cultural heritage.

The library has an agreement with Norwegian newspapers that permits training on copyrighted content. Husnes stated, "No private company has this."

The library was well positioned for this work because it has been digitizing its collection since 2005. It has amassed 20 PB of unique data stored in a 3-2-1 format three copies, two media types, one off-site, totaling about 60 PB overall. The digitization process for raw text, sound, moving pictures, still images, and web content involved extensive OCR scanning and generated a large amount of metadata, along with APIs for online access.

The Storage and Pipeline Challenge

The bulk of the data resides in a digital disk plus tape archive optimized for preservation. Husnes's task was to move this data to the LLM training system. He said the bottleneck was not compute but data quality, cleaning, and pipeline throughput.

There are two main processing stages. The first stage is in-house computation, using an Nvidia DGX H200 system, a 384 core CPU cluster, and multiple Huawei OceanStor Dorado all-flash arrays, totaling 2 PB of flash capacity. This low-latency storage serves the data pipelines and training preparation.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The pipeline includes steps for data ingestion, cleaning, deduplication, format normalization, validation, and preparation. Once data passes through the pipeline, it is sent to Norway's national supercomputer, the Sigma2 Olivia system, for actual training runs. The Olivia system is an HPE Cray Supercomputing EX system with 448 GPUs and 64,512 CPU cores, using a 5.3 PB Cray ClusterStor E1000 storage system.

One major challenge was overcoming the different needs of two storage systems. The 60 PB preservation system is optimized for durability and cost, not fast IO, and has high read latency because it is designed for infrequent access. The AI pipeline storage is designed for high-throughput, low-latency, parallel data IO. Husnes said he learned that nobody was talking about the problems involved in moving PB-scale datasets from an archive to and through an AI data pipeline. His team had to figure out how to do it themselves.

Ongoing Learning and Future Issues

The LLM training is still ongoing. Husnes summarized what his team continues to learn about:

Evaluation: There are no standard evaluation tools to assess a sovereign Norwegian LLM. The language has two written forms, multiple dialects, and historical changes. The team is building their own evaluation tool on the fly.

Governance: Who controls access to a sovereign LLM? Who decides what it can be used for? These are institutional and political questions with no easy answers.

Orchestration: Making three systems the preservation archive, the on-prem AI environment, and the national Sigma2 supercomputer work smoothly together is an ongoing project.

The takeaways here are that Huawei storage is playing a serious and significant role in the European market, and that any country developing a sovereign, local language LLM would do well to consult with Husnes and learn what is involved.

As Husnes put it, Norway is a small country solving a problem every non-English-speaking nation will face: how do you build AI that reflects your language, your culture, and your history? AI needs custodians, not just builders.

Related on Neura Market

More from Neura News

Funding

Prentis AI Lab Co-Founded by Reid Hoffman, Marc Pincus Seeks $100M

Prentis, a new AI research lab co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is in talks to raise $100 million at a $1 billion valuation. The startup focuses on computer use models that automate office workflows. It has already signed contracts worth up to $50 million with several customers and claims its Hive-32B model outperforms rivals like OpenAI's GPT-5.4 and Anthropic's Claude Opus 4.6 on key benchmarks.

Jul 24·4 min read
Industry

Cognition Acquires Poke to Give Devin Coding Agent a Personality

Cognition, the startup behind AI coding assistant Devin, has acquired Poke, an AI assistant known for its friendly, conversational style. The deal, valued in the low nine figures, aims to bring Poke's personality-driven interaction model to Devin, making the coding agent feel more like a colleague than a tool. Poke will also benefit from Cognition's models and infrastructure to become faster and more reliable.

Jul 24·3 min read
AI Models

Anthropic expands Claude voice mode to Opus and Sonnet models

Anthropic has expanded Claude's voice mode to run on its most powerful models, Opus and Sonnet, across mobile, desktop, and web platforms. Users can now switch between models mid-conversation, use voice commands in eleven languages, and connect to external tools like Gmail, Google Calendar, or Slack to compose and send emails by voice. The update positions Claude as a unique option for tool integration in voice AI, though competitors like OpenAI and Google offer more natural speech processing.

Jul 24·2 min read