📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal’s AMÁLIA, a €5.5 million European Portuguese LLM, is operational but faces unresolved questions about its openness, native data adequacy, and optimization goals. These issues highlight broader challenges for European sovereign-LLMs.
Portugal’s €5.5 million AMÁLIA large language model is now operational, marking a significant step in the country’s AI development. However, critical questions about its openness, native-language data sufficiency, and primary goals remain unresolved, raising concerns about the broader European sovereign-LLM movement.
The model, developed by a consortium of approximately 60 researchers across Portugal’s top institutions, was officially launched in October 2025 and is currently available to 450,000 academic users via the FCT’s IAedu platform. It is based on a continuation of the EuroLLM multilingual foundation, with the training process including 107 billion tokens, of which only about 5.8 billion are Portuguese-specific, sourced mainly from the national web archive Arquivo.pt.
While AMÁLIA outperforms previous open models on Portuguese benchmarks and beats Qwen 3-8B on most tests, it still trails Qwen on the ALBA benchmark, which is considered a primary measure of Portuguese language capabilities. The final version is expected in June 2026, and the team has not yet publicly addressed several structural questions that are central to evaluating the model’s true readiness and strategic positioning.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.

Portuguese Flash Cards – Learn Portuguese Language Vocabulary Words and Phrases – Basic Language for Beginners – Gift for Travelers, Kids, and Adults by Travelflips
- Language: Portuguese flash cards for beginners
- Learning Features: Phonetic pronunciation and English translation
- Quality: Durable cards with portable storage
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications for European Sovereign-Language AI Strategies
The questions surrounding AMÁLIA reflect broader issues facing European efforts to develop independent, native-language large language models. These concerns impact national AI policies, funding priorities, and the future competitiveness of European AI in a global context. How open these models truly are, whether native-language data is sufficient, and what they should optimize for will influence their adoption, trust, and strategic value.
European Sovereign-LLMs and Portugal’s Role in AI Development
Portugal’s investment in AMÁLIA is part of a wider European initiative to foster sovereign-language AI, with projects like Italy’s Minerva, Germany’s Aleph Alpha, and France’s Mistral. These efforts aim to reduce dependency on US and Chinese models, but often face similar structural questions about data openness, native-language training, and strategic objectives. The public discourse has largely focused on individual launches rather than the systemic challenges they reveal.
AMÁLIA’s development is notable for its use of a continuation approach based on a multilingual foundation, contrasting with models like Minerva, which are trained from scratch on native language data. The model’s current performance and the unresolved questions about its openness and training data highlight the ongoing debate over the best path for European AI sovereignty.
“While AMÁLIA is an impressive technical achievement, we must critically examine what it truly delivers in terms of openness, native-language data, and strategic focus.”
— Duarte O.Carmo, AI researcher
Unanswered Questions About AMÁLIA’s Openness and Goals
It remains unclear how open AMÁLIA will be in future versions, especially regarding access to underlying data and training processes. Additionally, the precise strategic objectives—whether focusing on academic, commercial, or governmental applications—are not yet publicly defined. The final performance and capabilities of the model by June 2026 could shift these considerations, but current gaps in transparency persist.
Next Milestones for AMÁLIA and European AI Sovereignty
The final version of AMÁLIA is scheduled for release in June 2026, which will provide a clearer picture of its capabilities and strategic positioning. Over the next 12-24 months, the project team is expected to address some of the structural questions, potentially releasing more transparency on data sources, openness policies, and optimization goals. Broader European initiatives will likely scrutinize these developments to assess their alignment with sovereignty objectives.
Key Questions
What are the main challenges facing AMÁLIA’s development?
The key challenges include questions about how open the model will be, whether the native Portuguese data used is sufficient for high-quality performance, and what strategic objectives it is designed to serve.
How does AMÁLIA compare to other European models?
AMÁLIA outperforms previous open models on Portuguese benchmarks and beats Qwen 3-8B on most tests, but it still lags on the primary ALBA benchmark. Its development approach, based on a continuation of a multilingual foundation, differs from models trained from scratch.
Why are these questions about openness and data important?
They determine the model’s transparency, trustworthiness, and strategic value, impacting national AI sovereignty and Europe’s competitive position in AI development.
What will happen after the final version is released?
The team is expected to clarify some structural questions, and European policymakers will evaluate whether AMÁLIA meets sovereignty and strategic goals, influencing future funding and development directions.
Source: ThorstenMeyerAI.com