Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

TL;DR

Rio de Janeiro’s local LLM, marketed as an original model, is confirmed to be a direct merge of existing models Nex and Qwen. This raises questions about its claimed originality and development process.

Rio de Janeiro’s so-called ‘homegrown’ large language model (LLM), Rio-3.5-Open-397B, is confirmed to be a direct blend of existing models Nex and Qwen, not an independently trained model as its developers claimed. This revelation questions the authenticity of Rio’s local AI initiative and raises concerns about transparency in AI development claims.

Investigations by researchers on Hacker News have uncovered that Rio’s LLM, marketed as an original 397-billion-parameter model trained by IplanRIO, is actually a weighted combination of two existing models: Nex and Qwen. The analysis shows that every weight tensor in Rio’s model is approximately a 0.6/0.4 blend of Nex and Qwen, across all layers and components, with no evidence of any unique training or fine-tuning conducted by Rio.

Further testing revealed that when the model’s system prompt is removed, it identifies itself as ‘Nex, from Nex-AGI’ 79% of the time and as ‘Rio’ only 0%, reciting the organization’s backstory verbatim. These findings strongly suggest the model is not original but a simple interpolation of existing models, contradicting claims of local training and development.

Implications for Local AI Development Claims

This discovery impacts Rio de Janeiro’s AI reputation, casting doubt on claims of developing a unique, homegrown LLM. It highlights potential issues of transparency and raises questions about the integrity of local AI initiatives that market themselves as independent innovations.

For the broader AI community, it underscores the importance of verifying model origins and training processes, especially when public funding or regional pride is involved. The revelation may influence future oversight and scrutiny of AI projects claiming local development.

Yahboom Jetson AGX Thor Developer Board 128GB 2070 TFLOPS AI Large Model Voice Module, USB 3.0 HUB, 15.6in Display, USB Camera

Yahboom Jetson AGX Thor Developer Board 128GB 2070 TFLOPS AI Large Model Voice Module, USB 3.0 HUB, 15.6in Display, USB Camera

【AI Performance for Edge Computing】 Powered by N-VIDI-A Jetson AGX Thor module with 128GB memory and 2070 TFLOPS…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Rio’s AI Initiatives and Model Claims

Rio de Janeiro has promoted its AI efforts as a regional innovation, with IplanRIO announcing the release of an ‘original’ 397-billion-parameter LLM, Rio-3.5-Open-397B. The model was presented as a locally trained, independent AI solution aimed at regional and public sector applications.

However, prior to this investigation, there was limited transparency about the training data, sources, or whether the model was built from scratch or adapted from existing models. The recent findings suggest that the model is a simple weighted merge of Nex and Qwen, with no evidence of separate training by Rio or IplanRIO.

“Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen across all layers.”

— an anonymous researcher

Fine-Tuning Large Language Models: From Custom Datasets to High-Performance AI Models Using Modern Toolchains

Fine-Tuning Large Language Models: From Custom Datasets to High-Performance AI Models Using Modern Toolchains

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Model Development and Claims

It is still unclear whether Rio or IplanRIO attempted any form of fine-tuning or additional training beyond the merging process. The extent of transparency and whether the model’s creators intended to misrepresent its origins remain unverified. Additionally, the implications for regional AI initiatives that rely on such claims are still unfolding as more details emerge.

AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Verification and Transparency

Further analysis and independent audits are likely to follow, aiming to clarify whether Rio’s model involved any original training or fine-tuning. Authorities or oversight bodies may scrutinize the project’s claims more closely, and Rio’s officials might respond publicly to these findings. The broader AI community will also monitor how regional projects handle transparency and representation of their models.

EXPLAINABLE AI : Techniques that Meet Auditors’ Needs : Building Transparent, Defensible, and Audit-Ready Artificial Intelligence for Modern Enterprises

EXPLAINABLE AI : Techniques that Meet Auditors’ Needs : Building Transparent, Defensible, and Audit-Ready Artificial Intelligence for Modern Enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Is Rio’s LLM truly an original model?

No, investigations indicate it is a weighted merge of Nex and Qwen models, not an independently trained or fine-tuned model.

Why does this matter for Rio’s AI reputation?

It raises concerns about transparency and the integrity of local AI development claims, which could impact regional trust and credibility.

Could Rio have trained the model independently?

Based on current evidence, there is no indication of independent training; the model appears to be a simple blend of existing models.

What are the broader implications for AI projects claiming local development?

This case highlights the need for transparency and verification in AI development claims, especially when public or regional funds are involved.

Will there be further investigations?

Likely, as independent researchers and authorities seek to confirm the model’s origins and ensure transparency in regional AI initiatives.

Source: Hacker News


You May Also Like

Police boast of hacking VPN where criminals “believed themselves to be safe”

Authorities dismantled the First VPN infrastructure, arresting the administrator and sharing intelligence that aids ongoing cybercrime investigations worldwide.

The app you need to clean up your computer

A new utility app has been launched claiming to help users identify and remove unnecessary files and processes to improve computer performance.

There Are No Instances in ATProto

ATProto does not have instances like Mastodon; it separates hosting from app aggregation, offering a different approach to decentralization.

Better Models: Worse Tools

Recent tests reveal that the latest Anthropic models, including Opus 4.8 and Sonnet 5, increasingly produce malformed tool calls, despite being more advanced.