TL;DR
A published headline claims Huawei Pangu Pro trained 505 billion parameters without Nvidia hardware while suggesting that supply-chain evidence complicates that account. The available material provides no technical report, chip inventory, supplier records or independent verification, leaving both assertions unsubstantiated.
A published report has linked Huawei Pangu Pro to a 505-billion-parameter training run completed without Nvidia accelerators, while its headline also suggests that unspecified supply-chain evidence complicates that account. No technical report, hardware inventory, supplier records or independent audit accompanied the available material, so the central claims remain unverified.
The report advances two related but separate propositions. It says Pangu Pro reached a scale of 505 billion parameters and that its training did not use Nvidia hardware. It then indicates that the supply chain tells a different story, without identifying the component, supplier or record behind that qualification.
The available material does not specify the model’s architecture or explain what the parameter figure represents. In a dense model, all parameters may participate in each operation; in a mixture-of-experts system, the reported total can be far larger than the number of active parameters used for each token. Without that distinction, training data volume, computing budget, evaluation results and cluster configuration, the scale and cost of the reported run cannot be established.
The phrase without Nvidia is also undefined. It could refer only to accelerators used in the main training run, or it could claim that Nvidia products were absent from experiments, evaluation, deployment and supporting infrastructure. The material provides no accelerator model, cluster inventory or methodology, and it does not show whether Huawei itself made the claim or whether the wording originated with the publication.
505 Billion Parameters. No Nvidia. No Proof—Yet.
A published headline attributes a massive, Nvidia-free training run to Huawei Pangu Pro. But the supplied material includes no technical report, hardware inventory, supplier record or independent audit—leaving both the model-scale claim and the supply-chain qualification unresolved.
Architecture and active-parameter count are not disclosed.
The scope of “without” and the accelerator inventory are undefined.
Consequential if documented; unsupported by the available evidence.
Two claims, several missing links
The report combines a model-scale assertion with a hardware-provenance assertion, then hints that supply-chain evidence complicates the story. Each proposition requires different documentation.
What does 505B measure?
No architecture is identified, so the figure could describe all parameters or only a differently defined model total.
Definition missingWhat ran the cluster?
No accelerator model, device count, topology, runtime, power demand or training methodology is provided.
Inventory missingWho made the assertion?
The material does not establish whether Huawei confirmed the wording or whether it originated with the publication.
Source unclearWhich dependency conflicts?
No component, supplier, procurement record or manufacturing dependency is named.
Evidence unspecifiedDid the model deliver?
Training data, benchmark results, evaluation protocol and measured capabilities are absent.
Results missingHow broad is “without”?
The phrase may cover only the final run—or experiments, evaluation, deployment and supporting systems too.
Scope undefinedParameter count is not compute cost
A model can advertise a very large total while activating only a subset for each token. Without knowing whether Pangu Pro is dense or mixture-of-experts, the headline number cannot establish training requirements.
Most parameters participate
In a dense architecture, broadly all model parameters may be involved in each forward operation.
- Total and active parameter counts are comparatively close.
- The reported scale more directly signals per-token compute demand.
- Training cost still depends on tokens, hardware and efficiency.
Total can greatly exceed active
A routed architecture may hold many experts while selecting only a subset for each token. Both figures are needed for a meaningful comparison.
| Published proposition | What would support it | Available material | Assessment |
|---|---|---|---|
| Pangu Pro has 505 billion parameters | Model card, architecture diagram and parameter definition | Headline-level figure only | Unverified |
| The model completed training at that scale | Run logs, token count, duration, checkpoints and compute budget | No training methodology disclosed | Unverified |
| The training run used no Nvidia accelerators | Cluster inventory, device identifiers and independent inspection | No accelerator inventory disclosed | Unverified |
| The supply chain tells a different story | Named components, suppliers and traceable procurement records | No conflicting dependency identified | Unresolved |
A cluster is more than its accelerator
Even if Nvidia accelerators were absent from the main run, technological independence would remain a stack-wide question involving manufacturing, memory, packaging, networking, software and infrastructure.
Chip design
Architecture, tooling and intellectual property
Fabrication
Process technology and production equipment
Memory
High-bandwidth capacity and suppliers
Packaging
Advanced integration and chip interconnects
Network
Switches, optical links and cluster fabric
Software
Compilers, libraries and orchestration
The key distinction: “No Nvidia accelerators in the final training run” is much narrower than “no Nvidia hardware anywhere in development, evaluation, deployment or supporting infrastructure.” The supplied material proves neither interpretation.
How the report could become testable
A credible verification chain must connect the claimed model to its training run, physical cluster, component provenance and independent review.
Model definition
Architecture, total parameters and active parameters per token
Training record
Data volume, checkpoints, runtime and compute expenditure
Cluster inventory
Accelerator types, quantities, topology and supporting systems
Supply records
Named components, suppliers, fabrication and procurement links
Independent audit
Technical review, reproducible evaluation and provenance checks
Five disclosures needed next
Bottom line: Until these records appear, the story remains a consequential report—not a verified account of a 505-billion-parameter, Nvidia-free training program.
The Stakes of Nvidia-Free Training
If documented, the reported run would indicate that Huawei assembled enough non-Nvidia computing capacity to train an extremely large AI system. That would matter to developers, hardware suppliers and policymakers tracking whether alternative accelerator ecosystems can support large-scale model development.
The supply-chain qualification could change how that achievement is interpreted. A training cluster depends on more than accelerator branding: fabrication, high-bandwidth memory, advanced packaging, networking, software, power delivery and cooling all shape its capabilities. A domestically branded processor may still rely on foreign-linked equipment, components or intellectual property elsewhere in the stack, but the available material identifies no such dependency in this case.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
- Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
- Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
- Interface: PCIe 3.0 x16 with 250W TDP
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Claim Spanning the Computing Stack
Parameter count alone is not a reliable measure of model capability or training independence. A credible account would connect the model architecture with the number and type of accelerators, training duration, energy use, interconnect design and measured performance. None of those details appears in the supplied report summary.
The same problem applies to supply-chain provenance. A conflict could involve processor design or fabrication, memory, packaging, optical links, switches, compilers or earlier development equipment. The headline does not say which layer is at issue, whether Nvidia products were allegedly present, or whether the concern instead involves another foreign supplier.
This distinction matters because a statement about the main training cluster is narrower than a statement about the entire development process. Nvidia hardware could be absent from the final run yet have been used during earlier experiments or evaluation. No evidence in the available material establishes either scenario.

SXM2 Single Passthrough X16 Directly Connecting Backplanes for Highly Bandwidth GPU CPU Data Transfer in Training an
- High Bandwidth PCIe x16 Connection: Direct PCIe x16 channel for maximum data transfer
- Robust Heat Dissipation Design: Metal-based architecture with efficient cooling paths
- Stable High-Power Operation: Supports elevated power levels for demanding workloads
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Missing Records Leave Claims Unresolved
It is not yet clear whether 505 billion refers to total or active parameters, whether training was completed as described, or how the model performed. The training corpus, compute budget, run duration, architecture and benchmark results are also not disclosed.
The hardware claim remains equally open. There is no disclosed cluster inventory, accelerator quantity, procurement record or independent inspection. The supposed supply-chain discrepancy is also unexplained, leaving readers unable to determine whether it concerns Nvidia equipment, another foreign dependency or a broader interpretation of technological independence.

Acer Veriton AI Mini Workstation Personal Computer GN100-UD11 Series
- Powerful AI Performance: 1 PFLOPS FP4 AI with NVIDIA GB10 Superchip
- Pre-installed NVIDIA DGX OS: Optimized for full NVIDIA AI stack
- High-Performance GPU & CPU: Blackwell GPU with 5th-gen Tensor Cores and 20-core Arm CPU
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Documentation Needed to Test the Claim
The report can be tested only if Huawei, the publication or another source releases a model card or technical paper, identifies the accelerators and cluster topology, and explains the scope of the phrase without Nvidia. Parameter definitions and reproducible evaluation results would clarify the model-scale claim.
Supply-chain evidence would also need to name the relevant components and suppliers, describe how they were connected to the training run and provide records that independent specialists can examine. Until that happens, the story remains a potentially consequential report, not a verified account of a 505-billion-parameter Nvidia-free training program.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
- Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
- Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
- Interface: PCIe 3.0 x16 with 250W TDP
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did Huawei confirm that Pangu Pro has 505 billion parameters?
The supplied material does not establish that Huawei issued or verified the figure. It appears in a published headline, but no Huawei technical paper or model card was included.
Was Pangu Pro trained entirely without Nvidia hardware?
That has not been independently established. The report does not identify the accelerators used or define whether the claim covers only the main run or all experiments, evaluation and deployment.
What does 505 billion parameters mean?
The meaning is unclear because the material does not say whether it describes total parameters or parameters active for each token. That difference can materially affect the computing requirements of mixture-of-experts models.
What supply-chain evidence contradicts the Nvidia-free claim?
No specific evidence is identified. The potential discrepancy could involve chips, fabrication, memory, packaging, networking or software, but the available material names no records or suppliers.
What evidence would verify the report?
Verification would require a hardware inventory and training methodology, architecture details, parameter definitions, performance results and traceable supply records. An independent technical audit would provide stronger support than headline wording alone.
Source: Thorsten Meyer AI