Huawei Pangu Pro Trains 505 Billion Parameters Without Nvidia: Supply Chain Tells Different Story – Tech Times

TL;DR

A published headline claims Huawei Pangu Pro trained 505 billion parameters without Nvidia hardware while suggesting that supply-chain evidence complicates that account. The available material provides no technical report, chip inventory, supplier records or independent verification, leaving both assertions unsubstantiated.

A published report has linked Huawei Pangu Pro to a 505-billion-parameter training run completed without Nvidia accelerators, while its headline also suggests that unspecified supply-chain evidence complicates that account. No technical report, hardware inventory, supplier records or independent audit accompanied the available material, so the central claims remain unverified.

The report advances two related but separate propositions. It says Pangu Pro reached a scale of 505 billion parameters and that its training did not use Nvidia hardware. It then indicates that the supply chain tells a different story, without identifying the component, supplier or record behind that qualification.

The available material does not specify the model’s architecture or explain what the parameter figure represents. In a dense model, all parameters may participate in each operation; in a mixture-of-experts system, the reported total can be far larger than the number of active parameters used for each token. Without that distinction, training data volume, computing budget, evaluation results and cluster configuration, the scale and cost of the reported run cannot be established.

The phrase without Nvidia is also undefined. It could refer only to accelerators used in the main training run, or it could claim that Nvidia products were absent from experiments, evaluation, deployment and supporting infrastructure. The material provides no accelerator model, cluster inventory or methodology, and it does not show whether Huawei itself made the claim or whether the wording originated with the publication.

At a glance
reportWhen: Reported; publication date and current…
The developmentA report has attributed a 505-billion-parameter, Nvidia-free training run to Huawei Pangu Pro while offering no documentation for that claim or its supply-chain qualification.
Huawei Pangu Pro: 505 Billion Parameters, Unverified
AI infrastructure / claim audit

505 Billion Parameters. No Nvidia. No Proof—Yet.

A published headline attributes a massive, Nvidia-free training run to Huawei Pangu Pro. But the supplied material includes no technical report, hardware inventory, supplier record or independent audit—leaving both the model-scale claim and the supply-chain qualification unresolved.

Claim 01 505B parameters

Architecture and active-parameter count are not disclosed.

Claim 02 Trained without Nvidia

The scope of “without” and the accelerator inventory are undefined.

Current assessment Unverified

Consequential if documented; unsupported by the available evidence.

Reported scale 505B Claimed parameter total
Core propositions 2 Model scale + hardware independence
Disclosed records 0 In the supplied material
Evidence status Open Independent verification required

The report combines a model-scale assertion with a hardware-provenance assertion, then hints that supply-chain evidence complicates the story. Each proposition requires different documentation.

Model scale

What does 505B measure?

No architecture is identified, so the figure could describe all parameters or only a differently defined model total.

Definition missing
Training hardware

What ran the cluster?

No accelerator model, device count, topology, runtime, power demand or training methodology is provided.

Inventory missing
Claim ownership

Who made the assertion?

The material does not establish whether Huawei confirmed the wording or whether it originated with the publication.

Source unclear
Supply chain

Which dependency conflicts?

No component, supplier, procurement record or manufacturing dependency is named.

Evidence unspecified
Performance

Did the model deliver?

Training data, benchmark results, evaluation protocol and measured capabilities are absent.

Results missing
Independence

How broad is “without”?

The phrase may cover only the final run—or experiments, evaluation, deployment and supporting systems too.

Scope undefined

Parameter count is not compute cost

A model can advertise a very large total while activating only a subset for each token. Without knowing whether Pangu Pro is dense or mixture-of-experts, the headline number cannot establish training requirements.

Dense model

Most parameters participate

In a dense architecture, broadly all model parameters may be involved in each forward operation.

  • Total and active parameter counts are comparatively close.
  • The reported scale more directly signals per-token compute demand.
  • Training cost still depends on tokens, hardware and efficiency.
Mixture of experts

Total can greatly exceed active

A routed architecture may hold many experts while selecting only a subset for each token. Both figures are needed for a meaningful comparison.

Architecture
Not shown
Active parameters
Not shown
Training tokens
Not shown
Compute budget
Not shown
Benchmarks
Not shown
Published proposition What would support it Available material Assessment
Pangu Pro has 505 billion parameters Model card, architecture diagram and parameter definition Headline-level figure only Unverified
The model completed training at that scale Run logs, token count, duration, checkpoints and compute budget No training methodology disclosed Unverified
The training run used no Nvidia accelerators Cluster inventory, device identifiers and independent inspection No accelerator inventory disclosed Unverified
The supply chain tells a different story Named components, suppliers and traceable procurement records No conflicting dependency identified Unresolved

A cluster is more than its accelerator

Even if Nvidia accelerators were absent from the main run, technological independence would remain a stack-wide question involving manufacturing, memory, packaging, networking, software and infrastructure.

01

Chip design

Architecture, tooling and intellectual property

02

Fabrication

Process technology and production equipment

03

Memory

High-bandwidth capacity and suppliers

04

Packaging

Advanced integration and chip interconnects

05

Network

Switches, optical links and cluster fabric

06

Software

Compilers, libraries and orchestration

The key distinction: “No Nvidia accelerators in the final training run” is much narrower than “no Nvidia hardware anywhere in development, evaluation, deployment or supporting infrastructure.” The supplied material proves neither interpretation.

How the report could become testable

A credible verification chain must connect the claimed model to its training run, physical cluster, component provenance and independent review.

A

Model definition

Architecture, total parameters and active parameters per token

B

Training record

Data volume, checkpoints, runtime and compute expenditure

C

Cluster inventory

Accelerator types, quantities, topology and supporting systems

D

Supply records

Named components, suppliers, fabrication and procurement links

E

Independent audit

Technical review, reproducible evaluation and provenance checks

Five disclosures needed next

Model card or technical paper Architecture, parameter definitions and training methodology
Complete hardware inventory Accelerators, quantities, networking and cluster topology
Defined Nvidia-free scope Main run only—or experiments, evaluation and deployment too
Reproducible performance results Benchmarks, evaluation conditions and measured outcomes
Traceable supply-chain evidence Components, suppliers, records and connection to the run
Independent technical review Inspection that goes beyond publication wording

Bottom line: Until these records appear, the story remains a consequential report—not a verified account of a 505-billion-parameter, Nvidia-free training program.

The Stakes of Nvidia-Free Training

If documented, the reported run would indicate that Huawei assembled enough non-Nvidia computing capacity to train an extremely large AI system. That would matter to developers, hardware suppliers and policymakers tracking whether alternative accelerator ecosystems can support large-scale model development.

The supply-chain qualification could change how that achievement is interpreted. A training cluster depends on more than accelerator branding: fabrication, high-bandwidth memory, advanced packaging, networking, software, power delivery and cooling all shape its capabilities. A domestically branded processor may still rely on foreign-linked equipment, components or intellectual property elsewhere in the stack, but the available material identifies no such dependency in this case.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)

  • Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
  • Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
  • Interface: PCIe 3.0 x16 with 250W TDP

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Claim Spanning the Computing Stack

Parameter count alone is not a reliable measure of model capability or training independence. A credible account would connect the model architecture with the number and type of accelerators, training duration, energy use, interconnect design and measured performance. None of those details appears in the supplied report summary.

The same problem applies to supply-chain provenance. A conflict could involve processor design or fabrication, memory, packaging, optical links, switches, compilers or earlier development equipment. The headline does not say which layer is at issue, whether Nvidia products were allegedly present, or whether the concern instead involves another foreign supplier.

This distinction matters because a statement about the main training cluster is narrower than a statement about the entire development process. Nvidia hardware could be absent from the final run yet have been used during earlier experiments or evaluation. No evidence in the available material establishes either scenario.

SXM2 Single Passthrough X16 Directly Connecting Backplanes for Highly Bandwidth GPU CPU Data Transfer in Training an

SXM2 Single Passthrough X16 Directly Connecting Backplanes for Highly Bandwidth GPU CPU Data Transfer in Training an

  • High Bandwidth PCIe x16 Connection: Direct PCIe x16 channel for maximum data transfer
  • Robust Heat Dissipation Design: Metal-based architecture with efficient cooling paths
  • Stable High-Power Operation: Supports elevated power levels for demanding workloads

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Missing Records Leave Claims Unresolved

It is not yet clear whether 505 billion refers to total or active parameters, whether training was completed as described, or how the model performed. The training corpus, compute budget, run duration, architecture and benchmark results are also not disclosed.

The hardware claim remains equally open. There is no disclosed cluster inventory, accelerator quantity, procurement record or independent inspection. The supposed supply-chain discrepancy is also unexplained, leaving readers unable to determine whether it concerns Nvidia equipment, another foreign dependency or a broader interpretation of technological independence.

Acer Veriton AI Mini Workstation Personal Computer GN100-UD11 Series

Acer Veriton AI Mini Workstation Personal Computer GN100-UD11 Series

  • Powerful AI Performance: 1 PFLOPS FP4 AI with NVIDIA GB10 Superchip
  • Pre-installed NVIDIA DGX OS: Optimized for full NVIDIA AI stack
  • High-Performance GPU & CPU: Blackwell GPU with 5th-gen Tensor Cores and 20-core Arm CPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Documentation Needed to Test the Claim

The report can be tested only if Huawei, the publication or another source releases a model card or technical paper, identifies the accelerators and cluster topology, and explains the scope of the phrase without Nvidia. Parameter definitions and reproducible evaluation results would clarify the model-scale claim.

Supply-chain evidence would also need to name the relevant components and suppliers, describe how they were connected to the training run and provide records that independent specialists can examine. Until that happens, the story remains a potentially consequential report, not a verified account of a 505-billion-parameter Nvidia-free training program.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)

  • Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
  • Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
  • Interface: PCIe 3.0 x16 with 250W TDP

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Did Huawei confirm that Pangu Pro has 505 billion parameters?

The supplied material does not establish that Huawei issued or verified the figure. It appears in a published headline, but no Huawei technical paper or model card was included.

Was Pangu Pro trained entirely without Nvidia hardware?

That has not been independently established. The report does not identify the accelerators used or define whether the claim covers only the main run or all experiments, evaluation and deployment.

What does 505 billion parameters mean?

The meaning is unclear because the material does not say whether it describes total parameters or parameters active for each token. That difference can materially affect the computing requirements of mixture-of-experts models.

What supply-chain evidence contradicts the Nvidia-free claim?

No specific evidence is identified. The potential discrepancy could involve chips, fabrication, memory, packaging, networking or software, but the available material names no records or suppliers.

What evidence would verify the report?

Verification would require a hardware inventory and training methodology, architecture details, parameter definitions, performance results and traceable supply records. An independent technical audit would provide stronger support than headline wording alone.

Source: Thorsten Meyer AI

You May Also Like

WinUI 3 Performance: A Leap Forward

Microsoft’s WinUI 3 demonstrates major performance enhancements, reducing latency and resource usage, promising a more responsive Windows app experience.

Why Your Streaming Looks Bad: The 3 Things to Check First

Great streaming starts with checking these 3 key factors—discover what might be causing your poor quality and how to fix it.

AI And Sovereignty: Why Nationality Is Not The Key

An analysis of how legal and geopolitical factors challenge the notion that nationality determines AI sovereignty, focusing on Canadian and European contexts.

Photo GIMP – A Patch for GIMP 3 for Photoshop Users

A new community patch, PhotoGIMP, transforms GIMP 3 to resemble Photoshop, making it easier for users switching from Adobe’s software. Available for Linux, Windows, and macOS.