Freyr Technology AI Book a consult

NVIDIA HGX B300 · Blackwell Ultra

The B300.
In Vancouver.

Eight NVIDIA Blackwell Ultra GPUs on a single baseboard, wired by fifth generation NVLink. The most capable AI node NVIDIA ships today, deployed in Vancouver and contracted to you.

An NVIDIA HGX B300 system, front three-quarter view
NVIDIA HGX B300. Eight Blackwell Ultra SXM GPUs on a single baseboard, all-to-all through a dual NVLink 5 switch system at 14.4 TB/s aggregate.

8x

Blackwell Ultra GPUs per node

2.1TB

HBM3e memory per node

144PF

FP4 tensor core, sparse

14.4TB/s

Aggregate NVLink bandwidth

The hardware

Eight GPUs that work as one.

The HGX B300 puts eight Blackwell Ultra SXM modules on a single baseboard and connects every GPU to every other GPU through a dual NVLink 5 switch system. Each link runs at 1.8 TB/s, and the fabric carries 14.4 TB/s in aggregate.

Because the interconnect is not a bottleneck, a model spread across all eight GPUs behaves more like one device with 2.1 terabytes of memory than like eight devices passing tensors over a network. That matters most for large mixture of experts models, long context windows and large training runs.

Blackwell Ultra adds roughly twice the attention layer performance of standard Blackwell and 1.5 times the dense FP4 throughput. For reasoning models and long sequence inference, that is where the gain shows up.

Aerial view of a city grid, roads meeting at intersections

On a B300 node the interconnect is wide enough that the eight GPUs behave as one pool of memory rather than eight separate destinations.

Specifications

NVIDIA HGX B300, per node.

Platform figures as published by NVIDIA for the HGX B300. Freyr deployments are configured to customer specification, so CPU, local storage and fabric topology are set during the design phase rather than fixed here.

NVIDIA HGX B300 platform specifications
Specification HGX B300
Form factor8x NVIDIA Blackwell Ultra SXM
Total GPU memory2.1 TB HBM3e
FP4 tensor core144 PFLOPS sparse · 108 PFLOPS dense
FP8 / FP6 tensor core72 PFLOPS sparse
INT8 tensor core3 POPS sparse
FP16 / BF16 tensor core36 PFLOPS sparse
TF32 tensor core18 PFLOPS sparse
FP32600 TFLOPS
FP64 / FP64 tensor core10 TFLOPS
NVLink generationFifth generation, NVLink 5 Switch
NVLink GPU-to-GPU bandwidth1.8 TB/s
Total NVLink bandwidth14.4 TB/s
Networking bandwidth1.6 TB/s · 8x NVIDIA ConnectX-8
Attention performance2x NVIDIA Blackwell
TenancySingle tenant, bare metal, no hypervisor
LocationVancouver, British Columbia · Tier III
TermReserved, one to five years
DeploymentApproximately five weeks from signature

Platform figures per the NVIDIA HGX platform specifications. Sparse and dense figures are noted where NVIDIA distinguishes them. The final four rows describe the Freyr deployment rather than the NVIDIA platform.

What Freyr adds

The chip is the same anywhere. The terms are not.

Anyone with an allocation can quote you a B300. What differs is who else is on the box, where it sits, and whether it is still yours in eighteen months.

01

Bare metal, single tenant

No hypervisor between your code and the hardware, and no other tenant on your baseboard. Nothing else competes for the node, so it performs the same in year three as on day one.

02

Vancouver facilities

Tier III colocation in Vancouver, on a grid that is roughly ninety percent hydroelectric. Data processed on Freyr hardware stays in the facility you contracted.

03

Configured before it ships

Storage profile, interconnect topology and CPU pairing are set during design rather than worked around after delivery. Roughly five weeks from signature to a cluster you can log into.

Use cases

Three workloads worth the node.

Reasoning and agentic inference

Blackwell Ultra roughly doubles attention layer throughput against standard Blackwell. Long chains of thought and multi step agent loops spend most of their time in that path, so the gain shows up in tokens per second.

Frontier and foundation model training

72 PFLOPS of FP8 and 2.1 TB of pooled HBM3e per node, with 14.4 TB/s of NVLink behind it. Models that would need heavy quantisation or sharding elsewhere fit and train here without it.

Regulated and sovereign workloads

Health, financial and public sector work that cannot leave the country it is regulated in, or share a physical host. Single tenant hardware in a named Vancouver facility gives your privacy officer a clear answer.

Tell us the model, the scale and the timeline.

We will come back with a node count and a configuration, and tell you if the B300 is not the right part for the job. Capacity goes to whoever commits first, and the 2026 deployment window is open.

John Kinash · Business Development, North America · john.kinash@freyrtech.ai · 778 872 8288