Reka Cloud

Reserve GPU Capacity

Managed training and inference workloads with serverless or dedicated compute. From 64 GPUs to multi-thousand-GPU islands, configured around your workload.

64 to 10k+

GPUs

1 mo+

Terms

24 / 7

Support

Capacity request
1 / 6

What size cluster are you looking for?

Total GPUs across the deployment. A range is fine.

The machines that matter, in stock.

Current NVIDIA training platforms from our own racks and a vetted provider network, with status updated as capacity comes online.

Rack-scale / NVL72 systems

Reservations open

GB300

NVIDIA NVL72

GPUs
72x Blackwell Ultra per rack
Memory
288 GB HBM3e per GPU
Fabric
NVLink 5 / Quantum-X800

Available now

GB200

NVIDIA NVL72

GPUs
72x Blackwell per rack
Memory
192 GB HBM3e per GPU
Fabric
NVLink 5 / Quantum-X800

Node-scale / HGX systems

B200

NVIDIA HGX

Available now

GPUs
8x B200 per node
Memory
180 GB HBM3e per GPU
Fabric
400G Quantum-2 InfiniBand

H200

NVIDIA HGX

Available now

GPUs
8x H200 per node
Memory
141 GB HBM3e per GPU
Fabric
400G Quantum-2 InfiniBand

H100

NVIDIA HGX

Available now

GPUs
8x H100 per node
Memory
80 GB HBM3 per GPU
Fabric
3.2 Tb/s InfiniBand per node

Need another node count, storage ratio, or generation? Share the workload and we will configure the cluster around it.

NVIDIA

First in line for every new generation.

NVIDIA is an investor and our closest engineering partner. In a market where frontier GPUs are allocated before they ship, that partnership is why the fleet above reads available, and it carries through how every cluster is designed, built, and supported.

01

Early allocation

Each new platform generation is allocated to us as it ships, so reservations are fulfilled on silicon timelines, not the resale market.

02

Reference architecture

Every cluster is built and burned in to NVIDIA reference designs and validated before handover.

03

Direct engineering support

When an issue needs NVIDIA's attention, it reaches NVIDIA engineering through the partnership, not a vendor queue.

Customer zero was our own frontier lab.

Before this platform ran an external workload, it trained Reka's multimodal frontier models end to end. Link flaps, straggler nodes, checkpoint stalls, and silent data corruption were found and fixed on our runs, not yours.

10,000+

GPUs in one training fabric

99.2%

measured training goodput

24 / 7

operator coverage

Built for training. Tuned for goodput.

The machine is the floor. What you are buying is everything that keeps it busy.

Bare metal

No hypervisor between your job and the silicon. You get the whole machine, and every FLOP you pay for.

Non-blocking fabric

Full-bisection, rail-optimized InfiniBand built for collective operations at cluster scale.

Parallel storage

High-throughput storage sits beside the compute so checkpoints do not gate step time.

Scheduler ready

Slurm or Kubernetes configured, images loaded, and a reference job completed before handoff.

Goodput SLAs

Hot spares, uptime commitments, and reporting on the hours your workload actually trained.

24/7 training engineers

Around-the-clock coverage from engineers who operate frontier training clusters themselves.

Reservation to first job, in four steps.

01

Scope

We size compute, fabric, storage, and regions around your workload.

02

Reserve

Lock the capacity against a firm go-live date, from one month to multi-year.

03

Provision

We configure your scheduler, images, networking, access, and observability.

04

Operate

Start on a burned-in cluster with root access and live engineering support.

What infrastructure teams ask first.

Is this bare metal or virtualized?+

Bare metal. Your jobs run directly on the machines, with root access and the full fabric available to your workload.

How quickly can a cluster be live?+

Standard HGX blocks can be ready in days. Custom scheduler, image, storage, and networking layouts typically take a few weeks after signature.

What reservation terms are available?+

One month to multi-year. Reserve for a single training run or lock a longer allocation with a defined go-live date.

What is included in the price?+

Compute, power, cooling, storage, fabric, and support are included in the GPU-hour rate, with no ingress or egress fees.

What happens when nodes fail?+

Hot spares remain in the same fabric, failed nodes are swapped quickly, and downtime is credited.

Who operates the infrastructure?+

Reka Cloud does, 24/7. The same engineering discipline used to train frontier multimodal models is applied to your cluster.

Put a go-live date on the cluster.

Share the workload and timeline. We will return with concrete capacity, topology, and commercial options.

Reserve capacity