Skip to content
GEOGRAPHYUSCanadaMexicoEuropeAsiaPLATFORMSH100H200B200B300GB300 NVL72MI300XSERVICESBare MetalManaged SlurmManaged KubernetesO&MSTANDARDTier IIISingle-tenant by defaultSPEEDDeployments in weeks
GEOGRAPHYUSCanadaMexicoEuropeAsiaPLATFORMSH100H200B200B300GB300 NVL72MI300XSERVICESBare MetalManaged SlurmManaged KubernetesO&MSTANDARDTier IIISingle-tenant by defaultSPEEDDeployments in weeks
ENTERPRISE GPU AS A SERVICE / BARE METAL / SLURM / KUBERNETES / O&M

Your own GPU cluster.
Fully managed.

Dedicated bare-metal GPU infrastructure with the scheduler, the orchestration, and the 24/7 operations included. You get root on real hardware; we keep it healthy, fast, and full.

Single-tenant by default, no hypervisor, no shared GPUs, no noisy neighbours. Deployed in Tier III facilities across the US, Canada, Mexico, Europe, and Asia.

Request a cluster quote
~2 MINUTES

Tell us the shape of the cluster. An engineer replies with a configuration and a deployment window, not a sales sequence.

Single-tenant bare metal, managed end to end. Configuration and availability confirmed in writing before you commit.

Single-tenant bare metal

Dedicated nodes and a dedicated fabric. Your throughput is never someone else's job queue.

The managed layer included

Slurm or Kubernetes stood up, tuned, and operated, so your researchers ship instead of babysitting infrastructure.

Weeks, not quarters

Design, build, burn-in, acceptance test, handover. A real deployment window in writing.

S01 / SERVICES

The cluster, and everything that keeps it running.

Take the hardware alone, or take the whole stack. Most teams take the whole stack, because the scheduler and the 3 a.m. link flap are the parts that actually cost engineering time.

01

Bare Metal GPU

Dedicated, single-tenant GPU nodes with root access and no virtualization tax. Non-blocking InfiniBand or RoCEv2, NVMe local scratch, and shared high-throughput storage.

Bare Metal GPU detail

Tenancy
Single-tenant, dedicated
Platforms
H100 / H200 / B200 / B300 / GB300 NVL72 / MI300X
Fabric
InfiniBand or RoCEv2, non-blocking
02

Managed Slurm

A research-grade scheduler stood up, tuned, and operated for you: partitions, QoS, fair-share accounting, container support, and job telemetry.

Managed Slurm detail

Scheduler
Slurm, tuned per workload
Containers
Pyxis / Enroot / Apptainer
Included
Accounting, telemetry, upgrades
03

Managed Kubernetes

Production Kubernetes with the NVIDIA GPU operator, autoscaling node pools, ingress, and full observability. Cluster lifecycle and upgrades are ours to carry.

Managed Kubernetes detail

Distribution
Upstream-conformant K8s
GPU
GPU operator, MIG or time-slicing
Included
Observability, ingress, lifecycle
04

Operations & Maintenance

24/7 NOC coverage, hardware sparing and RMA, firmware and driver lifecycle, fabric health, and thermal monitoring, on our fleet or on hardware you already own.

Operations & Maintenance detail

Coverage
24/7 NOC, SLA-backed
Hardware
On-site sparing and RMA handling
Scope
Our fleet or your own
S02 / CLUSTER PROFILES

Configured to the workload.

Three shapes cover most of what we deploy. Every one is delivered single-tenant, burned in, and acceptance tested before handover.

P.01

Training block

For pre-training and large-scale training runs that need tightly-coupled, high-density compute.

Hardware
H200 / B200 / B300-class HGX or NVL72 rack-scale
Fabric
InfiniBand or RoCEv2 non-blocking
Density
30-130 kW per rack, liquid
Managed layer
Slurm, typically
P.02

Inference fleet

For serving models in production, where scale and uptime matter more than peak interconnect.

Hardware
H200 / MI300X-class, 141-192 GB
Cooling
Air or liquid
Scale
64-1,024 GPUs
Managed layer
Kubernetes, typically
P.03

Fine-tuning cluster

For teams adapting existing models on a modest, air-cooled footprint with a shorter commitment.

Hardware
H100 / L40S-class, 8-128 GPUs
Cooling
Air-cooled friendly
Managed layer
Slurm or Kubernetes
Current platforms
GB300 NVL72B300 HGXB200 HGXH200 HGXH100 SXM/PCIeMI300Xnew-gen by allocation
Representative deployments / anonymized

256 H200-class GPUs on InfiniBand, managed Slurm, single-tenant Tier III hall.

Rack-scale NVL72 training block, liquid cooled, full O&M and fabric monitoring.

512-GPU inference fleet on managed Kubernetes with autoscaling node pools.

Illustrative configurations to gauge fit. Customer identities and pricing are never disclosed publicly.

S03 / HOW WE DELIVER

Nothing is handed over untested.

A cluster that passes a smoke test is not a cluster that survives a three-week training run. Four stages, each with an artifact you receive.

D.01

DESIGNConfigured against the workload

Node count, GPU platform, fabric topology, storage throughput, power density, and the managed layer, sized to the jobs you actually intend to run. Written down before anything is ordered.

YOU RECEIVE
CONFIGURATION SHEET
DEPLOYMENT WINDOW
D.02

BUILDRacked, cabled, burned in

Hardware racked and cabled in a Tier III facility matched to your density and geography, then burned in under sustained load until the weak components fail here rather than in your first run.

YOU RECEIVE
BURN-IN REPORT
CABLING AND TOPOLOGY MAP
D.03

VALIDATEAcceptance tested end to end

NCCL all-reduce across the full cluster, per-link bandwidth and error counters, storage throughput, thermal soak, and scheduler smoke tests. You see the numbers before you accept the cluster.

YOU RECEIVE
NCCL / FABRIC RESULTS
ACCEPTANCE CRITERIA SIGN-OFF
D.04

OPERATERun, monitored, kept current

24/7 monitoring of GPUs, fabric, power, and thermals. Sparing and RMA handled on site, firmware and drivers kept on a tested track, and a named engineer who knows your cluster.

YOU RECEIVE
24/7 NOC AND SLA
MONTHLY HEALTH REPORT

This is what prevents silent throughput loss, mystery job failures, and the slow drift that turns a fast cluster into an average one.

S04 / FACILITY STANDARD

Your cluster picks the building.

We deploy into independent Tier III facilities qualified for your specific cluster: your density, your geography, your term. Every facility clears the same bar before it carries a deployment.

Tier III

CERTIFIED OR VERIFIED EQUIVALENT DESIGN / QUALIFICATION FILE ON RECORD

N+1

CONCURRENTLY MAINTAINABLE POWER AND COOLING / TOPOLOGY REVIEWED

Density matched

AIR OR LIQUID, MATCHED TO THE CLUSTER / COOLING SPEC ON FILE

Carrier-neutral

FIBER, PROVIDERS AND LATENCY DOCUMENTED / NETWORK SHEET
Qualification file, one per facility
F.01
Tier certification / design dossier
REVIEWED
Power headroom / N+1 topology
REVIEWED
Density and cooling limits
DOCUMENTED
Fiber providers and latency
DOCUMENTED
Remote hands and escalation path
ON FILE
Site walk
BEFORE DEPLOYMENT

Facility detail shared under NDA / metros disclosed there

Multi-country
deployment footprint
US / Canada / Mexico / Europe / Asia
Single
tenant by default
no shared GPUs, no hypervisor
S05 / TIMELINE

Weeks, not procurement quarters.

WEEK 0

Requirements and configuration

One working session: workload, scale, density, interconnect, storage, term length, and the managed layer you want. You leave with a written configuration.

WEEK 1

Quote and cluster spec

A configuration sheet with the numbers in it: node counts, fabric, storage, the managed layer, the SLA, and the deployment window. Every number is in the quote; none are on the website.

WEEKS 2-4

Build, validate, handover

Racked, burned in, and acceptance tested in a Tier III facility. Bare-metal access, your scheduler or ours, and a direct line to the engineers who stood it up. Newest-generation platforms depend on allocation timing, and we say so in the quote.

Security posture

Single-tenant by default, so your workloads never share a hypervisor with strangers. Clusters run bare metal in facilities whose SOC 2 or ISO 27001 attestations we review during qualification and make available under NDA. Private networking, dedicated storage, and customer-controlled key material where the workload requires it.

OPERATIONS FOR YOUR OWN FLEET

Already own the hardware?

Owning GPUs is the easy part; keeping a fabric healthy at 3 a.m. is not. We take over operations for fleets you already own, deployment, scheduler, monitoring, sparing, firmware, and the on-call rotation, under an SLA, in your facility or one we qualify for you.

FOR DATA CENTER OPERATORS

We bring our customers to your floor.

Crystal Cloud contracts with the end customer and deploys directly into your facility. Requirements run from 2 to 100 MW, single-tenant, Tier III, air and liquid. You get one credit-verified counterparty and a filled hall, not a stream of leads to chase.

S06 / QUESTIONS

Questions.

Q.01

What does Crystal Cloud actually provide?

Enterprise GPU as a service: dedicated bare-metal GPU clusters with high-speed interconnect, plus the managed layer on top, Slurm, Kubernetes, or both, and 24/7 operations and maintenance. You get root access to your own hardware and an engineering team that runs it.

Q.02

Is the hardware shared with other customers?

No. Clusters are single-tenant bare metal by default. No hypervisor, no noisy neighbours, no shared GPUs. The nodes, the fabric, and the storage are yours for the term.

Q.03

How fast can a cluster be delivered?

Weeks, not procurement quarters. Design and configuration in week 0, build and burn-in through weeks 1 to 3, acceptance testing and handover after that. Newest-generation platforms depend on allocation timing, and we say so up front.

Q.04

Why is there no pricing on the site?

Every cluster is configured to the workload, the density, and the term, so the numbers live in a written quote rather than on a webpage. Tell us the shape of the cluster and an engineer replies with a configuration and a deployment window.

S07 / CONTACT

Start with the cluster you need.

Request a cluster quote
~2 MINUTES

Tell us the shape of the cluster. An engineer replies with a configuration and a deployment window, not a sales sequence.

Single-tenant bare metal, managed end to end. Configuration and availability confirmed in writing before you commit.