Real hardware.
Root access.
Dedicated GPU nodes on a dedicated fabric, provisioned for one customer at a time. No hypervisor between you and the silicon, no shared interconnect, no throughput that quietly disappears into someone else's job.
- Tenancy
- Single-tenant
- Access
- Root, bare metal
- Fabric
- InfiniBand / RoCEv2
- Scale
- 8 GPUs to multi-MW
- Facilities
- Tier III
What you get.
Dedicated nodes, provisioned to your image
You choose the OS, the CUDA and driver track, and the base image. We provision it, keep it on a tested version track, and hand you root. Nothing is virtualized and nothing is oversubscribed.
Non-blocking GPU fabric
InfiniBand NDR/HDR or RoCEv2 in a rail-optimized, non-blocking topology sized to the cluster. Full-cluster NCCL all-reduce results are part of the acceptance test you sign off on.
Storage that keeps the GPUs fed
Local NVMe scratch on every node plus a shared high-throughput parallel filesystem sized to your dataset and checkpoint pattern, on a separate storage network.
Private networking and isolation
Dedicated VLANs, private interconnect between your nodes, controlled egress, and optional direct connectivity back to your own network or cloud accounts.
Burn-in and acceptance testing
Sustained thermal and load soak, per-link error counters, storage throughput, and GPU health screening before handover. Weak components fail here, not in your first long run.
The technical sheet.
Configured per deployment. These are the ranges we build within; your quote names the exact parts.
- NVIDIA Blackwell
- GB300 NVL72 / B300 / B200 HGX
- NVIDIA Hopper
- H200 / H100 SXM and PCIe
- AMD
- MI300X
- Node form
- 8-GPU HGX or rack-scale NVL72
- CPU / RAM
- Dual socket, 1-2 TB typical
- Compute network
- InfiniBand NDR/HDR or RoCEv2
- Topology
- Rail-optimized, non-blocking
- Per-GPU bandwidth
- Up to 400 Gb/s
- In-node
- NVLink / NVSwitch
- Management
- Separate out-of-band network
- Local scratch
- NVMe, per node
- Shared
- Parallel filesystem, sized per dataset
- Object
- S3-compatible, optional
- Checkpointing
- Throughput sized to run cadence
- Standard
- Tier III certified or equivalent
- Density
- Up to 130 kW per rack
- Cooling
- Air or direct-to-chip liquid
- Redundancy
- N+1 power and cooling
- Geography
- US / Canada / Mexico / Europe / Asia
Where this fits.
Long, tightly-coupled training runs where a single degraded link costs days of wall-clock time.
Teams that cannot run on shared multi-tenant infrastructure for security, residency, or audit reasons.
Groups that want the metal and the fabric, and intend to run their own scheduler and tooling on top.
Who carries the pager.
We do. Every item below is our responsibility for the life of the engagement.
- 24/7 monitoring of GPU health, fabric errors, power draw, and thermals.
- On-site sparing and RMA handling, so a failed node is swapped rather than ticketed.
- Firmware, BMC, and driver lifecycle managed on a tested version track.
- Fabric health checks and link-flap remediation before they show up as job failures.
- A named engineer who knows your cluster, reachable directly, not through a queue.
Tell us the shape of the cluster.
An engineer replies with a configuration and a deployment window. No pricing on the website, no sales sequence.
contact@crystalcloud.ai / Los Angeles, CA