Someone has to
carry the pager.
Buying GPUs is the easy part. Keeping a fabric healthy, drivers current, and utilization high for three years is the expensive part. We take that on under an SLA, on clusters we deploy, or on hardware you already own.
- Coverage
- 24/7/365 NOC
- Response
- SLA-backed, named times
- Hardware
- On-site sparing and RMA
- Scope
- Our fleet or yours
- Reporting
- Monthly health report
What you get.
24/7 monitoring and on-call
GPUs, fabric, storage, power, and thermals watched continuously with alerting tuned to catch degradation before it becomes a failed job. A human is awake and responsible at all times.
Hardware sparing and RMA
Spares held at the facility so a failed GPU, PSU, or optic is swapped in hours rather than waiting on a vendor cycle. We own the RMA paperwork and the follow-up.
Firmware and driver lifecycle
BMC, BIOS, NIC, switch, and GPU driver versions kept on a tested track. Updates are validated in a staging window before they touch production, and we keep the version history.
Fabric health and remediation
Continuous link error and congestion monitoring, cable and optic replacement, and routing checks. Link flaps get fixed as maintenance events, not discovered as mystery job failures.
Capacity, thermal, and utilization reporting
A monthly report with real numbers: availability, incident history, GPU utilization, thermal headroom, and where throughput is being lost. Recommendations included, upsells not.
The technical sheet.
Configured per deployment. These are the ranges we build within; your quote names the exact parts.
- NOC
- 24/7/365
- Escalation
- Named engineer, direct line
- Response
- Severity-tiered, in the SLA
- Maintenance
- Scheduled windows, agreed with you
- Spares
- Held on site
- Swap
- Hours, not vendor cycles
- RMA
- Handled end to end
- Optics and cabling
- Stocked and replaced
- Drivers / CUDA
- Tested version track
- Firmware
- BMC, BIOS, NIC, switch
- Scheduler
- Slurm or Kubernetes, if in scope
- Rollback
- Version history retained
- GPU
- DCGM health and utilization
- Fabric
- Link errors, congestion, routing
- Environment
- Power draw and thermals
- Reporting
- Monthly, with incident history
Where this fits.
Teams that bought hardware and discovered that operating it is a full-time engineering function.
Providers who need a operations bench that scales with the fleet without hiring one from scratch.
Organizations running GPUs in their own or a colocated facility who want vendor-grade operations over them.
Who carries the pager.
We do. Every item below is our responsibility for the life of the engagement.
- Continuous monitoring across GPU, fabric, storage, power, and thermal telemetry.
- Incident response to a written SLA with severity tiers and named response times.
- On-site spares, component swaps, and full RMA ownership.
- Tested firmware and driver rollouts, with rollback paths retained.
- Monthly health, availability, and utilization reporting with concrete recommendations.
Tell us the shape of the cluster.
An engineer replies with a configuration and a deployment window. No pricing on the website, no sales sequence.
contact@crystalcloud.ai / Los Angeles, CA