Skip to content
GEOGRAPHYUSCanadaMexicoEuropeAsiaPLATFORMSH100H200B200B300GB300 NVL72MI300XSERVICESBare MetalManaged SlurmManaged KubernetesO&MSTANDARDTier IIISingle-tenant by defaultSPEEDDeployments in weeks
GEOGRAPHYUSCanadaMexicoEuropeAsiaPLATFORMSH100H200B200B300GB300 NVL72MI300XSERVICESBare MetalManaged SlurmManaged KubernetesO&MSTANDARDTier IIISingle-tenant by defaultSPEEDDeployments in weeks
MANAGED SLURM / SERVICE

Research-grade
scheduling, operated.

Slurm is the right scheduler for training work and the wrong thing for a research team to maintain. We stand it up on your dedicated cluster, tune it to how your people actually queue jobs, and keep it running.

At a glance
Scheduler
Slurm
Containers
Pyxis / Enroot / Apptainer
Accounting
slurmdbd, fair-share
Runs on
Your dedicated cluster
Upgrades
Managed, tested first
A01 / WHAT'S INCLUDED

What you get.

01

Cluster stood up and tuned

Controller and database nodes, high availability where the cluster warrants it, topology-aware scheduling so multi-node jobs land on the right rails, and prolog/epilog hooks that clean up after failed runs.

02

Partitions, QoS, and fair-share

Queue policy modelled on your team: reserved partitions for production runs, preemptible partitions for experiments, per-group fair-share so no single user starves the cluster.

03

Containers and environments

Pyxis and Enroot, or Apptainer, wired in so srun pulls a container without root. A curated set of base images for PyTorch, JAX, NCCL, and the CUDA track your cluster runs.

04

Job telemetry and accounting

Per-job GPU utilization, memory, and interconnect metrics in Grafana, plus slurmdbd accounting so you can see who used what and where throughput was left on the table.

05

Health checks and node draining

Automated node health checks catch a bad GPU or a flapping link and drain the node before it silently kills a multi-day run. Failed nodes are replaced from spares, not debugged in the queue.

A02 / SPECIFICATION

The technical sheet.

Configured per deployment. These are the ranges we build within; your quote names the exact parts.

Scheduler
Version track
Current stable, tested before rollout
Controller
HA pair available
Accounting DB
slurmdbd with backups
Topology
Rail-aware placement
Preemption
Configured per partition policy
Job environment
Containers
Pyxis / Enroot / Apptainer
MPI
OpenMPI, NCCL-tuned
Modules
Lmod or container-only, your call
Shared FS
Home, scratch, checkpoints
Observability
Metrics
Prometheus / Grafana
GPU
DCGM per-job utilization
Logs
Centralized, retained
Reporting
Monthly utilization summary
Access and policy
Auth
LDAP / OIDC / SSH keys
Login nodes
Dedicated, hardened
Quotas
Per user and per group
Health checks
Automated drain on failure
A03 / WHO IT'S FOR

Where this fits.

Research and training teams

Groups running many multi-node jobs where queue policy and fair-share matter more than a web console.

Teams without a platform group

Organizations whose scientists would otherwise be maintaining the scheduler themselves, badly, at night.

HPC-native organizations

Labs and enterprises already fluent in Slurm who want the same workflow on modern GPU hardware without running it.

A04 / HOW IT'S OPERATED

Who carries the pager.

We do. Every item below is our responsibility for the life of the engagement.

  • Scheduler, controller, and accounting database monitored and backed up 24/7.
  • Version upgrades staged and tested before they touch your production queue.
  • Queue policy reviewed with you as team size and workload mix change.
  • Node health checks, automatic draining, and spare-node replacement.
  • Direct escalation to the engineers who configured your cluster.
A05 / REQUEST A QUOTE

Tell us the shape of the cluster.

An engineer replies with a configuration and a deployment window. No pricing on the website, no sales sequence.

contact@crystalcloud.ai / Los Angeles, CA

Request a cluster quote
~2 MINUTES

Tell us the shape of the cluster. An engineer replies with a configuration and a deployment window, not a sales sequence.

Single-tenant bare metal, managed end to end. Configuration and availability confirmed in writing before you commit.