Production K8s
on dedicated GPUs.
The orchestration layer your engineers already know, running on hardware that is exclusively yours. We handle the GPU operator, the node pools, the upgrades, and the observability so your team ships services instead of debugging device plugins.
- Distribution
- Upstream-conformant
- GPU stack
- NVIDIA GPU operator
- Sharing
- MIG or time-slicing
- Scaling
- Autoscaling node pools
- Lifecycle
- Managed upgrades
What you get.
Cluster built and hardened
Highly available control plane, RBAC, network policy, secrets management, and private registries. Access through your existing identity provider, with kubeconfigs scoped per team.
GPU scheduling done properly
NVIDIA GPU operator with device plugin, DCGM exporter, and node feature discovery. MIG partitioning or time-slicing where you want density, whole GPUs where you want determinism.
Distributed training on K8s
Kubeflow training operators, Volcano or Kueue for gang scheduling, and topology-aware placement so multi-node jobs land across the right rails instead of wherever the scheduler felt like.
Ingress, networking, and storage
Ingress controllers, load balancing, and CNI configured for GPU workloads, plus CSI drivers wired to local NVMe and the shared parallel filesystem.
Observability and autoscaling
Prometheus, Grafana, and log aggregation from day one, with node-pool autoscaling and pod-level GPU metrics so capacity decisions come from data rather than guesswork.
The technical sheet.
Configured per deployment. These are the ranges we build within; your quote names the exact parts.
- Distribution
- Upstream-conformant Kubernetes
- Availability
- HA control plane
- Upgrades
- Staged, tested, scheduled with you
- Etcd
- Backed up and monitored
- Auth
- OIDC / SSO, RBAC per team
- Operator
- NVIDIA GPU operator
- Partitioning
- MIG, time-slicing, or whole GPU
- Metrics
- DCGM exporter
- Gang scheduling
- Volcano or Kueue
- Training
- Kubeflow training operators
- CNI
- Configured for GPU east-west traffic
- RDMA
- Available on InfiniBand clusters
- Ingress
- Controller plus load balancing
- CSI
- Local NVMe and shared filesystem
- Registry
- Private, in-cluster or external
- Monitoring
- Prometheus / Grafana
- Logging
- Centralized and retained
- Autoscaling
- Node pool and workload level
- Backups
- Cluster state, scheduled
Where this fits.
Groups serving models in production who need rollouts, autoscaling, and uptime rather than a batch queue.
Teams standardized on Kubernetes who want the same API surface over dedicated GPUs.
Organizations running experiments and production endpoints on the same fleet, with clear isolation between them.
Who carries the pager.
We do. Every item below is our responsibility for the life of the engagement.
- Control plane, etcd, and add-ons monitored, backed up, and patched 24/7.
- Kubernetes and GPU operator upgrades staged in a test cluster before production.
- Node failures drained and replaced from spares without manual intervention.
- Capacity and utilization reviewed monthly with recommendations, not upsells.
- Direct escalation to the engineers who built the cluster.
Tell us the shape of the cluster.
An engineer replies with a configuration and a deployment window. No pricing on the website, no sales sequence.
contact@crystalcloud.ai / Los Angeles, CA