Together AI has shipped a set of operational improvements to its GPU Clusters product, grouped into platform health and operational control. On the health side: passive health checks now continuously monitor live workloads for GPU failures, thermal throttling, and Xid errors; auto node repair surfaces recommended remediation actions (reboot, reprovision, failover, remove) for operator approval; and a rebuilt Slurm-on-Kubernetes stack (based on Slinky) fixes zombie processes, durable job accounting, reliable process cleanup, and accurate GPU state after reschedules. On the control side: a new cluster details view shows node health, live GPU utilization, and event history; external OIDC support enables per-user Kubernetes RBAC via existing identity providers (Google, Okta, Auth0, Entra ID); startup scripts allow self-serve node customization at boot, job start, and job end; and acceptance testing is now opt-in at cluster creation for larger or longer-running workloads.
Table of contents
Catching and fixing failures as they happenTogether Slurm-on-K8s 2.0: The Future of Slurm on KubernetesOperational controlWhat changes operationallyWhat’s nextShare this post