A declarative workflow orchestrator for massive GPU clusters of any kind, including slurm, K8S and docker. Topology-aware node and GPU placement, multi-node replicas and sweeps, readiness/failure probes, and batch submission — so one descriptor drives hundreds of GPUs instead of a pile of hand-written bash. -
View it on GitHub