Staff Engineer, Distributed GPU Clusters
Kindredventures is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will enable research at scale by provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads.
You will extend orchestration, implement topology-aware scheduling, and deliver a self-serve platform for researchers, with strong focus on reliability, observability, and cost efficiency.
#J-18808-Ljbffr