JOBSEARCHER

HPC Software Engineer

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

A well-funded and rapidly growing technology organization focused on building and scaling high-performance computing (HPC) and cloud infrastructure. The team operates at the forefront of distributed systems, enabling large-scale research, simulation, and data-driven decision-making. Their platform supports mission-critical workloads that drive innovation across advanced technical domains.The RoleWe're looking for a Software Engineer to join a team responsible for building and operating a large-scale HPC platform that powers complex batch workloads on Kubernetes.This role sits at the intersection of distributed systems, cloud infrastructure, and high-performance computing. You'll work on a modern scheduling platform designed to orchestrate workloads across multiple Kubernetes clusters at scale.You'll be joining a highly experienced engineering team working on cutting-edge infrastructure supporting ML and compute-intensive workloads.What You'll DoDesign and build backend systems using Go (Golang)Develop and operate highly scalable, distributed systems for large-scale workloadsBuild and manage containerized applications in Kubernetes environmentsOptimize and manage data across databases (primarily PostgreSQL and other data stores)Troubleshoot and tune Linux-based systems within a compute-heavy environmentDebug networking and system-level issues to improve performance and reliabilityDiagnose complex production issues across infrastructure and application layersApply strong software design principles and computer science fundamentals to your workContribute to CI/CD pipelines and engineering best practicesStay current with new technologies and apply them where relevantWhat We're Looking ForExperience building Kubernetes components (e.g. controllers, operators)Experience with event-driven architectures (Kafka, Pulsar, or similar)Background in high-performance computing, Kubernetes, or workflow orchestration systemsExperience running distributed systems in cloud environments (AWS preferred)Familiarity with monitoring/logging tools (e.g. Prometheus, Grafana)Experience with job scheduling systems (e.g. SLURM or similar)Why JoinWork on cutting-edge infrastructure at scaleTackle complex engineering challenges in distributed systems and HPCCollaborate with a high-caliber, deeply technical teamMake a direct impact on systems that power advanced research and innovationBenefitsLunch stipend (via delivery service)100% employer-covered medical, dental, and vision (for employees + families)16 weeks paid parental leave401(k) with company matchAdditional optional health and wellness benefitsGenerous PTO + company holidays