{"schemaVersion":"jobsearcher.job.v1","id":"e9af710f96b034ba2838c3cc","url":"https://jobsearcher.com/jobs/e9af710f96b034ba2838c3cc","canonicalUrl":"https://jobsearcher.com/jobs/e9af710f96b034ba2838c3cc","title":"Automated Testing Engineer, Compute","description":"Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.\n\nWe're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.\n\nWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.\n\nIf you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.\n\nAbout the Role:\n\nAs an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated integration testing framework to validate high-performance GPU training, ensuring that distributed workloads scale efficiently across multiple virtualized nodes. Your role is critical in ensuring the stability of the low-level infrastructure and validating the interconnect fabric that powers the world’s most demanding AI and HPC applications.\n\nSan Francisco, Sunnyvale (Onsite)\n\nWhat You’ll Be Working On:\n\nCI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.\n\nMulti-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.\n\nCluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.\n\nInterconnect & Fabric Testing: Validate high-speed interconnects—including NVLink, Infinity Fabric, InfiniBand, and RoCE—within virtualized environments to ensure low-latency, high-bandwidth communication.\n\nCollective Communication Benchmarking: Architect and run comprehensive test suites using nccl-tests and rccl-tests (e.g., AllReduce, AllGather) to verify performance across node boundaries.\n\nPerformance Bottleneck Analysis: Perform deep-dive analysis of regressions in CPU performance and multi-node communication, identifying root causes across the guest OS, hypervisor, and physical fabric.\n\nCreate automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.\n\nWhat You’ll Bring to the Team:\n\nEducation & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.\n\nExperience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.\n\nWorking knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.\n\nCI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.\n\nAutomation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.\n\nDistributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.\n\nNetworking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.\n\nSystem Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).\n\nBonus Points:\n\nExperience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.\n\nFamiliarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).\n\nKnowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).\n\nBenefits:\n\nCompetitive compensation and equity packages\n\nRestricted Stock Units\n\nPaid time off, paid holidays & leave of absence programs\n\nComprehensive health, dental & vision insurance\n\nEmployer contributions to HSA account\n\nPaid parental leave\n\nPaid life insurance, short-term and long-term disability\n\nProfessional development & tuition reimbursement\n\nMental health & wellness support\n\nCommuter benefits (parking & transit)\n\nCell phone stipend\n\n401(k) Retirement plan with company match up to 4% of salary\n\nVolunteer time off\n\nGlobal travel insurance & emergency assistance\n\nDaily meals allowance\n\nAdditional perks & programs specific to location\n\nCompensation:\n\nCompensation will be paid in the range of $172,500 - $210,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.\n\nCrusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.","company":"Crusoe","rawCompany":"crusoe","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-25T04:42:20.413Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1253.00","title":"Software Quality Assurance Analysts and Testers","slug":"software-quality-assurance-analysts-and-testers"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Automated Testing Engineer, Compute","description":"Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.\n\nWe're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.\n\nWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.\n\nIf you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.\n\nAbout the Role:\n\nAs an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated integration testing framework to validate high-performance GPU training, ensuring that distributed workloads scale efficiently across multiple virtualized nodes. Your role is critical in ensuring the stability of the low-level infrastructure and validating the interconnect fabric that powers the world’s most demanding AI and HPC applications.\n\nSan Francisco, Sunnyvale (Onsite)\n\nWhat You’ll Be Working On:\n\nCI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.\n\nMulti-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.\n\nCluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.\n\nInterconnect & Fabric Testing: Validate high-speed interconnects—including NVLink, Infinity Fabric, InfiniBand, and RoCE—within virtualized environments to ensure low-latency, high-bandwidth communication.\n\nCollective Communication Benchmarking: Architect and run comprehensive test suites using nccl-tests and rccl-tests (e.g., AllReduce, AllGather) to verify performance across node boundaries.\n\nPerformance Bottleneck Analysis: Perform deep-dive analysis of regressions in CPU performance and multi-node communication, identifying root causes across the guest OS, hypervisor, and physical fabric.\n\nCreate automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.\n\nWhat You’ll Bring to the Team:\n\nEducation & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.\n\nExperience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.\n\nWorking knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.\n\nCI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.\n\nAutomation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.\n\nDistributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.\n\nNetworking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.\n\nSystem Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).\n\nBonus Points:\n\nExperience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.\n\nFamiliarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).\n\nKnowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).\n\nBenefits:\n\nCompetitive compensation and equity packages\n\nRestricted Stock Units\n\nPaid time off, paid holidays & leave of absence programs\n\nComprehensive health, dental & vision insurance\n\nEmployer contributions to HSA account\n\nPaid parental leave\n\nPaid life insurance, short-term and long-term disability\n\nProfessional development & tuition reimbursement\n\nMental health & wellness support\n\nCommuter benefits (parking & transit)\n\nCell phone stipend\n\n401(k) Retirement plan with company match up to 4% of salary\n\nVolunteer time off\n\nGlobal travel insurance & emergency assistance\n\nDaily meals allowance\n\nAdditional perks & programs specific to location\n\nCompensation:\n\nCompensation will be paid in the range of $172,500 - $210,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.\n\nCrusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.","datePosted":"2026-08-25T04:42:20.413Z","dateModified":"2026-08-25T04:42:20.413Z","hiringOrganization":{"@type":"Organization","name":"Crusoe","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e9af710f96b034ba2838c3cc"},"url":"https://jobsearcher.com/jobs/e9af710f96b034ba2838c3cc"}}