{"schemaVersion":"jobsearcher.job.v1","id":"d3bd2b5979e5c722eba0a42b","url":"https://jobsearcher.com/jobs/d3bd2b5979e5c722eba0a42b","canonicalUrl":"https://jobsearcher.com/jobs/d3bd2b5979e5c722eba0a42b","title":"Senior GPU Systems Engineer","description":"Design and operate large-scale GPU clusters running NVIDIA H100 and B200 hardware for AI inference and training workloads.\nFollow us on LinkedIn\nEngineering Remote full time Remote-eligible Competitive\nAbout the role\nYobitel builds and operates the AI-native compute layer that our customers' training and inference workloads run on. We own the hardware rather than reselling someone else's: H100 fleets in production today, B200 capacity landing next. When a rack degrades at 3am, it is our pager and our problem.\nWe are hiring a senior systems engineer to keep that fleet fast, healthy and full. This is a deep systems role that sits close to the metal, spanning firmware and driver stacks, NUMA and topology tuning, RDMA fabric behaviour, and the performance regressions that only ever show up under real multi-node load. You will work alongside the platform and ML infrastructure teams, and your work sets the ceiling on what every workload above you can achieve.\nWhat success looks like\nIn your first 90 days, you have commissioned a rack end to end and written the runbook that lets someone else repeat it.\nNode-level performance regressions get caught by automated checks before customers feel them.\nFleet utilisation and time-to-recovery both improve, and you can show the numbers.\nResponsibilities\nCommission new GPU racks end to end: burn-in, hardware validation, firmware and driver baselines, topology verification, acceptance into the fleet.\nDiagnose failures at the hardware boundary, covering GPU faults, NVLink and NVSwitch degradation, PCIe and RDMA fabric issues, thermal and power events.\nTune node performance for real workloads: NUMA placement, CPU and interrupt affinity, huge pages, GPUDirect and RDMA paths.\nBuild and maintain automated health checking and burn-in tooling so a bad node is quarantined before a customer schedules onto it.\nOwn driver, CUDA toolkit and firmware version policy across the fleet, including safe staged rollout and rollback.\nInvestigate multi-node training and inference performance regressions with the ML infrastructure team, and drive them to root cause.\nCarry the on-call rotation for fleet health, and write the runbook the first time you solve something by hand.\nRequirements\nSubstantial production experience operating GPU or HPC compute at fleet scale, rather than single workstations.\nStrong Linux systems fundamentals: kernel and userspace boundaries, cgroups, scheduling, memory, storage and network stacks.\nHands-on familiarity with the NVIDIA stack: driver and CUDA toolkit management, NVML, DCGM, MIG, and GPU telemetry.\nPractical experience debugging high-performance networking, whether InfiniBand, RoCE or RDMA, including fabric-level diagnosis.\nComfort automating operational work in Python, Go or Bash, and treating tooling as a first-class deliverable.\nExperience carrying production on-call for infrastructure with real availability commitments.\nClear written communication, since much of this role is documenting what you learned so the next person does not relearn it.\nNice to have\nExperience bringing up a brand-new GPU generation, including the surprises that only appear on new silicon.\nFamiliarity with Kubernetes device plugins, GPU operators, or scheduling GPUs in a containerised environment.\nBackground in data-centre-side concerns such as rack power budgeting, liquid cooling or high-density thermal design.\nContributions to relevant open-source projects, or published benchmarking work.\nRole reference: YJ-00001","company":"Yobitel Communications","rawCompany":"yobitel communications","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-07T10:59:04.384Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior GPU Systems Engineer","description":"Design and operate large-scale GPU clusters running NVIDIA H100 and B200 hardware for AI inference and training workloads.\nFollow us on LinkedIn\nEngineering Remote full time Remote-eligible Competitive\nAbout the role\nYobitel builds and operates the AI-native compute layer that our customers' training and inference workloads run on. We own the hardware rather than reselling someone else's: H100 fleets in production today, B200 capacity landing next. When a rack degrades at 3am, it is our pager and our problem.\nWe are hiring a senior systems engineer to keep that fleet fast, healthy and full. This is a deep systems role that sits close to the metal, spanning firmware and driver stacks, NUMA and topology tuning, RDMA fabric behaviour, and the performance regressions that only ever show up under real multi-node load. You will work alongside the platform and ML infrastructure teams, and your work sets the ceiling on what every workload above you can achieve.\nWhat success looks like\nIn your first 90 days, you have commissioned a rack end to end and written the runbook that lets someone else repeat it.\nNode-level performance regressions get caught by automated checks before customers feel them.\nFleet utilisation and time-to-recovery both improve, and you can show the numbers.\nResponsibilities\nCommission new GPU racks end to end: burn-in, hardware validation, firmware and driver baselines, topology verification, acceptance into the fleet.\nDiagnose failures at the hardware boundary, covering GPU faults, NVLink and NVSwitch degradation, PCIe and RDMA fabric issues, thermal and power events.\nTune node performance for real workloads: NUMA placement, CPU and interrupt affinity, huge pages, GPUDirect and RDMA paths.\nBuild and maintain automated health checking and burn-in tooling so a bad node is quarantined before a customer schedules onto it.\nOwn driver, CUDA toolkit and firmware version policy across the fleet, including safe staged rollout and rollback.\nInvestigate multi-node training and inference performance regressions with the ML infrastructure team, and drive them to root cause.\nCarry the on-call rotation for fleet health, and write the runbook the first time you solve something by hand.\nRequirements\nSubstantial production experience operating GPU or HPC compute at fleet scale, rather than single workstations.\nStrong Linux systems fundamentals: kernel and userspace boundaries, cgroups, scheduling, memory, storage and network stacks.\nHands-on familiarity with the NVIDIA stack: driver and CUDA toolkit management, NVML, DCGM, MIG, and GPU telemetry.\nPractical experience debugging high-performance networking, whether InfiniBand, RoCE or RDMA, including fabric-level diagnosis.\nComfort automating operational work in Python, Go or Bash, and treating tooling as a first-class deliverable.\nExperience carrying production on-call for infrastructure with real availability commitments.\nClear written communication, since much of this role is documenting what you learned so the next person does not relearn it.\nNice to have\nExperience bringing up a brand-new GPU generation, including the surprises that only appear on new silicon.\nFamiliarity with Kubernetes device plugins, GPU operators, or scheduling GPUs in a containerised environment.\nBackground in data-centre-side concerns such as rack power budgeting, liquid cooling or high-density thermal design.\nContributions to relevant open-source projects, or published benchmarking work.\nRole reference: YJ-00001","datePosted":"2026-08-07T10:59:04.384Z","dateModified":"2026-08-07T10:59:04.384Z","hiringOrganization":{"@type":"Organization","name":"Yobitel Communications","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d3bd2b5979e5c722eba0a42b"},"url":"https://jobsearcher.com/jobs/d3bd2b5979e5c722eba0a42b"}}