{"schemaVersion":"jobsearcher.job.v1","id":"15f79cfb9ff1fa79d3bd09f0","url":"https://jobsearcher.com/jobs/15f79cfb9ff1fa79d3bd09f0","canonicalUrl":"https://jobsearcher.com/jobs/15f79cfb9ff1fa79d3bd09f0","title":"AI Kernel Cluster Engineer","description":"AI Kernel Cluster Engineer\n\nPosition Overview\n\nI'm partnering with a rapidly growing AI infrastructure company supporting large-scale GPU environments that power AI training and inference workloads.\n\nThis is an opportunity to take ownership of the health, stability, and performance of production GPU clusters running mission-critical AI workloads. You'll work at the intersection of Linux systems, GPU infrastructure, containerization, and high-performance networking, helping ensure advanced AI environments operate reliably and efficiently at scale. Working closely with engineering and operations teams, you'll play a key role in maintaining the infrastructure that powers next-generation AI applications.\n\nKey Responsibilities\n\nOwn the day-to-day health, performance, resource allocation, and operations of production GPU clusters.\nManage Ubuntu Linux systems, including kernel-level tuning, driver management, patching, and OS troubleshooting.\nAdminister and troubleshoot LXC container environments supporting AI workloads.\nMonitor and maintain InfiniBand and RoCEv2 networking across GPU cluster environments.\nDiagnose and resolve complex issues spanning Linux systems, containers, GPUs, and cluster networking, including escalated production incidents.\n\nQualifications\n\nRequired\n\n3+ years of experience supporting GPU clusters, HPC environments, or large-scale infrastructure platforms.\nDeep experience with Ubuntu Linux, including kernel-level operations, system tuning, driver management, and OS troubleshooting.\nHands-on experience managing LXC containers in production environments.\nExperience supporting GPU cluster networking, including InfiniBand, RoCEv2, and other high-performance networking technologies.\nExperience troubleshooting issues across Linux systems, containers, networking, and GPU infrastructure.\nFamiliarity with NVIDIA technologies including DCGM, UFM, NCCL, CUDA, and SHARP, as well as monitoring platforms such as Prometheus and Grafana.\nExperience with SLURM, Kubernetes, or similar workload scheduling platforms.\nProficiency with Bash, Python, or similar scripting languages for automation and diagnostics.\n\nNice to Have\n\nFamiliarity with AI/ML frameworks such as PyTorch, TensorFlow, or JAX.\nExperience with Ansible, Terraform, or similar infrastructure automation platforms.\nFamiliarity with TensorRT, ONNX, or AI model optimization tooling.\nAI, cloud, Linux, or networking certifications.\nExperience supporting large-scale AI infrastructure or GPU cloud environments.\n\nBenefits\n\n$185,000 to $225,000 base salary\n30% to 50% annual performance bonus\nRestricted Stock Units (RSUs)\nThis opportunity is ideal for an engineer who enjoys working close to the operating system, networking, and hardware layers while supporting high-performance GPU environments running large-scale AI workloads. You'll have the opportunity to solve complex infrastructure challenges, work with cutting-edge AI compute environments, and play a key role in supporting the next generation of AI applications.\n\n- For this position, you must be currently authorized to work in the United States without the need for sponsorship for a non-immigrant visa. CyberCoders will consider for Employment in the City of Los Angeles qualified Applicants with Criminal Histories in a manner consistent with the requirements of the Los Angeles Fair Chance Initiative for Hiring (Ban the Box) Ordinance.This job was first posted by CyberCoders on 08/19/2026 and applications will be accepted on an ongoing basis until the position is filled or closed.Everforth CyberCoders is proud to be an Equal Opportunity Employer\n\nAll qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, sexual orientation, gender identity or expression, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, status as a crime victim, disability, protected veteran status, or any other characteristic protected by law. Our hiring process includes AI screening for keywords and minimum qualifications, and a virtual recruiter as part of the application process. A human recruiter reviews all results. Click here for details on our virtual recruiter . Everforth CyberCoders will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable state and local law, including but not limited to the Los Angeles County Fair Chance Ordinance, the San Francisco Fair Chance Ordinance, and the California Fair Chance Act. Everforth CyberCoders is committed to working with and providing reasonable accommodation to individuals with physical and mental disabilities. Individuals needing special assistance or an accommodation while seeking employment can contact a member of our Human Resources team at Benefits@CyberCoders.com to make arrangements.\n\nSalary\n\n$185000 - $225000 USD per year","company":"Everforth Cybercoders","rawCompany":"everforth cybercoders","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-22T15:17:15.360Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1299.00","title":"Computer Occupations, All Other","slug":"computer-occupations-all-other"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Kernel Cluster Engineer","description":"AI Kernel Cluster Engineer\n\nPosition Overview\n\nI'm partnering with a rapidly growing AI infrastructure company supporting large-scale GPU environments that power AI training and inference workloads.\n\nThis is an opportunity to take ownership of the health, stability, and performance of production GPU clusters running mission-critical AI workloads. You'll work at the intersection of Linux systems, GPU infrastructure, containerization, and high-performance networking, helping ensure advanced AI environments operate reliably and efficiently at scale. Working closely with engineering and operations teams, you'll play a key role in maintaining the infrastructure that powers next-generation AI applications.\n\nKey Responsibilities\n\nOwn the day-to-day health, performance, resource allocation, and operations of production GPU clusters.\nManage Ubuntu Linux systems, including kernel-level tuning, driver management, patching, and OS troubleshooting.\nAdminister and troubleshoot LXC container environments supporting AI workloads.\nMonitor and maintain InfiniBand and RoCEv2 networking across GPU cluster environments.\nDiagnose and resolve complex issues spanning Linux systems, containers, GPUs, and cluster networking, including escalated production incidents.\n\nQualifications\n\nRequired\n\n3+ years of experience supporting GPU clusters, HPC environments, or large-scale infrastructure platforms.\nDeep experience with Ubuntu Linux, including kernel-level operations, system tuning, driver management, and OS troubleshooting.\nHands-on experience managing LXC containers in production environments.\nExperience supporting GPU cluster networking, including InfiniBand, RoCEv2, and other high-performance networking technologies.\nExperience troubleshooting issues across Linux systems, containers, networking, and GPU infrastructure.\nFamiliarity with NVIDIA technologies including DCGM, UFM, NCCL, CUDA, and SHARP, as well as monitoring platforms such as Prometheus and Grafana.\nExperience with SLURM, Kubernetes, or similar workload scheduling platforms.\nProficiency with Bash, Python, or similar scripting languages for automation and diagnostics.\n\nNice to Have\n\nFamiliarity with AI/ML frameworks such as PyTorch, TensorFlow, or JAX.\nExperience with Ansible, Terraform, or similar infrastructure automation platforms.\nFamiliarity with TensorRT, ONNX, or AI model optimization tooling.\nAI, cloud, Linux, or networking certifications.\nExperience supporting large-scale AI infrastructure or GPU cloud environments.\n\nBenefits\n\n$185,000 to $225,000 base salary\n30% to 50% annual performance bonus\nRestricted Stock Units (RSUs)\nThis opportunity is ideal for an engineer who enjoys working close to the operating system, networking, and hardware layers while supporting high-performance GPU environments running large-scale AI workloads. You'll have the opportunity to solve complex infrastructure challenges, work with cutting-edge AI compute environments, and play a key role in supporting the next generation of AI applications.\n\n- For this position, you must be currently authorized to work in the United States without the need for sponsorship for a non-immigrant visa. CyberCoders will consider for Employment in the City of Los Angeles qualified Applicants with Criminal Histories in a manner consistent with the requirements of the Los Angeles Fair Chance Initiative for Hiring (Ban the Box) Ordinance.This job was first posted by CyberCoders on 08/19/2026 and applications will be accepted on an ongoing basis until the position is filled or closed.Everforth CyberCoders is proud to be an Equal Opportunity Employer\n\nAll qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, sexual orientation, gender identity or expression, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, status as a crime victim, disability, protected veteran status, or any other characteristic protected by law. Our hiring process includes AI screening for keywords and minimum qualifications, and a virtual recruiter as part of the application process. A human recruiter reviews all results. Click here for details on our virtual recruiter . Everforth CyberCoders will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable state and local law, including but not limited to the Los Angeles County Fair Chance Ordinance, the San Francisco Fair Chance Ordinance, and the California Fair Chance Act. Everforth CyberCoders is committed to working with and providing reasonable accommodation to individuals with physical and mental disabilities. Individuals needing special assistance or an accommodation while seeking employment can contact a member of our Human Resources team at Benefits@CyberCoders.com to make arrangements.\n\nSalary\n\n$185000 - $225000 USD per year","datePosted":"2026-08-22T15:17:15.360Z","dateModified":"2026-08-22T15:17:15.360Z","hiringOrganization":{"@type":"Organization","name":"Everforth Cybercoders","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"15f79cfb9ff1fa79d3bd09f0"},"url":"https://jobsearcher.com/jobs/15f79cfb9ff1fa79d3bd09f0"}}