AI Kernel Cluster Engineer
AI Kernel Cluster Engineer
Position Overview
I'm partnering with a rapidly growing AI infrastructure company supporting large-scale GPU environments that power AI training and inference workloads.
This is an opportunity to take ownership of the health, stability, and performance of production GPU clusters running mission-critical AI workloads. You'll work at the intersection of Linux systems, GPU infrastructure, containerization, and high-performance networking, helping ensure advanced AI environments operate reliably and efficiently at scale. Working closely with engineering and operations teams, you'll play a key role in maintaining the infrastructure that powers next-generation AI applications.
Key Responsibilities
Own the day-to-day health, performance, resource allocation, and operations of production GPU clusters.
Manage Ubuntu Linux systems, including kernel-level tuning, driver management, patching, and OS troubleshooting.
Administer and troubleshoot LXC container environments supporting AI workloads.
Monitor and maintain InfiniBand and RoCEv2 networking across GPU cluster environments.
Diagnose and resolve complex issues spanning Linux systems, containers, GPUs, and cluster networking, including escalated production incidents.
Qualifications
Required
3+ years of experience supporting GPU clusters, HPC environments, or large-scale infrastructure platforms.
Deep experience with Ubuntu Linux, including kernel-level operations, system tuning, driver management, and OS troubleshooting.
Hands-on experience managing LXC containers in production environments.
Experience supporting GPU cluster networking, including InfiniBand, RoCEv2, and other high-performance networking technologies.
Experience troubleshooting issues across Linux systems, containers, networking, and GPU infrastructure.
Familiarity with NVIDIA technologies including DCGM, UFM, NCCL, CUDA, and SHARP, as well as monitoring platforms such as Prometheus and Grafana.
Experience with SLURM, Kubernetes, or similar workload scheduling platforms.
Proficiency with Bash, Python, or similar scripting languages for automation and diagnostics.
Nice to Have
Familiarity with AI/ML frameworks such as PyTorch, TensorFlow, or JAX.
Experience with Ansible, Terraform, or similar infrastructure automation platforms.
Familiarity with TensorRT, ONNX, or AI model optimization tooling.
AI, cloud, Linux, or networking certifications.
Experience supporting large-scale AI infrastructure or GPU cloud environments.
Benefits
$185,000 to $225,000 base salary
30% to 50% annual performance bonus
Restricted Stock Units (RSUs)
This opportunity is ideal for an engineer who enjoys working close to the operating system, networking, and hardware layers while supporting high-performance GPU environments running large-scale AI workloads. You'll have the opportunity to solve complex infrastructure challenges, work with cutting-edge AI compute environments, and play a key role in supporting the next generation of AI applications.
- For this position, you must be currently authorized to work in the United States without the need for sponsorship for a non-immigrant visa. CyberCoders will consider for Employment in the City of Los Angeles qualified Applicants with Criminal Histories in a manner consistent with the requirements of the Los Angeles Fair Chance Initiative for Hiring (Ban the Box) Ordinance.This job was first posted by CyberCoders on 08/19/2026 and applications will be accepted on an ongoing basis until the position is filled or closed.Everforth CyberCoders is proud to be an Equal Opportunity Employer
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, sexual orientation, gender identity or expression, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, status as a crime victim, disability, protected veteran status, or any other characteristic protected by law. Our hiring process includes AI screening for keywords and minimum qualifications, and a virtual recruiter as part of the application process. A human recruiter reviews all results. Click here for details on our virtual recruiter . Everforth CyberCoders will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable state and local law, including but not limited to the Los Angeles County Fair Chance Ordinance, the San Francisco Fair Chance Ordinance, and the California Fair Chance Act. Everforth CyberCoders is committed to working with and providing reasonable accommodation to individuals with physical and mental disabilities. Individuals needing special assistance or an accommodation while seeking employment can contact a member of our Human Resources team at Benefits@CyberCoders.com to make arrangements.
Salary
$185000 - $225000 USD per year