JOBSEARCHER

HPC Operations Engineer

NVIDIASomerville, MAL5 SeniorSeptember 15th, 2026
Overview In this role you enable reliable HPC operations for NVIDIA’s compute ecosystem, supporting researchers and engineers who run demanding workloads. You work with cross-functional teams to keep scheduling, storage, and compute resources healthy and responsive. You’ll triage incidents, improve operational processes, and contribute to documentation and runbooks that empower users and internal teams. This is a hands-on role with impact on system stability, performance, and user experience in a leading AI and semiconductor environment. You join a collaborative, innovative team that values practical improvements and scalable workflows. Compensation / Benefitsequitybenefits package ResponsibilitiesProvide first-line support for HPC users across scheduling, compute, storage, and access issuesTroubleshoot job failures, scheduler errors, resource constraints, and performance concerns; drive resolution or escalationTriage infrastructure incidents; gather diagnostics and involve SMEs when neededMonitor system health, queues, node status, and service availability for stable operationsComplete maintenance, patching, and configuration updates per proceduresDevelop and maintain operational documentation, runbooks, and user/internal guidelinesIdentify recurring issues and propose workflow refinements to improve team processes Key requirementsBachelor’s degree in Computer Science, Information Technology, Engineering, or related field, or equivalent experience2+ years supporting Linux-based production environmentsSolid Linux administration fundamentals (RHEL/CentOS and/or Ubuntu)Methodical trouble-shooting and escalation judgmentExperience interacting with users in a technical support or operations roleStrong written communication for documentation and proceduresAbility to follow established processes with attention to detailstrong written communicationattention to detailability to work with users and cross-functional teamsBash or Python scriptingExperience with workload schedulers (LSF, Slurm)Understanding of NFS, automounter, LDAP