JOBSEARCHER

HPC Systems Engineer

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

Job ID 2610670Location Charlottesville, VA, USDate Posted 2026-03-26Category Engineering and SciencesSubcategory Systems EngineerSchedule Full-TimeShift Day JobTravel NoMinimum Clearance Required Top_SecretClearance Level Must Be Able to Obtain TS/SCIPotential for Remote Work ORA_ON_SITEDescriptionSAIC is looking for a highly qualified HPC Systems Engineer to support the Army’s Golden Dome initiative. The engineer will support the deployment and sustainment of Linux-based High Performance Computing (HPC) cluster environments used for distributed compute workloads, simulation environments, and GPU-enabled processing.The environment will includemulti-node Linux compute clustersworkload scheduling platforms such as Slurm or PBScluster provisioning frameworks (e.g., xCAT, Warewulf)high-performance networking technologies including RDMA / InfiniBanddistributed parallel compute workloads utilizing MPI or OpenMPGPU-enabled compute resources supporting CUDA-based processingThe system will be used to support scientific computing, simulation workloads, and other distributed compute operations within a secure research environment.Candidates should be comfortable working within cluster-scale computing environments where performance, scheduler configuration, and distributed workload execution are critical operational factors.The HPC Systems Engineer will support the build-out, configuration, and sustainment of HPC cluster platforms.The role focuses oncluster platform configurationscheduler administrationdistributed compute troubleshootingperformance analysis across compute, storage, and network layersGPU compute workload supportautomation and operational toolingCandidates should have experience working with multi-node Linux cluster environments and distributed compute workloads.Core Technical CapabilitiesCandidates should demonstrate capability in most of the following areas.HPC Cluster PlatformsExperience supporting multi-node Linux compute clusters, including node integration, configuration, and operational sustainment.Experience with cluster provisioning tools such as xCAT, Warewulf, or similar node deployment systems is beneficial.Workload Scheduling PlatformsExperience supporting distributed compute workloads using schedulers such asSlurmPBS / PBS ProTorqueGrid EngineCandidates should understand queue configuration, job submission workflows, and scheduler troubleshooting.Candidates should understand how workload schedulers interact with distributed compute workloads and containerized execution environments.Linux Systems AdministrationStrong Linux administration experience includingcommand-line system administrationserver and compute node configurationsystem troubleshooting in distributed compute environmentsExperience With RHEL-based Environments Is Preferred.Distributed and Containerized WorkloadsExperience supporting distributed compute workloads utilizing parallel computing frameworks such asMPIOpenMPGPU compute frameworksCandidates should understand how workload schedulers interact with distributed compute workloads and containerized execution environments within HPC clusters.Familiarity with container technologies commonly used in HPC environments such asDockerPodmanSingularity / ApptainerCandidates should understand how containerized workloads interact with schedulers, GPU resources, and distributed compute environments.Experience supporting containerized HPC workloads or integrating container platforms with cluster infrastructure is desirable.HPC NetworkingFamiliarity with high-performance networking technologies includingRDMA networkingInfiniBandhigh-throughput cluster networking architecturesCandidates should be comfortable assisting with troubleshooting cluster communication or performance issues.GPU Compute EnvironmentsExperience supporting GPU-enabled compute environments and workloads utilizing CUDA frameworks is desirable.Automation and Operational ToolingExperience writing scripts or operational tooling using languages such asBashPython Automation experience supporting system administration or cluster operations is beneficial.QualificationsCandidates must meet the following requirements Bachelor degree in science/technology; 10 additional YoE can be substituted for degree8+ years of experience is requiredMinimum 6 years of experience administering Linux systems in enterprise, research computing, or distributed compute environmentsAn Active Top Secret clearance is required; an active TS/SCI clearance must be obtained prior to beginning work.100% onsite support in Charlottesville, VAExperience supporting distributed compute environments or HPC cluster platformsExperience working with workload schedulers such as Slurm, PBS, Torque, or similar systemsExperience administering Linux systems through command-line interfacesExperience with scripting or automation tools (Bash, Python, or similar)Ability to obtain required DoD 8140 (8570) IAT Level II certificationCandidates must have direct experience with HPC or distributed compute environments. Candidates With The Following Experience Are Strongly PreferredAdministration of multi-node HPC cluster environmentsExperience with parallel or distributed file systems such as Lustre, BeeGFS, or GPFSExperience supporting GPU-enabled compute environments and CUDA workloadsExperience with configuration management tools such as Ansible or PuppetExperience supporting research, laboratory, or mission computing environmentsExperience supporting systems within DoD/DoW or IC environments