HPC Platform Engineer, Software, Center for Quantum Computing
The Models, Quantum, & Silicon (MQS) Center for Quantum Computing (CQC) is a multi-disciplinary team of scientists, engineers, and technicians, on a mission to develop a fault‐tolerant quantum computer. We are looking to hire an HPC Platform Engineer to develop, automate, and maintain high‐performance computing (HPC) infrastructure on AWS that CQC scientists and engineers use for quantum computing hardware design and simulation. You will work closely with our experimental and theoretical physics teams to enable large‐scale HPC workloads with MPI‐based parallelism on EC2 instances, manage graphical environments for computer‐aided engineering applications, and accelerate the research computing lifecycle through automation and infrastructure‐as‐code. The ideal candidate will be able to translate high‐level science and simulation requirements into reliable deliverables (including cluster orchestration, job scheduling, CI/CD pipelines, reproducible environments, and artifact management) that are performant, scalable, and secure. Key job responsibilities Administer and automate cloud‐based HPC environments by deploying and maintaining clusters, managing OS and software stacks, building containers, provisioning users, and securing systems. Support computer‐aided engineering and computational science workflows (e.g., Palace) by improving HPC environment robustness, performance, and availability across instance types and regions. Anticipate and expand CQC computational capacity, leveraging AWS to accelerate the quantum hardware design cycle. Design reproducible development environments and own dependency, build, and release management across interdependent research software projects. Develop CI/CD pipelines and automate provisioning using infrastructure‐as‐code (e.g., Docker, AWS CDK). Maintain Amazon's high security bar while enabling fast‐paced development; implement observability and monitoring for rapid debugging and high uptime. We are looking for candidates with strong engineering principles, a bias for action, superior problem‐solving, and excellent communication skills. Working effectively within a team environment is essential. As an HPC Platform Engineer embedded in a research organization, you will have the opportunity to work on new ideas and stay abreast of the field of experimental quantum computation. A day in the life The majority of your time will be spent on projects that strengthen and scale our HPC platform and accelerate the computational workflows that underpin quantum hardware design. This requires working backwards from the needs of our science staff in the context of our larger experimental roadmap. You will translate science and software requirements into design proposals, balancing implementation complexity against time‐to‐delivery. Once a proposal has been reviewed and accepted, you'll drive implementation and coordinate with internal stakeholders to ensure a smooth rollout. About The Team You will be joining the Software team within the MQS Center of Quantum Computing. Our team is comprised of scientists and engineers who are building scalable software that enables quantum computing technologies. Basic Qualifications Experience in automating, deploying, and supporting large‐scale infrastructure. Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust. Experience with Linux/Unix. Experience with CI/CD pipelines and build processes. 2+ years of designing or architecting new and existing systems (design patterns, reliability and scaling). Experience using infrastructure‐as‐code to design and deploy cloud services. Preferred Qualifications Experience with distributed systems at scale. Experience in an AWS environment, including VPC, EC2, EBS, S3, SQS, CloudFormation, and Lambda. Experience with network fundamentals (DNS, DHCP, TCP/IP, routing, switching, HTTP). Experience in Kubernetes, Docker or containers ecosystem. Experience working with scientists in a research environment. Experience with high‐performance computing (HPC) infrastructure. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Company – Amazon Development Center U.S., Inc. Job ID: A10466636 #J-18808-Ljbffr