Machine Learning Infrastructure Engineer, GenAI Technology
Overview
In this role you design and operate high-performance GenAI infrastructure to accelerate model development and deployment. You collaborate with ML researchers to optimize training and inference, automate CI/CD and IaC pipelines, and monitor costs and security across cloud and on‑prem environments. You will help scale systems, improve reliability, and mentor teammates as part of Point72’s AI‑driven tech initiatives. This is a chance to shape production-grade GenAI platforms supporting a multi‑billion-dollar business.
Compensation / BenefitsFully-paid health care benefitsGenerous parental and family leaveVolunteer opportunitiesEmployee-led affinity groupsMental and physical wellness programsTuition assistance
ResponsibilitiesDesign and implement high-performance infrastructure for large-scale generative AI/ML workloadsDesign and operate distributed systems for training, tuning, inference, and data preprocessing pipelinesCollaborate with ML researchers to optimize compute, training throughput, and latencyDevelop and automate deployment, orchestration, and CI/CD pipelines using container orchestration and IaCImplement observability, monitoring, and cost-management in GPU/accelerator environmentsEvaluate, integrate, and benchmark new tech across cloud/on-prem for scalabilityDrive security, compliance, and runbooks for GenAI infrastructure including access controls and incident responseTroubleshoot and optimize performance across GPU/CPU stacksDocument architecture and operational practices and mentor engineers to accelerate production readiness
Key requirementsBachelor's or Master's in computer science, electrical engineering, or related technical field3–7 years of experience building and maintaining scalable compute or ML infra systemsDeep understanding of distributed systems, Kubernetes, and public cloud platforms (AWS, GCP, Azure)Hands-on experience with MLflow, Ray, Airflow, Kubeflow, and TerraformStrong understanding of reinforcement learning concepts and their infra implicationsProficiency in Python and systems-level programming (Go, C++, or Rust)Strong debugging, performance profiling, and optimization skills for GPU/CPU stacksExperience implementing monitoring, observability, and cost-optimization for GPU/accelerator environmentsExcellent collaboration and communication skills with a systems-thinking mindsetCommitment to the highest ethical standardscollaborationcommunicationsystems-thinkingKubernetesAWS/GCP/AzureMLflow