Machine Learning Engineer
Overview
Join NVIDIA to lead end-to-end AI system development and deployment. You will architect, deploy, and scale open-source models on distributed infrastructure, building robust pipelines and automated testing while ensuring secure CI/CD. You’ll operate GPU-driven workflows, manage orchestration with Kubernetes, Ray, or Slurm, and drive production-grade AI workloads through scalable platforms. This role combines AI application development with strong software engineering, giving you impact across model life cycles and system performance.
Compensation / Benefitsequitybenefitsbase salaryremote work optionscareer growthcompetitive compensation
ResponsibilitiesArchitect, deploy, and scale open-source models using distributed orchestration frameworks (Kubernetes, Ray, Slurm)Design and build ML systems and data pipelines; train, evaluate, and productionize AI agents; benchmark performanceRun model benchmarks and perform error/gap analysis; build analytics dashboards for stakeholdersOwn features from ideation to production; coordinate updates across repositories and communities
Key requirementsMaster’s or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience)3+ years of production-grade Python development with asynchronous design and clean architectureDeep experience with LangChain, Hugging Face libraries, vLLM, SGLang; TensorFlow, PyTorch, Scikit-learnProficient data analysis in Python (pandas, NumPy) and ability to translate results for varied audiencesHands-on model deployment, performance monitoring, and scaling with Kubernetes, Ray, or SlurmStrong understanding of GPU memory management and infrastructure tuning for high-throughput AI inferenceAdvanced GitLab CI/CD knowledge, automated tests, and vulnerability scanning in MR workflowsExperience with Python testing frameworks (PyTest), mocks, and AI-specific test generationAdvanced Git workflows, including rebasing, signing, and repo mirroringstrong communicationownership and accountabilitycollaboration with cross-functional teamsKubernetesRaySlurm