{"schemaVersion":"jobsearcher.job.v1","id":"6bbc2e3225fd898d5bf6d3e3","url":"https://jobsearcher.com/jobs/6bbc2e3225fd898d5bf6d3e3","canonicalUrl":"https://jobsearcher.com/jobs/6bbc2e3225fd898d5bf6d3e3","title":"Machine Learning Engineer","description":"Overview\nJoin NVIDIA to lead end-to-end AI system development and deployment. You will architect, deploy, and scale open-source models on distributed infrastructure, building robust pipelines and automated testing while ensuring secure CI/CD. You’ll operate GPU-driven workflows, manage orchestration with Kubernetes, Ray, or Slurm, and drive production-grade AI workloads through scalable platforms. This role combines AI application development with strong software engineering, giving you impact across model life cycles and system performance.\n\nCompensation / Benefitsequitybenefitsbase salaryremote work optionscareer growthcompetitive compensation\nResponsibilitiesArchitect, deploy, and scale open-source models using distributed orchestration frameworks (Kubernetes, Ray, Slurm)Design and build ML systems and data pipelines; train, evaluate, and productionize AI agents; benchmark performanceRun model benchmarks and perform error/gap analysis; build analytics dashboards for stakeholdersOwn features from ideation to production; coordinate updates across repositories and communities\nKey requirementsMaster’s or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience)3+ years of production-grade Python development with asynchronous design and clean architectureDeep experience with LangChain, Hugging Face libraries, vLLM, SGLang; TensorFlow, PyTorch, Scikit-learnProficient data analysis in Python (pandas, NumPy) and ability to translate results for varied audiencesHands-on model deployment, performance monitoring, and scaling with Kubernetes, Ray, or SlurmStrong understanding of GPU memory management and infrastructure tuning for high-throughput AI inferenceAdvanced GitLab CI/CD knowledge, automated tests, and vulnerability scanning in MR workflowsExperience with Python testing frameworks (PyTest), mocks, and AI-specific test generationAdvanced Git workflows, including rebasing, signing, and repo mirroringstrong communicationownership and accountabilitycollaboration with cross-functional teamsKubernetesRaySlurm","company":"NVIDIA","rawCompany":"nvidia","city":"Norfolk","state":"VA","isRemote":false,"isActive":true,"createdAt":"2026-09-15T03:50:48.748Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer","description":"Overview\nJoin NVIDIA to lead end-to-end AI system development and deployment. You will architect, deploy, and scale open-source models on distributed infrastructure, building robust pipelines and automated testing while ensuring secure CI/CD. You’ll operate GPU-driven workflows, manage orchestration with Kubernetes, Ray, or Slurm, and drive production-grade AI workloads through scalable platforms. This role combines AI application development with strong software engineering, giving you impact across model life cycles and system performance.\n\nCompensation / Benefitsequitybenefitsbase salaryremote work optionscareer growthcompetitive compensation\nResponsibilitiesArchitect, deploy, and scale open-source models using distributed orchestration frameworks (Kubernetes, Ray, Slurm)Design and build ML systems and data pipelines; train, evaluate, and productionize AI agents; benchmark performanceRun model benchmarks and perform error/gap analysis; build analytics dashboards for stakeholdersOwn features from ideation to production; coordinate updates across repositories and communities\nKey requirementsMaster’s or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience)3+ years of production-grade Python development with asynchronous design and clean architectureDeep experience with LangChain, Hugging Face libraries, vLLM, SGLang; TensorFlow, PyTorch, Scikit-learnProficient data analysis in Python (pandas, NumPy) and ability to translate results for varied audiencesHands-on model deployment, performance monitoring, and scaling with Kubernetes, Ray, or SlurmStrong understanding of GPU memory management and infrastructure tuning for high-throughput AI inferenceAdvanced GitLab CI/CD knowledge, automated tests, and vulnerability scanning in MR workflowsExperience with Python testing frameworks (PyTest), mocks, and AI-specific test generationAdvanced Git workflows, including rebasing, signing, and repo mirroringstrong communicationownership and accountabilitycollaboration with cross-functional teamsKubernetesRaySlurm","datePosted":"2026-09-15T03:50:48.748Z","dateModified":"2026-09-15T03:50:48.748Z","hiringOrganization":{"@type":"Organization","name":"NVIDIA","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Norfolk","addressRegion":"VA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6bbc2e3225fd898d5bf6d3e3"},"url":"https://jobsearcher.com/jobs/6bbc2e3225fd898d5bf6d3e3"}}