{"schemaVersion":"jobsearcher.job.v1","id":"4486b3bc2d459e4cdb243468","url":"https://jobsearcher.com/jobs/4486b3bc2d459e4cdb243468","canonicalUrl":"https://jobsearcher.com/jobs/4486b3bc2d459e4cdb243468","title":"Machine Learning Operations Engineer - Remote","description":"NAVA Software solutions is looking for a Machine Learning Operations EngineerDetails:Machine Learning Operations (MLOps) Engineer - AWS (with LLM Focus)Location: Remote workDuration: 12 monthsResponsibilities:LLM-Optimized MLOps Infrastructure: Design and implement MLOps infrastructure on AWS tailored for LLMs, leveraging services like SageMaker, EC2 (with GPU instances), S3, ECS/EKS, Lambda, and more.LLM Deployment Pipelines: Build and manage CI/CD pipelines specifically for LLM deployment, addressing unique challenges like model size, inference optimization, and versioning.LLMOps Practices: Implement LLMOps best practices for monitoring model performance, drift detection, prompt management, and feedback loops for continuous improvement.RESTful API Development: Design and develop RESTful APIs to expose LLM capabilities to other applications and services, ensuring scalability, security, and optimal performance.Model Optimization: Apply techniques like quantization, distillation, and pruning to optimize LLM models for efficient inference on AWS infrastructure.Monitoring and Observability: Establish comprehensive monitoring and alerting mechanisms to track LLM performance, latency, resource utilization, and potential biases.Prompt Engineering and Management: Develop strategies for prompt engineering and management to enhance LLM outputs and ensure consistency and safety.Collaboration: Work closely with data scientists, researchers, and software engineers to integrate LLM models into production systems effectively.Cost Optimization: Continuously optimize LLMOps processes and infrastructure for cost-efficiency while maintaining high performance and reliability.Qualifications: Experience: 3+ years of experience in MLOps or a related field, with hands-on experience in deploying and managing LLMs.AWS Expertise: Strong proficiency in AWS services relevant to MLOps and LLMs, including SageMaker, EC2 (with GPU instances), S3, ECS/EKS, Lambda, and API Gateway.LLM Knowledge: Deep understanding of LLM architectures (e.g., Transformers), training techniques, and inference optimization strategies.Programming Skills: Proficiency in Python and experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation), REST API frameworks (e.g., Flask, FastAPI), and LLM libraries (e.g., Hugging Face Transformers).Monitoring: Familiarity with monitoring and logging tools for LLMs, such as Prometheus, Grafana, and CloudWatch.Containerization: Experience with Docker and container orchestration (e.g., Kubernetes, ECS) for LLM deployment.Problem Solving: Excellent problem-solving and troubleshooting skills in the context of LLMs and MLOps.Communication: Strong communication and collaboration skills to effectively work with cross-functional teams","company":"Nava Software Solutions","rawCompany":"nava software solutions","city":"New York","state":"NY","isRemote":true,"isActive":false,"createdAt":"2026-06-03T00:04:35.030Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Operations Engineer - Remote","description":"NAVA Software solutions is looking for a Machine Learning Operations EngineerDetails:Machine Learning Operations (MLOps) Engineer - AWS (with LLM Focus)Location: Remote workDuration: 12 monthsResponsibilities:LLM-Optimized MLOps Infrastructure: Design and implement MLOps infrastructure on AWS tailored for LLMs, leveraging services like SageMaker, EC2 (with GPU instances), S3, ECS/EKS, Lambda, and more.LLM Deployment Pipelines: Build and manage CI/CD pipelines specifically for LLM deployment, addressing unique challenges like model size, inference optimization, and versioning.LLMOps Practices: Implement LLMOps best practices for monitoring model performance, drift detection, prompt management, and feedback loops for continuous improvement.RESTful API Development: Design and develop RESTful APIs to expose LLM capabilities to other applications and services, ensuring scalability, security, and optimal performance.Model Optimization: Apply techniques like quantization, distillation, and pruning to optimize LLM models for efficient inference on AWS infrastructure.Monitoring and Observability: Establish comprehensive monitoring and alerting mechanisms to track LLM performance, latency, resource utilization, and potential biases.Prompt Engineering and Management: Develop strategies for prompt engineering and management to enhance LLM outputs and ensure consistency and safety.Collaboration: Work closely with data scientists, researchers, and software engineers to integrate LLM models into production systems effectively.Cost Optimization: Continuously optimize LLMOps processes and infrastructure for cost-efficiency while maintaining high performance and reliability.Qualifications: Experience: 3+ years of experience in MLOps or a related field, with hands-on experience in deploying and managing LLMs.AWS Expertise: Strong proficiency in AWS services relevant to MLOps and LLMs, including SageMaker, EC2 (with GPU instances), S3, ECS/EKS, Lambda, and API Gateway.LLM Knowledge: Deep understanding of LLM architectures (e.g., Transformers), training techniques, and inference optimization strategies.Programming Skills: Proficiency in Python and experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation), REST API frameworks (e.g., Flask, FastAPI), and LLM libraries (e.g., Hugging Face Transformers).Monitoring: Familiarity with monitoring and logging tools for LLMs, such as Prometheus, Grafana, and CloudWatch.Containerization: Experience with Docker and container orchestration (e.g., Kubernetes, ECS) for LLM deployment.Problem Solving: Excellent problem-solving and troubleshooting skills in the context of LLMs and MLOps.Communication: Strong communication and collaboration skills to effectively work with cross-functional teams","datePosted":"2026-06-03T00:04:35.030Z","dateModified":"2026-06-03T00:04:35.030Z","hiringOrganization":{"@type":"Organization","name":"Nava Software Solutions","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"New York","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"4486b3bc2d459e4cdb243468"},"url":"https://jobsearcher.com/jobs/4486b3bc2d459e4cdb243468"}}