{"schemaVersion":"jobsearcher.job.v1","id":"cbba86c30cc0452e13a35e5f","url":"https://jobsearcher.com/jobs/cbba86c30cc0452e13a35e5f","canonicalUrl":"https://jobsearcher.com/jobs/cbba86c30cc0452e13a35e5f","title":"Machine Learning Engineer - Model Inference","description":"Join Apple Maps to help build the best map in the world! In this role on our ML Platform Team, you will leverage advanced deep learning and large language models to improve the search quality and overall customer experiences across our various Maps platforms. This role offers amazing opportunities to partner closely with research and product teams while taking ownership of projects and delivering measurable results at a global scale!\n\nDescription\n\nAs a member of our team, you will help design, build, and operate the services used to deploy and serve machine learning models at scale. You will help oversee the infrastructure that powers model inference, from developing high-performance serving systems to implementing optimization techniques that reduce latency, increase throughput, and improve hardware utilization. Get excited about collaborating closely with machine learning researchers, infrastructure engineers, and product teams to transform new models into reliable, production-ready experiences.\n\nThis role will require you to communicate technical ideas clearly, and to present design decisions and performance findings to both technical and cross-functional audiences. You will also participate in collaborative discussions, design reviews, and project planning meetings.\n\n.\n\nWe encourage our team-members to learn quickly, take ownership of meaningful projects, and contribute ideas that improve both the performance of our systems and the experiences of the people who use them!\n\n\",\"responsibilities\":\"Design, implement, test, and maintain scalable machine learning inference services.\n\nImprove inference latency, throughput, availability, and infrastructure efficiency.\n\nDevelop benchmarking and profiling tools to identify performance bottlenecks.\n\nDevelop techniques such as dynamic batching, caching, quantization, pruning, model compilation, and parallel execution.\n\nWork with machine learning frameworks, inference runtimes, GPUs, and other hardware accelerators.\n\nBuild monitoring, logging, alerting, and load-testing capabilities for production services.\n\nInvestigate reliability and performance issues across models, software runtimes, and infrastructure.\n\nWrite clear, maintainable code and participate in design reviews, code reviews, and operational support.\n\nPreferred Qualifications\n\nExperience with model serving technologies such as Triton, TensorRT, ONNX Runtime, vLLM, TensorFlow Serving, or TorchServe.\n\nFamiliarity with inference optimization techniques, including quantization, pruning, knowledge distillation, speculative decoding, kernel fusion, or continuous batching.\n\nUnderstanding of GPUs, accelerators, distributed systems, networking, or high-performance computing.\n\nFamiliarity with containers, Kubernetes, cloud infrastructure, and production observability tools.\n\nExperience benchmarking large language models, vision models, or other compute-intensive machine learning workloads.\n\nPossess curiosity about how software, models, and hardware interact to determine real-world performance.\n\nMinimum Qualifications\n\nBachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field plus at least 2 years of post graduate work experience.\n\nStrong programming skills in Python and at least one systems-oriented language such as C++, Rust, or Go.\n\nSolid understanding of data structures, algorithms, operating systems, and computer architecture.\n\nFamiliarity with machine learning fundamentals and modern deep learning frameworks such as PyTorch, TensorFlow, or JAX.\n\nExperience building, debugging, or evaluating software systems through coursework, internships, research, open-source contributions, or personal projects.\n\nAbility to analyze technical problems, communicate clearly, and work effectively with engineers across multiple disciplines.\n\nPay & Benefits\n\nAt Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $150,400 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.\n\nApple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits\n\nNote: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.","company":"Apple","rawCompany":"apple","city":"Cupertino","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-31T12:02:12.073Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer - Model Inference","description":"Join Apple Maps to help build the best map in the world! In this role on our ML Platform Team, you will leverage advanced deep learning and large language models to improve the search quality and overall customer experiences across our various Maps platforms. This role offers amazing opportunities to partner closely with research and product teams while taking ownership of projects and delivering measurable results at a global scale!\n\nDescription\n\nAs a member of our team, you will help design, build, and operate the services used to deploy and serve machine learning models at scale. You will help oversee the infrastructure that powers model inference, from developing high-performance serving systems to implementing optimization techniques that reduce latency, increase throughput, and improve hardware utilization. Get excited about collaborating closely with machine learning researchers, infrastructure engineers, and product teams to transform new models into reliable, production-ready experiences.\n\nThis role will require you to communicate technical ideas clearly, and to present design decisions and performance findings to both technical and cross-functional audiences. You will also participate in collaborative discussions, design reviews, and project planning meetings.\n\n.\n\nWe encourage our team-members to learn quickly, take ownership of meaningful projects, and contribute ideas that improve both the performance of our systems and the experiences of the people who use them!\n\n\",\"responsibilities\":\"Design, implement, test, and maintain scalable machine learning inference services.\n\nImprove inference latency, throughput, availability, and infrastructure efficiency.\n\nDevelop benchmarking and profiling tools to identify performance bottlenecks.\n\nDevelop techniques such as dynamic batching, caching, quantization, pruning, model compilation, and parallel execution.\n\nWork with machine learning frameworks, inference runtimes, GPUs, and other hardware accelerators.\n\nBuild monitoring, logging, alerting, and load-testing capabilities for production services.\n\nInvestigate reliability and performance issues across models, software runtimes, and infrastructure.\n\nWrite clear, maintainable code and participate in design reviews, code reviews, and operational support.\n\nPreferred Qualifications\n\nExperience with model serving technologies such as Triton, TensorRT, ONNX Runtime, vLLM, TensorFlow Serving, or TorchServe.\n\nFamiliarity with inference optimization techniques, including quantization, pruning, knowledge distillation, speculative decoding, kernel fusion, or continuous batching.\n\nUnderstanding of GPUs, accelerators, distributed systems, networking, or high-performance computing.\n\nFamiliarity with containers, Kubernetes, cloud infrastructure, and production observability tools.\n\nExperience benchmarking large language models, vision models, or other compute-intensive machine learning workloads.\n\nPossess curiosity about how software, models, and hardware interact to determine real-world performance.\n\nMinimum Qualifications\n\nBachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field plus at least 2 years of post graduate work experience.\n\nStrong programming skills in Python and at least one systems-oriented language such as C++, Rust, or Go.\n\nSolid understanding of data structures, algorithms, operating systems, and computer architecture.\n\nFamiliarity with machine learning fundamentals and modern deep learning frameworks such as PyTorch, TensorFlow, or JAX.\n\nExperience building, debugging, or evaluating software systems through coursework, internships, research, open-source contributions, or personal projects.\n\nAbility to analyze technical problems, communicate clearly, and work effectively with engineers across multiple disciplines.\n\nPay & Benefits\n\nAt Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $150,400 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.\n\nApple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits\n\nNote: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.","datePosted":"2026-07-31T12:02:12.073Z","dateModified":"2026-07-31T12:02:12.073Z","hiringOrganization":{"@type":"Organization","name":"Apple","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Cupertino","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"cbba86c30cc0452e13a35e5f"},"url":"https://jobsearcher.com/jobs/cbba86c30cc0452e13a35e5f"}}