{"schemaVersion":"jobsearcher.job.v1","id":"61c1e4f48a59b7eb13112990","url":"https://jobsearcher.com/jobs/61c1e4f48a59b7eb13112990","canonicalUrl":"https://jobsearcher.com/jobs/61c1e4f48a59b7eb13112990","title":"Software Engineer, Platform Reliability Engineering, AiDP","description":"AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning - including generative AI - along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.\n\nThe Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apple's enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs, builds, and operates mission-critical platforms and services spanning ML, GenAI, inference, and big data-enabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability, performance, and scalability across a diverse ecosystem of open source and cutting-edge technologies, serving some of Apple's most demanding workloads.\n\nDescription\n\nWe're seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You'll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you'll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure.\n\n\",\"responsibilities\":\"Design, build, and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scale\n\nOwn the full lifecycle of infrastructure and platform projects-from architectural design and implementation through deployment, monitoring, and optimization\n\nOperate and optimize high-throughput, mission-critical services to ensure reliability, performance, and cost-efficiency\n\nParticipate in on-call rotations to respond to production incidents; diagnose root causes, implement rapid fixes, and drive post-incident improvements\n\nLead cross-functional collaboration with engineering teams to define requirements, validate designs, and deliver customer-impacting features and improvements\n\nProactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resilience\n\nEstablish observability practices and continuously refine operational excellence standards across the platform\n\nPreferred Qualifications\n\n7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale.\n\nProficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments.\n\nFamiliarity with open source codebases; ability to read, understand, and explain complex system implementations\n\nStrong understanding of system architecture and proven ability to collaborate effectively across engineering teams\n\nHands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure).\n\nStrong foundational knowledge of Linux, databases, and security principles\n\nProactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services\n\nExcellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership\n\nDemonstrated track record of designing and operating systems at scale\n\nMinimum Qualifications\n\nBachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience\n\nProficiency in at least one systems programming language (Python, Go, Java, or similar)\n\nStrong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles\n\nHands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies)\n\nPay & Benefits\n\nAt Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.\n\nApple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits\n\nNote: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.","company":"Apple","rawCompany":"apple","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-04T10:33:05.046Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer, Platform Reliability Engineering, AiDP","description":"AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning - including generative AI - along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.\n\nThe Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apple's enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs, builds, and operates mission-critical platforms and services spanning ML, GenAI, inference, and big data-enabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability, performance, and scalability across a diverse ecosystem of open source and cutting-edge technologies, serving some of Apple's most demanding workloads.\n\nDescription\n\nWe're seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You'll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you'll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure.\n\n\",\"responsibilities\":\"Design, build, and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scale\n\nOwn the full lifecycle of infrastructure and platform projects-from architectural design and implementation through deployment, monitoring, and optimization\n\nOperate and optimize high-throughput, mission-critical services to ensure reliability, performance, and cost-efficiency\n\nParticipate in on-call rotations to respond to production incidents; diagnose root causes, implement rapid fixes, and drive post-incident improvements\n\nLead cross-functional collaboration with engineering teams to define requirements, validate designs, and deliver customer-impacting features and improvements\n\nProactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resilience\n\nEstablish observability practices and continuously refine operational excellence standards across the platform\n\nPreferred Qualifications\n\n7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale.\n\nProficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments.\n\nFamiliarity with open source codebases; ability to read, understand, and explain complex system implementations\n\nStrong understanding of system architecture and proven ability to collaborate effectively across engineering teams\n\nHands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure).\n\nStrong foundational knowledge of Linux, databases, and security principles\n\nProactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services\n\nExcellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership\n\nDemonstrated track record of designing and operating systems at scale\n\nMinimum Qualifications\n\nBachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience\n\nProficiency in at least one systems programming language (Python, Go, Java, or similar)\n\nStrong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles\n\nHands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies)\n\nPay & Benefits\n\nAt Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.\n\nApple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits\n\nNote: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.","datePosted":"2026-09-04T10:33:05.046Z","dateModified":"2026-09-04T10:33:05.046Z","hiringOrganization":{"@type":"Organization","name":"Apple","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"61c1e4f48a59b7eb13112990"},"url":"https://jobsearcher.com/jobs/61c1e4f48a59b7eb13112990"}}