{"schemaVersion":"jobsearcher.job.v1","id":"128c3df6f9f0ced0fef9aaa3","url":"https://jobsearcher.com/jobs/128c3df6f9f0ced0fef9aaa3","canonicalUrl":"https://jobsearcher.com/jobs/128c3df6f9f0ced0fef9aaa3","title":"Manufacturing System Development Engineer, Cloud AI/ML/storage server teams","description":"DESCRIPTION\nApplication deadline: Aug 1, 2026\n\nAmazon Web Services (AWS) Hardware Engineering designs and delivers next-generation cloud infrastructure. Our team builds custom accelerator systems that power AI, machine learning, and compute workloads at global scale.\n\nWe are seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostic tooling, and hardware debug for GPU-based server platforms at our ODM/CM manufacturing sites. In this role, you will be on the manufacturing floor debugging complex system failures, developing automation to improve yield and throughput, and building diagnostic tools that enable root cause identification at the line. You will bridge the gap between hardware design intent and manufacturing execution — ensuring our platforms are testable, diagnosable, and launch with exceptional quality.\n\nThis role requires someone equally comfortable writing code and debugging hardware. You will develop test automation, build diagnostic frameworks, and personally troubleshoot failures spanning firmware, kernel, drivers, PCIe, power, and GPU subsystems — all in a fast-paced manufacturing environment. When something fails at the line, you are the person who figures out why.\n\nDomestic and international travel (~25%)\n\nKey job responsibilities\nManufacturing Debug & Root Cause Analysis\n\nDebug complex system-level failures at the manufacturing line across compute, storage, GPU, networking, power, and thermal domains\nPerform root cause analysis correlating across firmware, kernel, driver, PCIe, signal integrity, and physical layers to isolate faults\nTroubleshoot Linux boot and runtime failures across x86 and ARM architectures, including NVMe, GPU, NIC, and accelerator subsystems\nDrive Root Cause Corrective Action (RCCA) for yield detractors, test escapes, and recurring manufacturing failures\nProvide on-site ODM/CM support during critical builds, EVT/DVT/PVT phases, and production ramp\n\nTest Software & Automation Development\n\nDesign, develop, and maintain manufacturing test software and diagnostic tools deployed at ODM/CM lines\nBuild automation that reduces manual triage — enabling faster fault isolation and higher first-pass yield\nDevelop and optimize system-level test flows (BFT, functional test, stress test, burn-in) for GPU accelerator platforms\nBuild, manage, and deploy CI/CD pipelines for rapid deployment of test code to manufacturing environments\nWrite scalable, robust code in Python, C/C++, or Java to solve manufacturing test and debug challenges\n\nManufacturing Process & Quality\n\nDefine and improve manufacturing test strategy including coverage, duration, fixture requirements, and pass/fail criteria\nAnalyze test data and yield trends to identify systemic issues; drive design and process improvements\nCollaborate on DFx reviews (DFT/DFM) to ensure new designs are testable and diagnosable at the manufacturing line\nDevelop diagnostic tooling requirements for ODM/CM enablement — ensuring partners can effectively screen and debug at scale\nResearch and implement automation techniques to improve manufacturing efficiency and reduce human intervention\n\nCross-Team Collaboration\n\nWork across hardware design, firmware, qualification, and manufacturing engineering teams to close the loop between line failures and design improvements\nEngage with ODMs and design partners on testability, diagnostic, and automation requirements during NPI\nCollaborate with internal teams on GPU module integration, test coverage, and manufacturing debug procedures\nPartner with fleet health teams to ensure manufacturing diagnostics align with production monitoring and field failure analysis\nBASIC QUALIFICATIONS\n2+ years of non-internship professional software development experience\n1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience\nExperience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby\n5+ years of software development experience with at least one modern language (Python, C/C++, Java)\n3+ years of experience debugging hardware systems — server, accelerator, storage, or high-tech platforms\nExperience with Linux/Unix systems including boot flow, kernel, drivers, and OS-level diagnostics\nHands-on experience troubleshooting hardware failures at a manufacturing line or lab environment\nExperience working with ODMs/CMs through product development and manufacturing lifecycle\nWillingness to travel domestically and internationally (~25%), including extended on-site manufacturing support\nPREFERRED QUALIFICATIONS\nExperience with GPU-based server or accelerator platform manufacturing and debug\nFamiliarity with server hardware architecture: PCIe topology, NVMe, BMC/IPMI, power delivery, thermal\nExperience developing manufacturing test automation or diagnostic frameworks at scale\nExperience with board-level debug (oscilloscope, logic analyzer)\nKnowledge of firmware, BIOS, BMC, and their interaction with manufacturing test flows\nExperience with manufacturing yield analysis, test optimization, and throughput improvement\nExperience building CI/CD pipelines for test software deployment\nFamiliarity with telemetry, log correlation, and failure pattern analysis in manufacturing environments\n\nAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.\n\nLos Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.\n\nOur inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.\n\nThe base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.\n\nUSA, CA, Cupertino - 148,700.00 - 201,200.00 USD annually\nUSA, CO, Denver - 129,200.00 - 174,800.00 USD annually\nUSA, WA, Seattle - 129,200.00 - 174,800.00 USD annually","company":"Amazon Data Services","rawCompany":"amazon data services","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-08-03T19:35:17.149Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Manufacturing System Development Engineer, Cloud AI/ML/storage server teams","description":"DESCRIPTION\nApplication deadline: Aug 1, 2026\n\nAmazon Web Services (AWS) Hardware Engineering designs and delivers next-generation cloud infrastructure. Our team builds custom accelerator systems that power AI, machine learning, and compute workloads at global scale.\n\nWe are seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostic tooling, and hardware debug for GPU-based server platforms at our ODM/CM manufacturing sites. In this role, you will be on the manufacturing floor debugging complex system failures, developing automation to improve yield and throughput, and building diagnostic tools that enable root cause identification at the line. You will bridge the gap between hardware design intent and manufacturing execution — ensuring our platforms are testable, diagnosable, and launch with exceptional quality.\n\nThis role requires someone equally comfortable writing code and debugging hardware. You will develop test automation, build diagnostic frameworks, and personally troubleshoot failures spanning firmware, kernel, drivers, PCIe, power, and GPU subsystems — all in a fast-paced manufacturing environment. When something fails at the line, you are the person who figures out why.\n\nDomestic and international travel (~25%)\n\nKey job responsibilities\nManufacturing Debug & Root Cause Analysis\n\nDebug complex system-level failures at the manufacturing line across compute, storage, GPU, networking, power, and thermal domains\nPerform root cause analysis correlating across firmware, kernel, driver, PCIe, signal integrity, and physical layers to isolate faults\nTroubleshoot Linux boot and runtime failures across x86 and ARM architectures, including NVMe, GPU, NIC, and accelerator subsystems\nDrive Root Cause Corrective Action (RCCA) for yield detractors, test escapes, and recurring manufacturing failures\nProvide on-site ODM/CM support during critical builds, EVT/DVT/PVT phases, and production ramp\n\nTest Software & Automation Development\n\nDesign, develop, and maintain manufacturing test software and diagnostic tools deployed at ODM/CM lines\nBuild automation that reduces manual triage — enabling faster fault isolation and higher first-pass yield\nDevelop and optimize system-level test flows (BFT, functional test, stress test, burn-in) for GPU accelerator platforms\nBuild, manage, and deploy CI/CD pipelines for rapid deployment of test code to manufacturing environments\nWrite scalable, robust code in Python, C/C++, or Java to solve manufacturing test and debug challenges\n\nManufacturing Process & Quality\n\nDefine and improve manufacturing test strategy including coverage, duration, fixture requirements, and pass/fail criteria\nAnalyze test data and yield trends to identify systemic issues; drive design and process improvements\nCollaborate on DFx reviews (DFT/DFM) to ensure new designs are testable and diagnosable at the manufacturing line\nDevelop diagnostic tooling requirements for ODM/CM enablement — ensuring partners can effectively screen and debug at scale\nResearch and implement automation techniques to improve manufacturing efficiency and reduce human intervention\n\nCross-Team Collaboration\n\nWork across hardware design, firmware, qualification, and manufacturing engineering teams to close the loop between line failures and design improvements\nEngage with ODMs and design partners on testability, diagnostic, and automation requirements during NPI\nCollaborate with internal teams on GPU module integration, test coverage, and manufacturing debug procedures\nPartner with fleet health teams to ensure manufacturing diagnostics align with production monitoring and field failure analysis\nBASIC QUALIFICATIONS\n2+ years of non-internship professional software development experience\n1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience\nExperience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby\n5+ years of software development experience with at least one modern language (Python, C/C++, Java)\n3+ years of experience debugging hardware systems — server, accelerator, storage, or high-tech platforms\nExperience with Linux/Unix systems including boot flow, kernel, drivers, and OS-level diagnostics\nHands-on experience troubleshooting hardware failures at a manufacturing line or lab environment\nExperience working with ODMs/CMs through product development and manufacturing lifecycle\nWillingness to travel domestically and internationally (~25%), including extended on-site manufacturing support\nPREFERRED QUALIFICATIONS\nExperience with GPU-based server or accelerator platform manufacturing and debug\nFamiliarity with server hardware architecture: PCIe topology, NVMe, BMC/IPMI, power delivery, thermal\nExperience developing manufacturing test automation or diagnostic frameworks at scale\nExperience with board-level debug (oscilloscope, logic analyzer)\nKnowledge of firmware, BIOS, BMC, and their interaction with manufacturing test flows\nExperience with manufacturing yield analysis, test optimization, and throughput improvement\nExperience building CI/CD pipelines for test software deployment\nFamiliarity with telemetry, log correlation, and failure pattern analysis in manufacturing environments\n\nAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.\n\nLos Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.\n\nOur inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.\n\nThe base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.\n\nUSA, CA, Cupertino - 148,700.00 - 201,200.00 USD annually\nUSA, CO, Denver - 129,200.00 - 174,800.00 USD annually\nUSA, WA, Seattle - 129,200.00 - 174,800.00 USD annually","datePosted":"2026-08-03T19:35:17.149Z","dateModified":"2026-08-03T19:35:17.149Z","hiringOrganization":{"@type":"Organization","name":"Amazon Data Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"128c3df6f9f0ced0fef9aaa3"},"url":"https://jobsearcher.com/jobs/128c3df6f9f0ced0fef9aaa3"}}