{"schemaVersion":"jobsearcher.job.v1","id":"828b1fb3c78f8767c860f450","url":"https://jobsearcher.com/jobs/828b1fb3c78f8767c860f450","canonicalUrl":"https://jobsearcher.com/jobs/828b1fb3c78f8767c860f450","title":"Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams","description":"DESCRIPTION\nAs a Cloud Hardware Development Engineer, you will be an end-to-end owner of storage and/or accelerator (AI/ML/GPU) server platforms — from New Product Introduction (NPI) through fleet health in production. You own the full lifecycle: design, development, qualification, launch, and ongoing operational excellence of servers running at scale in the AWS fleet.\n\nYou will work closely with internal customers to understand their technical needs and business goals, leveraging your experience with server design and the knowledge of various teams to architect solutions we deploy at scale. To deliver your products, you will work with an interdisciplinary team of component, firmware, power, mechanical, electrical, test, qualification, manufacturing engineers, and lead our ODM (design and manufacturing partners) to bring these servers to the data center. After launch, you own the fleet — monitoring quality, driving reliability improvements, and ensuring servers continue to meet customer requirements throughout their\noperational life.\n\nThis role demands deep technical curiosity and the willingness to jump in and personally solve the hardest problems. When a complex system failure occurs — whether during NPI qualification or in a production fleet of hundreds of thousands of servers — you roll up your sleeves, dive into the details across hardware, firmware, software, and physical layers, and drive to root cause. You don't wait for someone else to figure it out.\n\nYou will own end-to-end system reliability — proactively identifying deficiencies and driving toward zero-touch operations where automation detects, diagnoses, and resolves issues before customer impact. You will decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features, leading delivery yourself and through others in parallel.\n\nThis is a fast-paced, intellectually challenging position. You'll work with thought leaders in multiple technology areas, hold high standards for yourself and everyone you work with, and constantly look for ways to improve your products' performance, quality, and cost. We're changing an industry, and we want individuals who are ready for this challenge and want to reach beyond what is possible today.\n\nKey job responsibilities\nNPI — New Product Introduction\n\nOwn the end-to-end NPI lifecycle for storage and/or accelerator (AI/ML/GPU) server platforms — from architecture definition through design, qualification, manufacturing ramp, and launch\nLead technical solutions for complex server and rack system architectural challenges\nWork with ODM/manufacturing partners to develop, validate, and manufacture server products at scale\nDevelop functional specifications, design verification plans, and test procedures\nDrive qualification and readiness milestones, ensuring new platforms meet performance, reliability, and cost targets before fleet deployment\nIdentify and resolve technical risks early in the development cycle — don't let problems reach production\n\nFleet Health, Diagnostics & Automation\n\nOwn fleet health for the server platforms you launch — reliability doesn't end at ship\nDesign and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact\nDrive toward zero-touch operations — help build detection, diagnoses, and remediation of faults without human intervention\nDebug complex system failures in time-sensitive settings — personally diving deep when the problem demands it\nPerform root cause analysis correlating across firmware, kernel, driver, thermal, power, and physical layers\n\nSystems Design & Technical Depth\n\nApply expertise across hardware, software, system design, x86 architecture, processes, and operations (compute, storage, network, GPU)\nDesign and implement solutions to address system-level issues at large scale\nDecompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features\nCollaborate with hardware, software, manufacturing, supply chain, and product management teams\n\nCross-Team Collaboration\n\nWork closely with internal customers to ensure new server hardware meets data path and control path requirements\nIdentify early any potential problems onboarding new servers into customer ecosystems\nCollaborate across Hardware Engineering, component, firmware, test, qualification, and integration teams\nPartner with datacenter operations to close the loop between field failures and design improvements\n\nA day in the life\nYour day-to-day responsibilities include interfacing with internal and external customers to understand product requirements and facilitate system development on top of your server designs. You will learn operational challenges facing our existing fleet with the goal of improving the current customer experience and developing improved systems for future designs. You will work directly with vendors and ODM (manufacture partners) to scale your product. Some days you're reviewing a new platform design with your ODM; other days you're deep in logs and telemetry data chasing a failure mode across the fleet. You thrive\non that range.\nBASIC QUALIFICATIONS\nExperience in developing functional specifications, design verification plans and functional test procedures\nBachelor's degree or above in electrical engineering, computer engineering, or equivalent\nExperience in English-language communication skills, both written and verbal\nExperience with design & innovation and research & development\nKnowledge of operating systems, hardware, storage, network, security, database administration and cloud infrastructure\nExperience in server technologies such as, thermal, mechanical, power, and signal integrity\n5+ years of professional work (non-internship) experience\nPREFERRED QUALIFICATIONS\n5+ years of hardware design and validation of components, subsystems and systems experience\nExperience in server technologies: board design, high-speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking\nExperience developing and executing test procedures for mechanical or electrical systems/components\nExperience working with ODMs/manufacturer through the product development and manufacturing lifecycle\nExperience building predictive failure detection or proactive remediation systems at fleet scale\nExperience with storage/compute/GPU/accelerator platforms including integration, diagnostics, or performance validation\nFamiliarity with PCIe topology, NVLink, NVMe, and accelerator interconnects\nExperience with large-scale datacenter or cloud environments\n\nAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.\n\nLos Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.\n\nOur inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.\n\nThe base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.\n\nUSA, CA, Cupertino - 157,300.00 - 212,800.00 USD annually\nUSA, TX, Austin - 136,000.00 - 184,000.00 USD annually\nUSA, WA, Seattle - 136,000.00 - 184,000.00 USD annually","company":"Amazon Data Services","rawCompany":"amazon data services","city":"Austin","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-07-20T11:59:44.252Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams","description":"DESCRIPTION\nAs a Cloud Hardware Development Engineer, you will be an end-to-end owner of storage and/or accelerator (AI/ML/GPU) server platforms — from New Product Introduction (NPI) through fleet health in production. You own the full lifecycle: design, development, qualification, launch, and ongoing operational excellence of servers running at scale in the AWS fleet.\n\nYou will work closely with internal customers to understand their technical needs and business goals, leveraging your experience with server design and the knowledge of various teams to architect solutions we deploy at scale. To deliver your products, you will work with an interdisciplinary team of component, firmware, power, mechanical, electrical, test, qualification, manufacturing engineers, and lead our ODM (design and manufacturing partners) to bring these servers to the data center. After launch, you own the fleet — monitoring quality, driving reliability improvements, and ensuring servers continue to meet customer requirements throughout their\noperational life.\n\nThis role demands deep technical curiosity and the willingness to jump in and personally solve the hardest problems. When a complex system failure occurs — whether during NPI qualification or in a production fleet of hundreds of thousands of servers — you roll up your sleeves, dive into the details across hardware, firmware, software, and physical layers, and drive to root cause. You don't wait for someone else to figure it out.\n\nYou will own end-to-end system reliability — proactively identifying deficiencies and driving toward zero-touch operations where automation detects, diagnoses, and resolves issues before customer impact. You will decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features, leading delivery yourself and through others in parallel.\n\nThis is a fast-paced, intellectually challenging position. You'll work with thought leaders in multiple technology areas, hold high standards for yourself and everyone you work with, and constantly look for ways to improve your products' performance, quality, and cost. We're changing an industry, and we want individuals who are ready for this challenge and want to reach beyond what is possible today.\n\nKey job responsibilities\nNPI — New Product Introduction\n\nOwn the end-to-end NPI lifecycle for storage and/or accelerator (AI/ML/GPU) server platforms — from architecture definition through design, qualification, manufacturing ramp, and launch\nLead technical solutions for complex server and rack system architectural challenges\nWork with ODM/manufacturing partners to develop, validate, and manufacture server products at scale\nDevelop functional specifications, design verification plans, and test procedures\nDrive qualification and readiness milestones, ensuring new platforms meet performance, reliability, and cost targets before fleet deployment\nIdentify and resolve technical risks early in the development cycle — don't let problems reach production\n\nFleet Health, Diagnostics & Automation\n\nOwn fleet health for the server platforms you launch — reliability doesn't end at ship\nDesign and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact\nDrive toward zero-touch operations — help build detection, diagnoses, and remediation of faults without human intervention\nDebug complex system failures in time-sensitive settings — personally diving deep when the problem demands it\nPerform root cause analysis correlating across firmware, kernel, driver, thermal, power, and physical layers\n\nSystems Design & Technical Depth\n\nApply expertise across hardware, software, system design, x86 architecture, processes, and operations (compute, storage, network, GPU)\nDesign and implement solutions to address system-level issues at large scale\nDecompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features\nCollaborate with hardware, software, manufacturing, supply chain, and product management teams\n\nCross-Team Collaboration\n\nWork closely with internal customers to ensure new server hardware meets data path and control path requirements\nIdentify early any potential problems onboarding new servers into customer ecosystems\nCollaborate across Hardware Engineering, component, firmware, test, qualification, and integration teams\nPartner with datacenter operations to close the loop between field failures and design improvements\n\nA day in the life\nYour day-to-day responsibilities include interfacing with internal and external customers to understand product requirements and facilitate system development on top of your server designs. You will learn operational challenges facing our existing fleet with the goal of improving the current customer experience and developing improved systems for future designs. You will work directly with vendors and ODM (manufacture partners) to scale your product. Some days you're reviewing a new platform design with your ODM; other days you're deep in logs and telemetry data chasing a failure mode across the fleet. You thrive\non that range.\nBASIC QUALIFICATIONS\nExperience in developing functional specifications, design verification plans and functional test procedures\nBachelor's degree or above in electrical engineering, computer engineering, or equivalent\nExperience in English-language communication skills, both written and verbal\nExperience with design & innovation and research & development\nKnowledge of operating systems, hardware, storage, network, security, database administration and cloud infrastructure\nExperience in server technologies such as, thermal, mechanical, power, and signal integrity\n5+ years of professional work (non-internship) experience\nPREFERRED QUALIFICATIONS\n5+ years of hardware design and validation of components, subsystems and systems experience\nExperience in server technologies: board design, high-speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking\nExperience developing and executing test procedures for mechanical or electrical systems/components\nExperience working with ODMs/manufacturer through the product development and manufacturing lifecycle\nExperience building predictive failure detection or proactive remediation systems at fleet scale\nExperience with storage/compute/GPU/accelerator platforms including integration, diagnostics, or performance validation\nFamiliarity with PCIe topology, NVLink, NVMe, and accelerator interconnects\nExperience with large-scale datacenter or cloud environments\n\nAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.\n\nLos Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.\n\nOur inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.\n\nThe base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.\n\nUSA, CA, Cupertino - 157,300.00 - 212,800.00 USD annually\nUSA, TX, Austin - 136,000.00 - 184,000.00 USD annually\nUSA, WA, Seattle - 136,000.00 - 184,000.00 USD annually","datePosted":"2026-07-20T11:59:44.252Z","dateModified":"2026-07-20T11:59:44.252Z","hiringOrganization":{"@type":"Organization","name":"Amazon Data Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"828b1fb3c78f8767c860f450"},"url":"https://jobsearcher.com/jobs/828b1fb3c78f8767c860f450"}}