{"schemaVersion":"jobsearcher.job.v1","id":"beb3d2d4a1d8bb44e620322e","url":"https://jobsearcher.com/jobs/beb3d2d4a1d8bb44e620322e","canonicalUrl":"https://jobsearcher.com/jobs/beb3d2d4a1d8bb44e620322e","title":"System Design Engineer - AI Cluster Software Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE:\nThis is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster.\n\nTHE PERSON:\nDemonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools\nFive or more years of professional software development experience, including substantial experience building and supporting web applications\nProficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue\nExperience developing backend services and APIs using a modern server-side language or framework\nExperience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure\nExperience with automated testing, source control, code review, debugging, and production support\nWorking knowledge of web application security, authentication, authorization, and secure secrets handling\n\nKEY RESPONSIBILITIES:\nPartner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions\nHands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities\nDesign and development of cohesive interface code between disparate third party tools\nOwn features from requirements and design through deployment and ongoing maintenance\nWork in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)\nParticipate in code reviews and retros; adapt to changing requirements and priorities\n\nPREFERRED EXPERIENCE:\nStrong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).\nClear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps\nAMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning\nOrchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity\nAutomation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)\nStorage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies\nHands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial\n\nACADEMIC CREDENTIALS:\nBachelors or Masters degree in computer science or software/computer engineering\n\nLOCATIONS:\nAustin, TX\nSanta Clara, CA\nSecaucus, NJ\nMarkham, Ontario, CA\n\nThis role is not eligible for visa sponsorship.\n\n#LI-CB1\n#LI-Hybrid\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","company":"Advanced Micro Devices","rawCompany":"advanced micro devices","city":"Austin","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-03T22:52:07.698Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"System Design Engineer - AI Cluster Software Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE:\nThis is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster.\n\nTHE PERSON:\nDemonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools\nFive or more years of professional software development experience, including substantial experience building and supporting web applications\nProficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue\nExperience developing backend services and APIs using a modern server-side language or framework\nExperience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure\nExperience with automated testing, source control, code review, debugging, and production support\nWorking knowledge of web application security, authentication, authorization, and secure secrets handling\n\nKEY RESPONSIBILITIES:\nPartner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions\nHands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities\nDesign and development of cohesive interface code between disparate third party tools\nOwn features from requirements and design through deployment and ongoing maintenance\nWork in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)\nParticipate in code reviews and retros; adapt to changing requirements and priorities\n\nPREFERRED EXPERIENCE:\nStrong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).\nClear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps\nAMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning\nOrchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity\nAutomation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)\nStorage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies\nHands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial\n\nACADEMIC CREDENTIALS:\nBachelors or Masters degree in computer science or software/computer engineering\n\nLOCATIONS:\nAustin, TX\nSanta Clara, CA\nSecaucus, NJ\nMarkham, Ontario, CA\n\nThis role is not eligible for visa sponsorship.\n\n#LI-CB1\n#LI-Hybrid\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","datePosted":"2026-08-03T22:52:07.698Z","dateModified":"2026-08-03T22:52:07.698Z","hiringOrganization":{"@type":"Organization","name":"Advanced Micro Devices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"beb3d2d4a1d8bb44e620322e"},"url":"https://jobsearcher.com/jobs/beb3d2d4a1d8bb44e620322e"}}