{"schemaVersion":"jobsearcher.job.v1","id":"72518881f7230577941343b8","url":"https://jobsearcher.com/jobs/72518881f7230577941343b8","canonicalUrl":"https://jobsearcher.com/jobs/72518881f7230577941343b8","title":"Fellow Software Development Engineer, GEMM Optimization","description":"Overview:\n\nADVANCE YOUR CAREER. ADVANCE THE WORLD.\n\nAt AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.\n\nWhether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.\n\nResponsibilities:\n\nTHE ROLE:\n\nAMD is seeking a Fellow Software Development Engineer to architect and develop features in the tensile-lite assembly code generator for General Matrix Multiplication (GEMM) and GEMM + X (including attention). GEMM performance is one of the most critical levers for AI workload efficiency and this role sits at the center of AMD's effort to deliver best-in-class matrix-multiply performance across current and next-generation hardware.\n\nYou will own the deep, hardware-aware optimization work that turns AMD's GPU architecture into delivered performance: understanding the hardware specification at the instruction and pipeline level, partnering directly with hardware architects to influence future designs, and innovating algorithms to extract maximum throughput.\n\nTHE PERSON:\n\nThis role is a strong fit for an engineer who wants to work at the intersection of GPU microarchitecture, low-level code generation, and applied AI performance — someone equally comfortable reading a hardware spec, writing hand-tuned assembly, and explaining a performance tradeoff to a customer or an architect.\n\nKEY RESPONSIBILITIES\n\nAnalyze AMD GPU hardware specifications in depth and work closely with hardware architects to align kernel design with current silicon capabilities and to influence requirements for future generations.\nDevelop support for new ISA and HW/SW optimization features in the code-generator\nInnovate and implement new algorithms for implementing GEMM and GEMM + X\nProfile and root-cause GEMM performance bottlenecks across the stack\nPartner with customers and internal stakeholders to understand real-world workload requirements, reproduce performance issues, and deliver targeted performance\n\nPREFERRED EXPERIENCE:\n\nStrong command of GPU computer architecture: compute units, register files, cache/LDS hierarchies, memory bandwidth, matrix cores/WMMA-style instructions, and instruction scheduling/latency hiding.\nDeep expertise in GEMM algorithms and their mapping onto GPU hardware (tiling, blocking, register/LDS allocation, wave scheduling, memory hierarchy utilization).\nStrong software engineering fundamentals with a track record of building production-quality, high-performance software.\nDemonstrated experience optimizing GPU compute kernels, ideally GEMM and attention\nProficiency in C++ and assembly-level GPU programming; working knowledge of Python for tooling/automation.\nSolid understanding of parallel computing, memory hierarchies, and hardware-software performance tradeoffs.\nClear written and verbal communication skills; ability to work effectively with hardware architects, customers, and cross-functional engineering teams.\nDirect experience with any GPU compiler backend\nExperience with AMD GPU architectures (CDNA/RDNA) or comparable competitive architectures (NVIDIA Hopper/Blackwell, etc.)\nContributions to industry standards, publications, or patents related to GPU compute or matrix-multiply optimization\n\nPREFERRED ACADEMIC CREDENTIALS:\n\nBS/MS/PHD in CS/CE or related field with deep relevant experience\n\nLOCATION: San Jose, CA\n\nQualifications:\n\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","company":"Advanced Micro Devices","rawCompany":"advanced micro devices","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-04T11:46:27.451Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"334413","title":"Semiconductor and Related Device Manufacturing","slug":"semiconductor-and-related-device-manufacturing"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Fellow Software Development Engineer, GEMM Optimization","description":"Overview:\n\nADVANCE YOUR CAREER. ADVANCE THE WORLD.\n\nAt AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.\n\nWhether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.\n\nResponsibilities:\n\nTHE ROLE:\n\nAMD is seeking a Fellow Software Development Engineer to architect and develop features in the tensile-lite assembly code generator for General Matrix Multiplication (GEMM) and GEMM + X (including attention). GEMM performance is one of the most critical levers for AI workload efficiency and this role sits at the center of AMD's effort to deliver best-in-class matrix-multiply performance across current and next-generation hardware.\n\nYou will own the deep, hardware-aware optimization work that turns AMD's GPU architecture into delivered performance: understanding the hardware specification at the instruction and pipeline level, partnering directly with hardware architects to influence future designs, and innovating algorithms to extract maximum throughput.\n\nTHE PERSON:\n\nThis role is a strong fit for an engineer who wants to work at the intersection of GPU microarchitecture, low-level code generation, and applied AI performance — someone equally comfortable reading a hardware spec, writing hand-tuned assembly, and explaining a performance tradeoff to a customer or an architect.\n\nKEY RESPONSIBILITIES\n\nAnalyze AMD GPU hardware specifications in depth and work closely with hardware architects to align kernel design with current silicon capabilities and to influence requirements for future generations.\nDevelop support for new ISA and HW/SW optimization features in the code-generator\nInnovate and implement new algorithms for implementing GEMM and GEMM + X\nProfile and root-cause GEMM performance bottlenecks across the stack\nPartner with customers and internal stakeholders to understand real-world workload requirements, reproduce performance issues, and deliver targeted performance\n\nPREFERRED EXPERIENCE:\n\nStrong command of GPU computer architecture: compute units, register files, cache/LDS hierarchies, memory bandwidth, matrix cores/WMMA-style instructions, and instruction scheduling/latency hiding.\nDeep expertise in GEMM algorithms and their mapping onto GPU hardware (tiling, blocking, register/LDS allocation, wave scheduling, memory hierarchy utilization).\nStrong software engineering fundamentals with a track record of building production-quality, high-performance software.\nDemonstrated experience optimizing GPU compute kernels, ideally GEMM and attention\nProficiency in C++ and assembly-level GPU programming; working knowledge of Python for tooling/automation.\nSolid understanding of parallel computing, memory hierarchies, and hardware-software performance tradeoffs.\nClear written and verbal communication skills; ability to work effectively with hardware architects, customers, and cross-functional engineering teams.\nDirect experience with any GPU compiler backend\nExperience with AMD GPU architectures (CDNA/RDNA) or comparable competitive architectures (NVIDIA Hopper/Blackwell, etc.)\nContributions to industry standards, publications, or patents related to GPU compute or matrix-multiply optimization\n\nPREFERRED ACADEMIC CREDENTIALS:\n\nBS/MS/PHD in CS/CE or related field with deep relevant experience\n\nLOCATION: San Jose, CA\n\nQualifications:\n\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","datePosted":"2026-09-04T11:46:27.451Z","dateModified":"2026-09-04T11:46:27.451Z","hiringOrganization":{"@type":"Organization","name":"Advanced Micro Devices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"72518881f7230577941343b8"},"url":"https://jobsearcher.com/jobs/72518881f7230577941343b8"}}