{"schemaVersion":"jobsearcher.job.v1","id":"e9f4b0ea8d9bc7e96f45c1a8","url":"https://jobsearcher.com/jobs/e9f4b0ea8d9bc7e96f45c1a8","canonicalUrl":"https://jobsearcher.com/jobs/e9f4b0ea8d9bc7e96f45c1a8","title":"Principal SoC Performance Architect-Microbenchmarks","description":"OverviewAMD is looking for an outstanding technical contributor to drive performance analysis, characterization, and optimization of next-generation Data Center GPU (DCGPU) platforms. This role focuses on extracting maximum performance across the full system stack—including hardware, firmware, drivers, runtime, libraries, and workloads—through deep architectural understanding and data-driven methodologies. The engineer will develop and maintain microbenchmarks and system-level workloads spanning pre-silicon and post-silicon environments to enable performance validation, debug, and optimization.The PersonAs a passionate and technically strong Principal SoC Performance Engineer, you will work on highly parallel SoC architectures, leveraging deep understanding of GPU compute, memory hierarchy, and interconnects to analyze and optimize performance across AI and HPC workloads. You will be responsible for building microbenchmark suites and workload-driven analysis frameworks that expose performance characteristics of GPU subsystems (compute, memory, IO, interconnect, collectives) and ensure continuity across pre-silicon models, emulation, and post-silicon systems. The ideal candidate combines strong hardware/software co-design expertise with hands-on experience in performance analysis, profiling, and system-level debugging. You are expected to identify bottlenecks across the entire stack—from kernels to runtime to hardware—and translate insights into actionable improvements for both current and future architectures. You thrive in a fast-paced environment, are highly data-driven, and have a deep curiosity for understanding \"why\" performance behaves the way it does.Key ResponsibilitiesPerformance Analysis & Optimization: Analyze and optimize performance of DCGPU systems across AI training, inference, and HPC workloads.Identify bottlenecks across hardware, firmware, drivers, runtime, libraries, and applications.Perform deep kernel-level and system-level profiling to understand performance behavior.Provide actionable insights to architecture, software, and design teams to improve performance.Microbenchmark & Workload DevelopmentDesign and develop targeted microbenchmarks to characterize GPU subsystems (compute, memory, interconnect, collectives).Build representative system-level workloads reflecting real-world AI/HPC use cases.Ensure microbenchmarks correlate to application-level performance and architectural intent.Maintain and evolve benchmark suites across multiple GPU generations.Pre-Silicon & Post-Silicon ContinuityEnable performance validation in pre-silicon environments (simulation/emulation/models).Correlate performance data across pre-silicon models and post-silicon measurements.Develop methodologies to reuse workloads and microbenchmarks across the full lifecycle.Support bring-up and early silicon performance characterization.Full-Stack Performance EngineeringWork across the entire software stack: compiler, runtime, libraries, drivers, and firmware.Collaborate with ROCm / AI frameworks / kernel teams to improve performance.Analyze interactions between workload characteristics and hardware execution.Optimize key kernels (e.g., GEMMs, collectives, attention) and system-level behavior.Tooling & InfrastructureDevelop and enhance performance measurement, profiling, and analysis tools.Enable scalable, repeatable workflows for benchmarking and analysis.Build automation for performance regression tracking and reporting.Contribute to unified infrastructure spanning pre-silicon and post-silicon environments.Cross-Functional CollaborationPartner with SoC architecture, GPU IP, software, and system teams.Influence design decisions using data-driven performance insights.Collaborate with competitive analysis teams to understand gaps vs. industry platforms.Performance Modeling & InsightsDevelop strong intuition and/or models for performance scaling and limits.Translate performance data into architectural feedback for future GPU designs.Support competitive benchmarking and performance projections.Preferred Experience10–15+ years of experience in performance engineering for GPUs, HPC systems, or highly parallel SoCs.Strong understanding of GPU architecture, parallel computing, and memory hierarchies.Experience with microbenchmark development and system-level workload analysis.Hands-on experience with performance profiling tools (rocprof, Nsight, perf, etc.).Experience analyzing AI/HPC workloads (LLMs, training, inference, communication libraries like RCCL/NCCL).Strong background in hardware/software co-design and performance optimization.Familiarity with pre-silicon (simulation/emulation/models) and post-silicon performance workflows.Programming expertise in C/C++, Python; experience with GPU programming models (HIP, CUDA, OpenCL).Strong analytical and debugging skills with a data-driven mindset.Experience working across full software stack (compiler → runtime → kernels → system).Exposure to performance modeling, scaling analysis, or competitive benchmarking is a plus.Position RequirementsProven experience working on highly parallel compute systems or SoCs (GPUs preferred).Experience developing and maintaining microbenchmarks tied to architectural features.Strong exposure to performance analysis across pre-silicon and post-silicon environments.Solid understanding of GPU compute, memory systems, and interconnect architectures.Experience with profiling, tracing, and performance counter analysis.Ability to debug complex system-level performance issues across multiple layers.MS/PhD in Computer Engineering, Computer Science, or related field.Excellent communication skills and ability to present complex performance insights clearly.Academic CredentialsBachelor's or Master's degree in related discipline preferred.This role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's \"Responsible AI Policy\" is available here. This posting is for an existing vacancy. #J-18808-Ljbffr","company":"AMD","rawCompany":"amd","city":"Austin","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-04T23:46:00.709Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"17-2199.00","title":"Engineers, All Other","slug":"engineers-all-other"}],"industries":[{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"334413","title":"Semiconductor and Related Device Manufacturing","slug":"semiconductor-and-related-device-manufacturing"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal SoC Performance Architect-Microbenchmarks","description":"OverviewAMD is looking for an outstanding technical contributor to drive performance analysis, characterization, and optimization of next-generation Data Center GPU (DCGPU) platforms. This role focuses on extracting maximum performance across the full system stack—including hardware, firmware, drivers, runtime, libraries, and workloads—through deep architectural understanding and data-driven methodologies. The engineer will develop and maintain microbenchmarks and system-level workloads spanning pre-silicon and post-silicon environments to enable performance validation, debug, and optimization.The PersonAs a passionate and technically strong Principal SoC Performance Engineer, you will work on highly parallel SoC architectures, leveraging deep understanding of GPU compute, memory hierarchy, and interconnects to analyze and optimize performance across AI and HPC workloads. You will be responsible for building microbenchmark suites and workload-driven analysis frameworks that expose performance characteristics of GPU subsystems (compute, memory, IO, interconnect, collectives) and ensure continuity across pre-silicon models, emulation, and post-silicon systems. The ideal candidate combines strong hardware/software co-design expertise with hands-on experience in performance analysis, profiling, and system-level debugging. You are expected to identify bottlenecks across the entire stack—from kernels to runtime to hardware—and translate insights into actionable improvements for both current and future architectures. You thrive in a fast-paced environment, are highly data-driven, and have a deep curiosity for understanding \"why\" performance behaves the way it does.Key ResponsibilitiesPerformance Analysis & Optimization: Analyze and optimize performance of DCGPU systems across AI training, inference, and HPC workloads.Identify bottlenecks across hardware, firmware, drivers, runtime, libraries, and applications.Perform deep kernel-level and system-level profiling to understand performance behavior.Provide actionable insights to architecture, software, and design teams to improve performance.Microbenchmark & Workload DevelopmentDesign and develop targeted microbenchmarks to characterize GPU subsystems (compute, memory, interconnect, collectives).Build representative system-level workloads reflecting real-world AI/HPC use cases.Ensure microbenchmarks correlate to application-level performance and architectural intent.Maintain and evolve benchmark suites across multiple GPU generations.Pre-Silicon & Post-Silicon ContinuityEnable performance validation in pre-silicon environments (simulation/emulation/models).Correlate performance data across pre-silicon models and post-silicon measurements.Develop methodologies to reuse workloads and microbenchmarks across the full lifecycle.Support bring-up and early silicon performance characterization.Full-Stack Performance EngineeringWork across the entire software stack: compiler, runtime, libraries, drivers, and firmware.Collaborate with ROCm / AI frameworks / kernel teams to improve performance.Analyze interactions between workload characteristics and hardware execution.Optimize key kernels (e.g., GEMMs, collectives, attention) and system-level behavior.Tooling & InfrastructureDevelop and enhance performance measurement, profiling, and analysis tools.Enable scalable, repeatable workflows for benchmarking and analysis.Build automation for performance regression tracking and reporting.Contribute to unified infrastructure spanning pre-silicon and post-silicon environments.Cross-Functional CollaborationPartner with SoC architecture, GPU IP, software, and system teams.Influence design decisions using data-driven performance insights.Collaborate with competitive analysis teams to understand gaps vs. industry platforms.Performance Modeling & InsightsDevelop strong intuition and/or models for performance scaling and limits.Translate performance data into architectural feedback for future GPU designs.Support competitive benchmarking and performance projections.Preferred Experience10–15+ years of experience in performance engineering for GPUs, HPC systems, or highly parallel SoCs.Strong understanding of GPU architecture, parallel computing, and memory hierarchies.Experience with microbenchmark development and system-level workload analysis.Hands-on experience with performance profiling tools (rocprof, Nsight, perf, etc.).Experience analyzing AI/HPC workloads (LLMs, training, inference, communication libraries like RCCL/NCCL).Strong background in hardware/software co-design and performance optimization.Familiarity with pre-silicon (simulation/emulation/models) and post-silicon performance workflows.Programming expertise in C/C++, Python; experience with GPU programming models (HIP, CUDA, OpenCL).Strong analytical and debugging skills with a data-driven mindset.Experience working across full software stack (compiler → runtime → kernels → system).Exposure to performance modeling, scaling analysis, or competitive benchmarking is a plus.Position RequirementsProven experience working on highly parallel compute systems or SoCs (GPUs preferred).Experience developing and maintaining microbenchmarks tied to architectural features.Strong exposure to performance analysis across pre-silicon and post-silicon environments.Solid understanding of GPU compute, memory systems, and interconnect architectures.Experience with profiling, tracing, and performance counter analysis.Ability to debug complex system-level performance issues across multiple layers.MS/PhD in Computer Engineering, Computer Science, or related field.Excellent communication skills and ability to present complex performance insights clearly.Academic CredentialsBachelor's or Master's degree in related discipline preferred.This role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's \"Responsible AI Policy\" is available here. This posting is for an existing vacancy. #J-18808-Ljbffr","datePosted":"2026-08-04T23:46:00.709Z","dateModified":"2026-08-04T23:46:00.709Z","hiringOrganization":{"@type":"Organization","name":"AMD","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e9f4b0ea8d9bc7e96f45c1a8"},"url":"https://jobsearcher.com/jobs/e9f4b0ea8d9bc7e96f45c1a8"}}