{"schemaVersion":"jobsearcher.job.v1","id":"b5e027aa899efce30082c006","url":"https://jobsearcher.com/jobs/b5e027aa899efce30082c006","canonicalUrl":"https://jobsearcher.com/jobs/b5e027aa899efce30082c006","title":"AI Accelerator Software Distinguished Engineer- Framework Integration","description":"Description\n\nInvent the future with us.\n\nAmpere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.\n\nAs a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.\n\nJoin us at Ampere and work alongside a passionate and growing team - we’d love to have you apply!\n\nAbout the Role\n\nAs an AI Accelerator Software Distinguished Engineer – Framework Integration, you will lead end-to-end technical strategy and delivery for high-performance deep learning inference across Ampere accelerator platforms. You will set direction for how major ML frameworks are enabled and optimized for our hardware, ensuring high throughput, low latency, and efficient compute/memory utilization for current and next-generation AI workloads spanning data centers to edge.\n\nThis role is distinguished by deep technical ownership, architecture leadership, and cross-team influence—driving outcomes from performance modeling and integration strategy through production-ready runtime and kernel behavior.\n\nWhat You’ll Achieve:\n\nFramework integration leadership (PyTorch / ONNX / llama.cpp)\nOwn and advance integration of major deep learning frameworks—PyTorch, ONNX, llama.cpp, and related tooling—into the Ampere deep learning accelerator backend, enabling robust execution of real-world model graphs and operators.\nFull-stack acceleration across the SW/HW execution path\nDrive acceleration across the end-to-end stack, including (as applicable):\ninference serving and orchestration enablement\nframework-to-runtime integration layers\ncompiler/graph lowering and optimization\nruntime library and execution management\nuser-mode execution paths and performance-critical interfaces\ncompute kernel development and micro-optimizations\nprofiling, benchmarking, and continuous performance tuning\nModel enablement: performance + accuracy\nLead efforts to improve performance and correctness for models using popular frameworks and serving stacks such as vLLM and SGLang, ensuring stable behavior under production inference patterns (prefill/decode, batching, KV cache behavior, scheduling, etc.).\nHardware/software co-design and optimization\nProvide technical direction for HW/SW co-optimization of existing and evolving AI architectures to:\nmaximize computational efficiency\nincrease sustained throughput\nreduce latency and variance\nimprove scalability across cores, memory hierarchies, and system configurations\nraise the ceiling on what Ampere platforms can deliver\nBuild and evolve state-of-the-art AI accelerator software\nContribute to and shape the architecture of software/hardware AI co-processors and accelerators, defining reusable components, reference implementations, and performance guardrails.\nCross-functional technical collaboration and influence\nPartner with compiler/runtime/kernel, platform, and systems teams to integrate and validate AI capabilities in Ampere’s platforms and accelerators from cloud to edge.\n\nMentorship and technical excellence\nSet engineering standards through code reviews, design reviews, benchmark methodologies, and mentorship—raising the overall bar for quality, maintainability, and performance.\n\nAbout You:\n\nEducation & experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or a related technical field & 15 years of related experience; or MS degree & 12 years; or PhD & 8 years\n\nDeep framework expertise: Proven experience with software development focused on PyTorch, ONNX, and llama.cpp, including integration, graph/operator enablement, and performance-focused engineering.\n\nLinux accelerator runtime / driver experience (preferred): Experience building or extending user-mode drivers and/or runtime libraries for GPUs or deep learning accelerators on Linux is a plus.\n\nStrong systems programming + performance tuning: Deep expertise in Python and C/C++, with a strong track record in performance engineering (profiling, optimization, throughput/latency analysis, memory behavior).\n\nSolid ML/AI fundamentals: Strong understanding of AI/ML concepts (neural networks, data processing frameworks), and familiarity with modern model families including Transformers and Diffusion architectures.\n\nFluent with modern AI development tools (preferred): Comfortable using modern AI programming tools and workflows such as Codex or Claude Code.\nWhat We’ll Offer:\n\nAt Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $279,500 and $419,500, except in the San Francisco Bay Area where the range is between $293,000 and $440,000.\n\nOur benefits include health, wellness, and financial programs that support employees through every stage of life.\n\nBenefit highlights include:\nPremium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.\nUnlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.\nA variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.\n\nAnd there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.\n\n#LI-Hybrid\n\n#LI-Hybrid#LI-DR\n#LI-Hybrid\n\nAmpere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.","company":"Ampere Computing","rawCompany":"ampere computing","city":"Portland","state":"ME","isRemote":false,"isActive":false,"createdAt":"2026-08-03T15:56:25.104Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Accelerator Software Distinguished Engineer- Framework Integration","description":"Description\n\nInvent the future with us.\n\nAmpere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.\n\nAs a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.\n\nJoin us at Ampere and work alongside a passionate and growing team - we’d love to have you apply!\n\nAbout the Role\n\nAs an AI Accelerator Software Distinguished Engineer – Framework Integration, you will lead end-to-end technical strategy and delivery for high-performance deep learning inference across Ampere accelerator platforms. You will set direction for how major ML frameworks are enabled and optimized for our hardware, ensuring high throughput, low latency, and efficient compute/memory utilization for current and next-generation AI workloads spanning data centers to edge.\n\nThis role is distinguished by deep technical ownership, architecture leadership, and cross-team influence—driving outcomes from performance modeling and integration strategy through production-ready runtime and kernel behavior.\n\nWhat You’ll Achieve:\n\nFramework integration leadership (PyTorch / ONNX / llama.cpp)\nOwn and advance integration of major deep learning frameworks—PyTorch, ONNX, llama.cpp, and related tooling—into the Ampere deep learning accelerator backend, enabling robust execution of real-world model graphs and operators.\nFull-stack acceleration across the SW/HW execution path\nDrive acceleration across the end-to-end stack, including (as applicable):\ninference serving and orchestration enablement\nframework-to-runtime integration layers\ncompiler/graph lowering and optimization\nruntime library and execution management\nuser-mode execution paths and performance-critical interfaces\ncompute kernel development and micro-optimizations\nprofiling, benchmarking, and continuous performance tuning\nModel enablement: performance + accuracy\nLead efforts to improve performance and correctness for models using popular frameworks and serving stacks such as vLLM and SGLang, ensuring stable behavior under production inference patterns (prefill/decode, batching, KV cache behavior, scheduling, etc.).\nHardware/software co-design and optimization\nProvide technical direction for HW/SW co-optimization of existing and evolving AI architectures to:\nmaximize computational efficiency\nincrease sustained throughput\nreduce latency and variance\nimprove scalability across cores, memory hierarchies, and system configurations\nraise the ceiling on what Ampere platforms can deliver\nBuild and evolve state-of-the-art AI accelerator software\nContribute to and shape the architecture of software/hardware AI co-processors and accelerators, defining reusable components, reference implementations, and performance guardrails.\nCross-functional technical collaboration and influence\nPartner with compiler/runtime/kernel, platform, and systems teams to integrate and validate AI capabilities in Ampere’s platforms and accelerators from cloud to edge.\n\nMentorship and technical excellence\nSet engineering standards through code reviews, design reviews, benchmark methodologies, and mentorship—raising the overall bar for quality, maintainability, and performance.\n\nAbout You:\n\nEducation & experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or a related technical field & 15 years of related experience; or MS degree & 12 years; or PhD & 8 years\n\nDeep framework expertise: Proven experience with software development focused on PyTorch, ONNX, and llama.cpp, including integration, graph/operator enablement, and performance-focused engineering.\n\nLinux accelerator runtime / driver experience (preferred): Experience building or extending user-mode drivers and/or runtime libraries for GPUs or deep learning accelerators on Linux is a plus.\n\nStrong systems programming + performance tuning: Deep expertise in Python and C/C++, with a strong track record in performance engineering (profiling, optimization, throughput/latency analysis, memory behavior).\n\nSolid ML/AI fundamentals: Strong understanding of AI/ML concepts (neural networks, data processing frameworks), and familiarity with modern model families including Transformers and Diffusion architectures.\n\nFluent with modern AI development tools (preferred): Comfortable using modern AI programming tools and workflows such as Codex or Claude Code.\nWhat We’ll Offer:\n\nAt Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $279,500 and $419,500, except in the San Francisco Bay Area where the range is between $293,000 and $440,000.\n\nOur benefits include health, wellness, and financial programs that support employees through every stage of life.\n\nBenefit highlights include:\nPremium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.\nUnlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.\nA variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.\n\nAnd there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.\n\n#LI-Hybrid\n\n#LI-Hybrid#LI-DR\n#LI-Hybrid\n\nAmpere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.","datePosted":"2026-08-03T15:56:25.104Z","dateModified":"2026-08-03T15:56:25.104Z","hiringOrganization":{"@type":"Organization","name":"Ampere Computing","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Portland","addressRegion":"ME","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"b5e027aa899efce30082c006"},"url":"https://jobsearcher.com/jobs/b5e027aa899efce30082c006"}}