{"schemaVersion":"jobsearcher.job.v1","id":"819a585367be623f0f81bbe4","url":"https://jobsearcher.com/jobs/819a585367be623f0f81bbe4","canonicalUrl":"https://jobsearcher.com/jobs/819a585367be623f0f81bbe4","title":"AI Accelerator Software Principal Engineer- Framework Integration","description":"Description\n\nInvent the future with us.\n\nAmpere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.\n\nAs a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.\n\nJoin us at Ampere and work alongside a passionate and growing team - we’d love to have you apply!\n\nAbout the Role\nAs an AI Accelerator Software Principal Engineer – Framework Integration, you will lead the design and delivery of high-performance, low-latency deep learning inference solutions. You’ll help advance Ampere’s AI software stack by integrating and optimizing popular deep learning frameworks so models can run efficiently across data center and edgeenvironments—meeting the performance and efficiency requirements of next-generation AI workloads.\nYou will operate at the intersection of software engineering, performance engineering, and hardware-aware optimization, contributing to the full stack from model execution to accelerator-ready kernel performance.\nWhat You’ll Achieve:\nFramework integration for accelerator backends\nIntegrate and optimize deep learning frameworks—such as PyTorch, ONNX, and llama.cpp—for the Ampere deep learning accelerator backend, enabling efficient and correct execution across a wide set of model types.\nEnd-to-end deep learning performance acceleration\nGo deep into the full software/hardware execution stack, including:\ninference serving and orchestration\nframework integration layers\ncompiler and graph/runtime support\nruntime libraries and user-mode execution paths\ncompute kernel development\nprofiling, benchmarking, and performance tuning\nModel enablement with quality and speed\nImprove both performance and accuracy for models using popular frameworks, and ensure compatibility with serving ecosystems such as vLLM and SGLang—helping deliver production-ready inference behavior.\nHardware/software co-design and optimization\nPartner with hardware and platform teams to co-optimize AI execution for better outcomes:\nincreased throughput\nreduced latency\nimproved scalability\nbetter resource utilization (compute/memory/IO)\nhigher sustained performance under realistic workloads\nBuild state-of-the-art AI software components\nContribute to the development of software and hardware AI co-processors/accelerators, delivering reusable libraries, optimized execution paths, and robust integration with existing tooling.\nCross-functional collaboration\nWork closely with cross-functional teams (compiler/runtime, kernels, platform, and product engineering) to integrate AI capabilities into Ampere’s cloud-native processor platforms and accelerators.\nAbout You:\nEducation & experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years\nCore framework experience: Strong experience building with or integrating AI frameworks such as PyTorch, llama.cpp, and ONNX.\nLinux + accelerator/runtime expertise (preferred): Experience with developing user-mode drivers, runtime libraries, or low-level integration for GPUs or deep learning accelerators in Linux is a plus.\nStrong systems programming & performance skills:\nExpert in Python and C/C++\nStrong background in performance profiling and tuning (latency/throughput, memory behavior, kernel efficiency)\nDeep ML understanding: Solid understanding of AI/ML concepts including neural networks and data processing frameworks. Experience with modern deep model architectures such as Transformers and Diffusion models is preferred.\nModern AI tooling fluency (preferred): Fluent with modern AI programming tools such as Codex or Claude Code, and comfortable accelerating development workflows.\nWhat We’ll Offer:\n\nAt Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $182,000 and $273,000, except in the San Francisco Bay Area where the range is between $195,000 and $292,000.\n\nOur benefits include health, wellness, and financial programs that support employees through every stage of life.\n\nBenefit highlights include:\nPremium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.\nUnlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.\nA variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.\n\nAnd there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.\n\n#LI-Hybrid\n\n#LI-Hybrid#LI-DR\n#LI-Hybrid\n\nAmpere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.","company":"Ampere Computing","rawCompany":"ampere computing","city":"Portland","state":"ME","isRemote":false,"isActive":false,"createdAt":"2026-08-03T15:56:25.102Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Accelerator Software Principal Engineer- Framework Integration","description":"Description\n\nInvent the future with us.\n\nAmpere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.\n\nAs a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.\n\nJoin us at Ampere and work alongside a passionate and growing team - we’d love to have you apply!\n\nAbout the Role\nAs an AI Accelerator Software Principal Engineer – Framework Integration, you will lead the design and delivery of high-performance, low-latency deep learning inference solutions. You’ll help advance Ampere’s AI software stack by integrating and optimizing popular deep learning frameworks so models can run efficiently across data center and edgeenvironments—meeting the performance and efficiency requirements of next-generation AI workloads.\nYou will operate at the intersection of software engineering, performance engineering, and hardware-aware optimization, contributing to the full stack from model execution to accelerator-ready kernel performance.\nWhat You’ll Achieve:\nFramework integration for accelerator backends\nIntegrate and optimize deep learning frameworks—such as PyTorch, ONNX, and llama.cpp—for the Ampere deep learning accelerator backend, enabling efficient and correct execution across a wide set of model types.\nEnd-to-end deep learning performance acceleration\nGo deep into the full software/hardware execution stack, including:\ninference serving and orchestration\nframework integration layers\ncompiler and graph/runtime support\nruntime libraries and user-mode execution paths\ncompute kernel development\nprofiling, benchmarking, and performance tuning\nModel enablement with quality and speed\nImprove both performance and accuracy for models using popular frameworks, and ensure compatibility with serving ecosystems such as vLLM and SGLang—helping deliver production-ready inference behavior.\nHardware/software co-design and optimization\nPartner with hardware and platform teams to co-optimize AI execution for better outcomes:\nincreased throughput\nreduced latency\nimproved scalability\nbetter resource utilization (compute/memory/IO)\nhigher sustained performance under realistic workloads\nBuild state-of-the-art AI software components\nContribute to the development of software and hardware AI co-processors/accelerators, delivering reusable libraries, optimized execution paths, and robust integration with existing tooling.\nCross-functional collaboration\nWork closely with cross-functional teams (compiler/runtime, kernels, platform, and product engineering) to integrate AI capabilities into Ampere’s cloud-native processor platforms and accelerators.\nAbout You:\nEducation & experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years\nCore framework experience: Strong experience building with or integrating AI frameworks such as PyTorch, llama.cpp, and ONNX.\nLinux + accelerator/runtime expertise (preferred): Experience with developing user-mode drivers, runtime libraries, or low-level integration for GPUs or deep learning accelerators in Linux is a plus.\nStrong systems programming & performance skills:\nExpert in Python and C/C++\nStrong background in performance profiling and tuning (latency/throughput, memory behavior, kernel efficiency)\nDeep ML understanding: Solid understanding of AI/ML concepts including neural networks and data processing frameworks. Experience with modern deep model architectures such as Transformers and Diffusion models is preferred.\nModern AI tooling fluency (preferred): Fluent with modern AI programming tools such as Codex or Claude Code, and comfortable accelerating development workflows.\nWhat We’ll Offer:\n\nAt Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $182,000 and $273,000, except in the San Francisco Bay Area where the range is between $195,000 and $292,000.\n\nOur benefits include health, wellness, and financial programs that support employees through every stage of life.\n\nBenefit highlights include:\nPremium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.\nUnlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.\nA variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.\n\nAnd there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.\n\n#LI-Hybrid\n\n#LI-Hybrid#LI-DR\n#LI-Hybrid\n\nAmpere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.","datePosted":"2026-08-03T15:56:25.102Z","dateModified":"2026-08-03T15:56:25.102Z","hiringOrganization":{"@type":"Organization","name":"Ampere Computing","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Portland","addressRegion":"ME","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"819a585367be623f0f81bbe4"},"url":"https://jobsearcher.com/jobs/819a585367be623f0f81bbe4"}}