{"schemaVersion":"jobsearcher.job.v1","id":"f95e0ae4af7b924410d7e898","url":"https://jobsearcher.com/jobs/f95e0ae4af7b924410d7e898","canonicalUrl":"https://jobsearcher.com/jobs/f95e0ae4af7b924410d7e898","title":"Software Engineer, Model Inference, DeepMind","description":"Note: By applying to this position you will have an opportunity to share your preferred working location from the following: London, UK; Mountain View, CA, USA.\nMinimum qualifications:\nBachelor’s degree or equivalent practical experience.\n8 years of experience in software development.\n2 years of experience in deploying and maintaining machine learning models in a live production environment.\nExperience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).\nExperience designing, building, or optimizing model serving infrastructure or inference backends.\n\nPreferred qualifications:\nExperience with developing serving infrastructure.\nExperience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL).\nExperience profiling software to identify performance bottlenecks.\nExperience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism).\nFamiliarity with writing performance-optimized kernels.\nUnderstanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling).\nAbout the job\n\nAt Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.\n\nIn this role, you will be at the forefront of bringing AI research to life. You'll work directly with researchers and engineers to optimize and deploy large language models (LLMs) like Gemini onto Google's production infrastructure, impacting users across a different range of applications. This involves a blend of technical expertise and collaborative problem-solving to ensure both efficiency and quality throughout the entire LLM deployment lifecycle.\n\nThe role includes opportunities for both IC and TL opportunities, and is open to both Software Engineering and Research Engineering backgrounds. There are opportunities across multiple teams, so applicants with both specialist and generalist interests within serving are encouraged to apply.\n\nArtificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.\n\nWe are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.\nIndividual pay is determined by factors including job-related skills, experience, and relevant education or training.\n\nUS: $207000 - $300000 (USD) + 20% bonus target + equity + benefits\n\nLearn more about benefits at Google.\nResponsibilities\nCollaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.\nWork with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.\nIdentify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.\nGain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.\nLeverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).\nGoogle is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.","company":"Deepmind","rawCompany":"deepmind","city":"Mountain View","state":"HI","isRemote":false,"isActive":false,"createdAt":"2026-08-19T17:37:29.615Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer, Model Inference, DeepMind","description":"Note: By applying to this position you will have an opportunity to share your preferred working location from the following: London, UK; Mountain View, CA, USA.\nMinimum qualifications:\nBachelor’s degree or equivalent practical experience.\n8 years of experience in software development.\n2 years of experience in deploying and maintaining machine learning models in a live production environment.\nExperience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).\nExperience designing, building, or optimizing model serving infrastructure or inference backends.\n\nPreferred qualifications:\nExperience with developing serving infrastructure.\nExperience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL).\nExperience profiling software to identify performance bottlenecks.\nExperience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism).\nFamiliarity with writing performance-optimized kernels.\nUnderstanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling).\nAbout the job\n\nAt Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.\n\nIn this role, you will be at the forefront of bringing AI research to life. You'll work directly with researchers and engineers to optimize and deploy large language models (LLMs) like Gemini onto Google's production infrastructure, impacting users across a different range of applications. This involves a blend of technical expertise and collaborative problem-solving to ensure both efficiency and quality throughout the entire LLM deployment lifecycle.\n\nThe role includes opportunities for both IC and TL opportunities, and is open to both Software Engineering and Research Engineering backgrounds. There are opportunities across multiple teams, so applicants with both specialist and generalist interests within serving are encouraged to apply.\n\nArtificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.\n\nWe are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.\nIndividual pay is determined by factors including job-related skills, experience, and relevant education or training.\n\nUS: $207000 - $300000 (USD) + 20% bonus target + equity + benefits\n\nLearn more about benefits at Google.\nResponsibilities\nCollaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.\nWork with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.\nIdentify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.\nGain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.\nLeverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).\nGoogle is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.","datePosted":"2026-08-19T17:37:29.615Z","dateModified":"2026-08-19T17:37:29.615Z","hiringOrganization":{"@type":"Organization","name":"Deepmind","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Mountain View","addressRegion":"HI","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f95e0ae4af7b924410d7e898"},"url":"https://jobsearcher.com/jobs/f95e0ae4af7b924410d7e898"}}