{"schemaVersion":"jobsearcher.job.v1","id":"6a2b84ba2d0b55b833ae88ee","url":"https://jobsearcher.com/jobs/6a2b84ba2d0b55b833ae88ee","canonicalUrl":"https://jobsearcher.com/jobs/6a2b84ba2d0b55b833ae88ee","title":"Lead Machine Learning Engineer, Inference & Performance","description":"About Egen:\nEgen is a fast-growing and entrepreneurial company with a data-first mindset. We bring together the best engineering talent working with the most advanced technology platforms, including Google Cloud and Salesforce, to help clients drive action and impact through data and insights. We are committed to being a place where the best people choose to work so they can apply their engineering and technology expertise to envision what is next for how data and platforms can change the world for the better. We are dedicated to learning, thrive on solving tough problems, and continually innovate to achieve fast, effective results. If this describes you, we want you on our team.\nWant to learn more about life at Egen? Check out these resources in addition to the job description.\nMeet Egen\nLife at Egen\nCulture and Values at Egen\nCareer Development at Egen\nBenefits at Egen\nAbout the opportunity:\nAs a Senior AI Engineer, you will be at the forefront of our Generative AI initiatives. We treat AI as a software engineering discipline. You will be responsible for the full lifecycle of our AI features—specifically document intelligence and RAG pipelines—taking them from initial prototype to robust, scalable production services. You will solve for real-world constraints like latency, error handling, and cost optimization.\nYou’ll collaborate with a diverse range of clients to translate business needs into high-performance AI architectures. This role requires a blend of deep technical expertise in LLMs and a disciplined Software Engineering approach to ensure our solutions are robust, ethical, and scalable.\nWhat You Will Do:\nOptimize Inference: Build and tune production LLM serving with vLLM and SGLang—maximizing throughput and minimizing latency through batching, paged attention, quantization, and KV-cache strategies\nProfile & Accelerate Training: Instrument and profile training runs to find bottlenecks, then resolve them with the right attention implementations (e.g. FlashAttention) tuned to the underlying hardware (H200, GB200)\nEngineer for the Hardware: Apply a working understanding of GPU architecture and attention internals to choose the right approach per accelerator, rather than relying on defaults\nServe at Scale: Deploy and operate multiple models within shared GPU clusters on GKE, with autoscaling, efficient bin-packing, and graceful handling of mixed workloads\nDrive Efficiency: Own GPU utilization as a first-class metric—measure it, improve throughput-per-dollar, and continuously raise the ceiling on what our fleet can deliver\nCollaborate & Consult: Work directly with clients to understand performance, latency, and cost requirements, and translate them into pragmatic serving and training architectures\nYour Technical Toolkit:\nCore Languages: Mastery of Python and shell scripting; comfort reading and reasoning about lower-level (CUDA-adjacent) performance code is a strong plus\nInference Frameworks: Hands-on experience with vLLM, SGLash, or comparable high-performance serving stacks\nGPU & Model Internals: Solid grasp of GPU architecture, the fundamentals of LLM inference, and the attention mechanism—including where the bottlenecks live and how FlashAttention and similar techniques address them across hardware generations (H200, GB200)\nProfiling: Fluency with profiling tools to diagnose training and inference bottlenecks (compute-bound vs. memory-bound, kernel-level analysis)\nInfrastructure: Strong Kubernetes (GKE) experience—deploying and autoscaling multiple models on shared GPU clusters on Google Cloud\nMindset: A strong software engineering foundation—you write clean, maintainable code, measure before optimizing, and understand the full SDLC\nBasic Qualifications:\nBachelor's or Master's degree in Computer Science, Engineering, or a related technical field\n5+ years of experience in ML/AI engineering, with a meaningful portion focused on performance, infrastructure, or systems\nProven track record of deploying and optimizing models in a production environment\nDemonstrated experience profiling and improving GPU utilization for training and/or inference\nExperience with Classic Machine Learning (neural nets, training, tuning) is a strong plus\nKnowledge of Data Engineering and SQL\nPersonal Attributes:\nOwnership: You take pride in your work and see optimizations through from profile to production\nCuriosity: Hardware and serving frameworks change fast; you are a lifelong learner who stays ahead of the curve\nRigor: You measure before you optimize and let data, not intuition, guide where you spend effort\nConsultative Spirit: You enjoy interacting with clients and can translate technical complexity into business value\nEthics: You prioritize responsible AI development and data privacy\n$159,300 - $250,100 a year\nThis position may be hired at multiple levels. Final leveling is determined during the interview process based on a candidate's experience, skills, and interview outcomes, and the applicable salary range will align with the final level assigned.\n\nCompensation is determined based on factors including experience, expertise, interview performance, and geographic location, where applicable.\nCompensation & Benefits:\nThis role is eligible for our competitive salary and comprehensive benefits package to support your well-being:\nComprehensive Health Insurance\nPaid Leave (Vacation/PTO)\nPaid Holidays\nSick Leave\nParental Leave\nBereavement Leave\n401 (k) Employer Match\nEmployee Referral Bonuses\nCheck out our complete list of benefits here - >https://egen.ai/people/#benefits\nImportant: All roles are subject to standard hiring verification practices, which may include background checks, employment verification, and other relevant checks.\nEEO and Accommodations:\nEgen is an equal opportunity employer and is committed to inclusion, diversity, and equity in the workplace. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veterans’ status, or any other characteristic protected by federal, state, or local laws. Egen will also consider qualified applications with criminal histories, consistent with legal requirements. Egen welcomes and encourages applications from individuals with disabilities. Reasonable accommodations are available for candidates during all aspects of the selection process. Please advise the talent acquisition team if you require accommodations during the interview process.\nWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.","company":"Egen Solutions","rawCompany":"egen solutions","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T17:12:18.719Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead Machine Learning Engineer, Inference & Performance","description":"About Egen:\nEgen is a fast-growing and entrepreneurial company with a data-first mindset. We bring together the best engineering talent working with the most advanced technology platforms, including Google Cloud and Salesforce, to help clients drive action and impact through data and insights. We are committed to being a place where the best people choose to work so they can apply their engineering and technology expertise to envision what is next for how data and platforms can change the world for the better. We are dedicated to learning, thrive on solving tough problems, and continually innovate to achieve fast, effective results. If this describes you, we want you on our team.\nWant to learn more about life at Egen? Check out these resources in addition to the job description.\nMeet Egen\nLife at Egen\nCulture and Values at Egen\nCareer Development at Egen\nBenefits at Egen\nAbout the opportunity:\nAs a Senior AI Engineer, you will be at the forefront of our Generative AI initiatives. We treat AI as a software engineering discipline. You will be responsible for the full lifecycle of our AI features—specifically document intelligence and RAG pipelines—taking them from initial prototype to robust, scalable production services. You will solve for real-world constraints like latency, error handling, and cost optimization.\nYou’ll collaborate with a diverse range of clients to translate business needs into high-performance AI architectures. This role requires a blend of deep technical expertise in LLMs and a disciplined Software Engineering approach to ensure our solutions are robust, ethical, and scalable.\nWhat You Will Do:\nOptimize Inference: Build and tune production LLM serving with vLLM and SGLang—maximizing throughput and minimizing latency through batching, paged attention, quantization, and KV-cache strategies\nProfile & Accelerate Training: Instrument and profile training runs to find bottlenecks, then resolve them with the right attention implementations (e.g. FlashAttention) tuned to the underlying hardware (H200, GB200)\nEngineer for the Hardware: Apply a working understanding of GPU architecture and attention internals to choose the right approach per accelerator, rather than relying on defaults\nServe at Scale: Deploy and operate multiple models within shared GPU clusters on GKE, with autoscaling, efficient bin-packing, and graceful handling of mixed workloads\nDrive Efficiency: Own GPU utilization as a first-class metric—measure it, improve throughput-per-dollar, and continuously raise the ceiling on what our fleet can deliver\nCollaborate & Consult: Work directly with clients to understand performance, latency, and cost requirements, and translate them into pragmatic serving and training architectures\nYour Technical Toolkit:\nCore Languages: Mastery of Python and shell scripting; comfort reading and reasoning about lower-level (CUDA-adjacent) performance code is a strong plus\nInference Frameworks: Hands-on experience with vLLM, SGLash, or comparable high-performance serving stacks\nGPU & Model Internals: Solid grasp of GPU architecture, the fundamentals of LLM inference, and the attention mechanism—including where the bottlenecks live and how FlashAttention and similar techniques address them across hardware generations (H200, GB200)\nProfiling: Fluency with profiling tools to diagnose training and inference bottlenecks (compute-bound vs. memory-bound, kernel-level analysis)\nInfrastructure: Strong Kubernetes (GKE) experience—deploying and autoscaling multiple models on shared GPU clusters on Google Cloud\nMindset: A strong software engineering foundation—you write clean, maintainable code, measure before optimizing, and understand the full SDLC\nBasic Qualifications:\nBachelor's or Master's degree in Computer Science, Engineering, or a related technical field\n5+ years of experience in ML/AI engineering, with a meaningful portion focused on performance, infrastructure, or systems\nProven track record of deploying and optimizing models in a production environment\nDemonstrated experience profiling and improving GPU utilization for training and/or inference\nExperience with Classic Machine Learning (neural nets, training, tuning) is a strong plus\nKnowledge of Data Engineering and SQL\nPersonal Attributes:\nOwnership: You take pride in your work and see optimizations through from profile to production\nCuriosity: Hardware and serving frameworks change fast; you are a lifelong learner who stays ahead of the curve\nRigor: You measure before you optimize and let data, not intuition, guide where you spend effort\nConsultative Spirit: You enjoy interacting with clients and can translate technical complexity into business value\nEthics: You prioritize responsible AI development and data privacy\n$159,300 - $250,100 a year\nThis position may be hired at multiple levels. Final leveling is determined during the interview process based on a candidate's experience, skills, and interview outcomes, and the applicable salary range will align with the final level assigned.\n\nCompensation is determined based on factors including experience, expertise, interview performance, and geographic location, where applicable.\nCompensation & Benefits:\nThis role is eligible for our competitive salary and comprehensive benefits package to support your well-being:\nComprehensive Health Insurance\nPaid Leave (Vacation/PTO)\nPaid Holidays\nSick Leave\nParental Leave\nBereavement Leave\n401 (k) Employer Match\nEmployee Referral Bonuses\nCheck out our complete list of benefits here - >https://egen.ai/people/#benefits\nImportant: All roles are subject to standard hiring verification practices, which may include background checks, employment verification, and other relevant checks.\nEEO and Accommodations:\nEgen is an equal opportunity employer and is committed to inclusion, diversity, and equity in the workplace. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veterans’ status, or any other characteristic protected by federal, state, or local laws. Egen will also consider qualified applications with criminal histories, consistent with legal requirements. Egen welcomes and encourages applications from individuals with disabilities. Reasonable accommodations are available for candidates during all aspects of the selection process. Please advise the talent acquisition team if you require accommodations during the interview process.\nWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.","datePosted":"2026-08-03T17:12:18.719Z","dateModified":"2026-08-03T17:12:18.719Z","hiringOrganization":{"@type":"Organization","name":"Egen Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6a2b84ba2d0b55b833ae88ee"},"url":"https://jobsearcher.com/jobs/6a2b84ba2d0b55b833ae88ee"}}