{"schemaVersion":"jobsearcher.job.v1","id":"0ea2984252ec4262e8c28beb","url":"https://jobsearcher.com/jobs/0ea2984252ec4262e8c28beb","canonicalUrl":"https://jobsearcher.com/jobs/0ea2984252ec4262e8c28beb","title":"Machine Learning Engineer","description":"Location: Bay area (frequent customer interaction)\r\nTeam: Inference & Reinforcement Learning Platform\r\nAbout the Role\r\nWe're looking for a Machine Learning Engineer (MLE) to work directly with customers and partners to design, deploy, and validate inference and reinforcement learning (RL) proof-of-concepts on GMI's GPU infrastructure.\r\nThis is a high-impact, hybrid engineering role that sits at the intersection of platform engineering, applied ML, and customer success. You'll be embedded with customers during early-stage deployments—turning research ideas, datasets, and business requirements into working, performant systems on real GPU clusters.\r\nIf you enjoy being close to users, debugging real systems, and shipping results fast (not just writing docs), this role is for you.\r\nWhat You'll Do\r\nOwn customer POCs end-to-end\r\nDeploy and optimize LLM inference , RL training , and post-training workflows on GMI clusters\r\nTranslate customer requirements into concrete system designs and experiments\r\nForward-deploy with customers\r\nWork hands-on with research teams, startups, and enterprise customers\r\nDebug performance, stability, and correctness issues in real environments\r\nStand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve–style architectures)\r\nOptimize latency, throughput, GPU utilization, and cost efficiency\r\nRL & post-training POCs\r\nSupport RLHF / RFT / SFT workflows using customer-provided datasets\r\nIntegrate SDKs, training APIs, and cluster resources to shorten \"idea ? experiment\" cycles\r\nPerformance & reliability\r\nDiagnose GPU, networking, and distributed system bottlenecks\r\nRun benchmarks, profiling, and stress tests on multi-GPU / multi-node setups\r\nFeedback loop to product\r\nFeed real-world customer learnings back into GMI's platform, SDKs, and APIs\r\nHelp shape reference architectures, cookbooks, and best practices\r\nWhat We're Looking For\r\nCore Requirements\r\nStrong software engineering background (Python required; Go / Rust a plus)\r\nHands-on experience with ML inference or training systems\r\nFamiliarity with distributed systems and GPUs (multi-GPU, multi-node)\r\nComfort working directly with customers and ambiguous requirements\r\nAbility to debug end-to-end systems (code, infra, networking, performance)\r\nNice to Have\r\nExperience with\r\nRL or post-training workflows (RLHF, RFT, SFT)\r\nPyTorch, DeepSpeed, Megatron-LM, or similar\r\nGPU performance profiling and optimization\r\nPrior experience as\r\nSolutions Engineer\r\nApplied Research Engineer\r\nWhat Makes This Role Special\r\nYou're close to real users and real GPUs —not abstract roadmaps\r\nYou'll work on cutting-edge inference and RL workloads , not toy demos\r\nYou'll influence product direction through direct customer feedback\r\nFast iteration, high ownership, and visible impact\r\nWho Thrives Here\r\nEngineers who like shipping over theorizing\r\nPeople who enjoy being the \"last mile\" problem solver\r\nBuilders who want exposure to both deep systems and applied ML\r\nThose excited by early-stage POCs that turn into real production systems\r\nJ-18808-Ljbffr","company":"Gmi Cloud","rawCompany":"gmi cloud","city":"Mountain View","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-04-09T09:08:27.626Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer","description":"Location: Bay area (frequent customer interaction)\r\nTeam: Inference & Reinforcement Learning Platform\r\nAbout the Role\r\nWe're looking for a Machine Learning Engineer (MLE) to work directly with customers and partners to design, deploy, and validate inference and reinforcement learning (RL) proof-of-concepts on GMI's GPU infrastructure.\r\nThis is a high-impact, hybrid engineering role that sits at the intersection of platform engineering, applied ML, and customer success. You'll be embedded with customers during early-stage deployments—turning research ideas, datasets, and business requirements into working, performant systems on real GPU clusters.\r\nIf you enjoy being close to users, debugging real systems, and shipping results fast (not just writing docs), this role is for you.\r\nWhat You'll Do\r\nOwn customer POCs end-to-end\r\nDeploy and optimize LLM inference , RL training , and post-training workflows on GMI clusters\r\nTranslate customer requirements into concrete system designs and experiments\r\nForward-deploy with customers\r\nWork hands-on with research teams, startups, and enterprise customers\r\nDebug performance, stability, and correctness issues in real environments\r\nStand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve–style architectures)\r\nOptimize latency, throughput, GPU utilization, and cost efficiency\r\nRL & post-training POCs\r\nSupport RLHF / RFT / SFT workflows using customer-provided datasets\r\nIntegrate SDKs, training APIs, and cluster resources to shorten \"idea ? experiment\" cycles\r\nPerformance & reliability\r\nDiagnose GPU, networking, and distributed system bottlenecks\r\nRun benchmarks, profiling, and stress tests on multi-GPU / multi-node setups\r\nFeedback loop to product\r\nFeed real-world customer learnings back into GMI's platform, SDKs, and APIs\r\nHelp shape reference architectures, cookbooks, and best practices\r\nWhat We're Looking For\r\nCore Requirements\r\nStrong software engineering background (Python required; Go / Rust a plus)\r\nHands-on experience with ML inference or training systems\r\nFamiliarity with distributed systems and GPUs (multi-GPU, multi-node)\r\nComfort working directly with customers and ambiguous requirements\r\nAbility to debug end-to-end systems (code, infra, networking, performance)\r\nNice to Have\r\nExperience with\r\nRL or post-training workflows (RLHF, RFT, SFT)\r\nPyTorch, DeepSpeed, Megatron-LM, or similar\r\nGPU performance profiling and optimization\r\nPrior experience as\r\nSolutions Engineer\r\nApplied Research Engineer\r\nWhat Makes This Role Special\r\nYou're close to real users and real GPUs —not abstract roadmaps\r\nYou'll work on cutting-edge inference and RL workloads , not toy demos\r\nYou'll influence product direction through direct customer feedback\r\nFast iteration, high ownership, and visible impact\r\nWho Thrives Here\r\nEngineers who like shipping over theorizing\r\nPeople who enjoy being the \"last mile\" problem solver\r\nBuilders who want exposure to both deep systems and applied ML\r\nThose excited by early-stage POCs that turn into real production systems\r\nJ-18808-Ljbffr","datePosted":"2026-04-09T09:08:27.626Z","dateModified":"2026-04-09T09:08:27.626Z","hiringOrganization":{"@type":"Organization","name":"Gmi Cloud","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Mountain View","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0ea2984252ec4262e8c28beb"},"url":"https://jobsearcher.com/jobs/0ea2984252ec4262e8c28beb"}}