{"schemaVersion":"jobsearcher.job.v1","id":"8e049d43a8fc2c7b1c6945ca","url":"https://jobsearcher.com/jobs/8e049d43a8fc2c7b1c6945ca","canonicalUrl":"https://jobsearcher.com/jobs/8e049d43a8fc2c7b1c6945ca","title":"Machine Learning Engineer (Inference)","description":"Machine Learning Engineer (Inference)San Francisco, On-Site$200,000-$300,000 + equityWhy this roleEarly-stage infra company building a next-gen AI cloud (neocloud) — rethinking how models run across heterogeneous hardware.You’ll own the layer that actually executes models in production.🧠 What you’ll doBuild end-to-end inference systems (request → runtime → response)Optimise for latency, throughput, and concurrency under real loadDesign batching, scheduling, and queuing systemsManage KV cache + memory at scaleDebug performance across model → runtime → hardware⚙️ The fun technical bitsDeep dives into LLM inference (prefill, decode, attention)Solving tail latency + throughput trade-offsWorking across systems, ML, and hardware layersOptimising across GPUs + next-gen acceleratorsHands-on with vLLM, TensorRT-LLM, or custom runtimes🎯 What they wantExperience with ML inference / model serving systemsStrong systems or backend engineering fundamentalsComfortable with performance, memory, and scaling challengesPython + C++","company":"Acceler8 Talent","rawCompany":"acceler8 talent","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-05-03T02:52:04.171Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer (Inference)","description":"Machine Learning Engineer (Inference)San Francisco, On-Site$200,000-$300,000 + equityWhy this roleEarly-stage infra company building a next-gen AI cloud (neocloud) — rethinking how models run across heterogeneous hardware.You’ll own the layer that actually executes models in production.🧠 What you’ll doBuild end-to-end inference systems (request → runtime → response)Optimise for latency, throughput, and concurrency under real loadDesign batching, scheduling, and queuing systemsManage KV cache + memory at scaleDebug performance across model → runtime → hardware⚙️ The fun technical bitsDeep dives into LLM inference (prefill, decode, attention)Solving tail latency + throughput trade-offsWorking across systems, ML, and hardware layersOptimising across GPUs + next-gen acceleratorsHands-on with vLLM, TensorRT-LLM, or custom runtimes🎯 What they wantExperience with ML inference / model serving systemsStrong systems or backend engineering fundamentalsComfortable with performance, memory, and scaling challengesPython + C++","datePosted":"2026-05-03T02:52:04.171Z","dateModified":"2026-05-03T02:52:04.171Z","hiringOrganization":{"@type":"Organization","name":"Acceler8 Talent","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8e049d43a8fc2c7b1c6945ca"},"url":"https://jobsearcher.com/jobs/8e049d43a8fc2c7b1c6945ca"}}