{"schemaVersion":"jobsearcher.job.v1","id":"ed819c2108feeeb0303fcbcd","url":"https://jobsearcher.com/jobs/ed819c2108feeeb0303fcbcd","canonicalUrl":"https://jobsearcher.com/jobs/ed819c2108feeeb0303fcbcd","title":"Cloud Inference Engineer","description":"Qualifications CUDA + GPU inference optimization\r\nvLLM, SGLang, or TensorRT-LLM experience\r\nKV caching, paged attention, batching, token streaming, etc.\r\nDistributed compute (with GPUs is a super plus)\r\nNo degree required\r\nCompany Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.\r\nRole Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.\r\nDay To Day Responsibilities Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.\r\nConducting model performance reviews\r\nImprove scheduler, batcher, autoscaling; profile latency, cost, utilization\r\nSometimes write kernels\r\nJ-18808-Ljbffr","company":"SupportFinity","rawCompany":"supportfinity","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-15T01:03:55.031Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Cloud Inference Engineer","description":"Qualifications CUDA + GPU inference optimization\r\nvLLM, SGLang, or TensorRT-LLM experience\r\nKV caching, paged attention, batching, token streaming, etc.\r\nDistributed compute (with GPUs is a super plus)\r\nNo degree required\r\nCompany Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.\r\nRole Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.\r\nDay To Day Responsibilities Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.\r\nConducting model performance reviews\r\nImprove scheduler, batcher, autoscaling; profile latency, cost, utilization\r\nSometimes write kernels\r\nJ-18808-Ljbffr","datePosted":"2026-07-15T01:03:55.031Z","dateModified":"2026-07-15T01:03:55.031Z","hiringOrganization":{"@type":"Organization","name":"SupportFinity","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"ed819c2108feeeb0303fcbcd"},"url":"https://jobsearcher.com/jobs/ed819c2108feeeb0303fcbcd"}}