Cloud Inference Engineer
QualificationsCUDA + GPU inference optimizationvLLM, SGLang, or TensorRT-LLM experienceKV caching, paged attention, batching, token streaming, etc.Distributed compute (with GPUs is a super plus)No degree requiredCompanyLuminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.RoleFounding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.Day To Day ResponsibilitiesDeploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.Conducting model performance reviewsImprove scheduler, batcher, autoscaling; profile latency, cost, utilizationSometimes write kernels and, yes, occasional tasteful shitposting