JOBSEARCHER

Inference Software Engineer

About the CompanySeries A model-serving startup, ~35 people, $40M raised · San Francisco · Hybrid, 3 days · $180,000 – $280,000 + equity. They serve open-weight models for companies who can't send data to a third-party API. Their customers care about two numbers: p99 latency and cost per million tokens. This role owns both.About the RoleThe team runs a fork of vLLM across an H100 fleet. The last engineer to join cut p99 by 40% on the largest workload by reworking the scheduler. That’s the standard.ResponsibilitiesKernel-level optimisation — CUDA and Triton, custom ops, fused attentionContinuous batching, KV-cache strategy and speculative decoding, tuned per workload rather than globallyQuantisation trade-offs where the customer notices quality loss before you doMulti-GPU and multi-node serving, plus the profiling to prove any of it workedQualificationsBackgrounds that translate: HPC, graphics, embedded, compiler work, or performance engineering somewhere latency was a product requirement. Several of the strongest people in this space had never touched ML until it became a systems problem.Required SkillsFour engineers on the inference team today, going to eight this year.Pay range and compensation package$180,000 – $280,000 + equityEqual Opportunity StatementWe are committed to diversity and inclusivity.