JOBSEARCHER

Global Inference Library Engineer

Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details - We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments. What We’re Looking For Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software Strong programming experience with Python and C++, Rust, or similar systems languages Experience with LLM inference frameworks and model-serving infrastructure Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies Experience developing, integrating, or optimizing performance-critical compute kernels Understanding of modern transformer and LLM architectures Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management Experience benchmarking and profiling AI workloads across different hardware environments Strong understanding of GPU or accelerator architecture and performance characteristics Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable A bit about us: - We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures. Why join us? - Well-funded by leading tech investors Cutting edge technical problems with complex solutions Lucrative Equity in a seed stage startup Competitive compensation Excellent benefits (healthcare, vision, dental) #techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3