Global Inference Library Engineer
Experience: Senior Level
Salary: $175,000 - $250,000 per year
Job Details
-
We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.
What We’re Looking For
Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software
Strong programming experience with Python and C++, Rust, or similar systems languages
Experience with LLM inference frameworks and model-serving infrastructure
Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies
Experience developing, integrating, or optimizing performance-critical compute kernels
Understanding of modern transformer and LLM architectures
Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management
Experience benchmarking and profiling AI workloads across different hardware environments
Strong understanding of GPU or accelerator architecture and performance characteristics
Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable
A bit about us:
-
We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.
Why join us?
-
Well-funded by leading tech investors
Cutting edge technical problems with complex solutions
Lucrative Equity in a seed stage startup
Competitive compensation
Excellent benefits (healthcare, vision, dental)
#techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3