{"schemaVersion":"jobsearcher.job.v1","id":"a80417dc9d022ff411da036c","url":"https://jobsearcher.com/jobs/a80417dc9d022ff411da036c","canonicalUrl":"https://jobsearcher.com/jobs/a80417dc9d022ff411da036c","title":"Global Inference Library Engineer","description":"Experience: Senior Level\nSalary: $175,000 - $250,000 per year\n\nJob Details\n-\n\nWe’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.\n\nWhat We’re Looking For\n\nStrong experience building AI/ML infrastructure, inference systems, or high-performance computing software\nStrong programming experience with Python and C++, Rust, or similar systems languages\nExperience with LLM inference frameworks and model-serving infrastructure\nHands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies\nExperience developing, integrating, or optimizing performance-critical compute kernels\nUnderstanding of modern transformer and LLM architectures\nFamiliarity with inference concepts including batching, attention, KV caching, quantization, and memory management\nExperience benchmarking and profiling AI workloads across different hardware environments\nStrong understanding of GPU or accelerator architecture and performance characteristics\nExperience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable\n\nA bit about us:\n-\n\nWe're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.\n\nWhy join us?\n-\n\nWell-funded by leading tech investors\nCutting edge technical problems with complex solutions\nLucrative Equity in a seed stage startup\nCompetitive compensation\nExcellent benefits (healthcare, vision, dental)\n\n#techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3","company":"Leoforce","rawCompany":"leoforce","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-22T15:56:27.822Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Global Inference Library Engineer","description":"Experience: Senior Level\nSalary: $175,000 - $250,000 per year\n\nJob Details\n-\n\nWe’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.\n\nWhat We’re Looking For\n\nStrong experience building AI/ML infrastructure, inference systems, or high-performance computing software\nStrong programming experience with Python and C++, Rust, or similar systems languages\nExperience with LLM inference frameworks and model-serving infrastructure\nHands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies\nExperience developing, integrating, or optimizing performance-critical compute kernels\nUnderstanding of modern transformer and LLM architectures\nFamiliarity with inference concepts including batching, attention, KV caching, quantization, and memory management\nExperience benchmarking and profiling AI workloads across different hardware environments\nStrong understanding of GPU or accelerator architecture and performance characteristics\nExperience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable\n\nA bit about us:\n-\n\nWe're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.\n\nWhy join us?\n-\n\nWell-funded by leading tech investors\nCutting edge technical problems with complex solutions\nLucrative Equity in a seed stage startup\nCompetitive compensation\nExcellent benefits (healthcare, vision, dental)\n\n#techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3","datePosted":"2026-08-22T15:56:27.822Z","dateModified":"2026-08-22T15:56:27.822Z","hiringOrganization":{"@type":"Organization","name":"Leoforce","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a80417dc9d022ff411da036c"},"url":"https://jobsearcher.com/jobs/a80417dc9d022ff411da036c"}}