JOBSEARCHER

Inference Systems Engineer GPU Kernels & LLMs

Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the kernel level, optimize perf, and design parallelism across heterogeneous hardware to maximize throughput in production. You will work with state-of-the-art engines, profiling tools, and advanced GPU techniques while collaborating with a hands-on team in a fast-paced SF office environment. #J-18808-Ljbffr