JOBSEARCHER

MTS Inference: GPU Kernel & Performance Architect

Sail Research in San Francisco is looking for a motivated software engineer to optimize token processing at every layer of the stack. You will modify inference engines and analyze GPU performance, ensuring efficient hardware utilization. The ideal candidate has a solid understanding of LLM mechanics and interests in cutting-edge MLSys research. The company offers benefits like free meals and a Studio Display for every employee. J-18808-Ljbffr