JOBSEARCHER

Senior ML Researcher: Efficient Inference & Quantization

MakerMaker in San Francisco seeks a senior research engineer to advance efficiency in ML models, focusing on quantization, speculative decoding, and efficient training-time techniques. The role blends model architecture with inference performance, delivering production-ready improvements. You will run large-scale experiments, co-design with inference engineers, and push findings to production while publishing where appropriate. Collaboration with researchers is essential. J-18808-Ljbffr