JOBSEARCHER

Machine Learning Engineer, Inference Infrastructure

About the RoleWe are seeking a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide. ResponsibilitiesArchitect, build, and scale low-latency distributed inference serving systems for massive generative models. Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures. Collaborate with research teams to ensure smooth transition of new model architectures into production. Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures. RequirementsBS, MS, or Ph.D. in Computer Science or related technical field. 3+ years of industry experience building large-scale distributed systems or ML infrastructure. Deep proficiency in C++ and Python. Extensive experience with CUDA, Triton, or deep learning hardware accelerators. Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron). BenefitsTop-tier compensation including equity. Full medical, dental, and vision coverage with zero employee contribution options. Unlimited paid time off and flexible hybrid work policy. Catered daily lunches and wellness stipends.#J-18808-Ljbffr