Staff ML Systems Engineer - Diffusion LLM Serving
Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable.
You will extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load balancing, autoscaling, and traffic routing for model endpoints.
J-18808-Ljbffr