LLM Inference & Optimization Engineer
Together AI is building scalable AI inference infrastructure to support large language and vision models. You will design and optimize distributed inference engines, focus on low latency, high throughput, and co-design with hardware teams to accelerate GPU/accelerator performance.
The role emphasizes research-driven development, collaboration across teams, and delivering end-to-end model serving pipelines in a fast-paced startup environment in San Francisco.
#J-18808-Ljbffr