LLM Inference Architect - Frameworks & Optimization
Together AI is building state-of-the-art infrastructure to enable scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, shaping future AI deployment across diverse applications. #J-18808-Ljbffr