Inference Performance Engineer: Latency & Cost Optimization
OpenAI is looking for a performance modeler in San Francisco who will analyze inference stack performance and build cost-to-serve estimates. In this role, candidates should have expertise in performance profiling and enjoy reasoning about distributed systems.
The position offers a compensation range of $295K to $555K and requires collaboration with engineering and research teams to enhance performance and address system bottlenecks.
#J-18808-Ljbffr