Performance Engineer: Optimize GPU Inference at Scale
Morph is hiring a performance engineer to make the entire system faster, cheaper, and more reliable. You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served, in a small team with enormous compute and immediate production impact.
You’ll trace latency and throughput from the API layer down to individual kernels, optimize batching, routing, quantization, and distributed execution, and build benchmarks and observability to make
#J-18808-Ljbffr