JOBSEARCHER

Real-Time GPU Optimization Engineer - Inference

Techire AiMillbrae, CAL4 MidSeptember 11th, 2026
techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers. You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available. #J-18808-Ljbffr