JOBSEARCHER

Senior GPU Kernel Engineer: Inference Throughput

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You will write and tune CUDA kernels, optimize tensor cores, and push end-to-end latency down while maintaining accuracy. You will lead benchmarking workflows (MLPerf), mentor engineers, and collaborate with cross-functional partners. Experience with CUDA, C++, Python, and GPU architectures is required; familiarity with vLLM, TensorRT-LLM, llm-d, SGLang is a #J-18808-Ljbffr