{"schemaVersion":"jobsearcher.job.v1","id":"a28f66f399da7b202170765f","url":"https://jobsearcher.com/jobs/a28f66f399da7b202170765f","canonicalUrl":"https://jobsearcher.com/jobs/a28f66f399da7b202170765f","title":"GPU Optimization Engineer","description":"GPU Optimisation Engineer — Real-Time Inference\nWant to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads?\nThis team is building low-latency AI systems where milliseconds actually matter. The target isn’t “faster than baseline.” It’s sub-50ms time-to-first-token at 100+ concurrent requests on a single H100 — while maintaining model quality.\nThey’re hiring a GPU Optimisation Engineer who understands GPUs at an architectural level. Someone who knows where performance is really lost: memory hierarchy, kernel launch overhead, occupancy limits, scheduling inefficiencies, KV cache behaviour, attention paths. The work sits close to the metal, inside inference execution — not general infra, not model research.\nYou’ll operate across the kernel and runtime layers, profiling large-scale speech and multimodal models end-to-end and removing bottlenecks wherever they appear.\nWhat you’ll work on Profiling GPU bottlenecks across memory bandwidth, kernel fusion, quantisation, and scheduling\n\nWriting and tuning custom CUDA / Triton kernels for performance-critical paths\n\nImproving attention, decoding, and KV cache efficiency in inference runtimes\n\nModifying and extending vLLM-style systems to better suit real-time workloads\n\nOptimising models to fit GPU memory constraints without degrading output quality\n\nBenchmarking across NVIDIA GPUs (with exposure to AMD and other accelerators over time)\n\nPartnering directly with research to turn new model ideas into fast, production-ready inference\n\nThis is hands-on optimisation work across the stack. No layers of bureaucracy. No “platform ownership” theatre. Just deep performance engineering applied to models that are actively evolving.\nWhat tends to work well Strong experience with CUDA and/or Triton\n\nDeep understanding of GPU execution (memory hierarchy, scheduling, occupancy, concurrency)\n\nExperience optimising inference latency and throughput for large generative models\n\nFamiliarity with attention kernels, decoding paths, or LLM-style runtimes\n\nComfort profiling with low-level GPU tooling\n\nThe company is revenue-generating, its models are used by global enterprises, and the SF R&D team is expanding following a recent raise. This is growth hiring, not backfill.\nPackage & location Base salary: up to ~$300,000 (negotiable based on depth)\n\nEquity: Meaningful stock\n\nLocation: San Francisco preferred (relocation and visa sponsorship can be provided)\n\nIf you care about real-time constraints, GPU architecture, and squeezing every last millisecond out of large models, this is worth a conversation.\nAll applicants will receive a response.\n\n#J-18808-Ljbffr","company":"Trades Workforce Solutions","rawCompany":"trades workforce solutions","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-04-09T09:37:09.077Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"GPU Optimization Engineer","description":"GPU Optimisation Engineer — Real-Time Inference\nWant to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads?\nThis team is building low-latency AI systems where milliseconds actually matter. The target isn’t “faster than baseline.” It’s sub-50ms time-to-first-token at 100+ concurrent requests on a single H100 — while maintaining model quality.\nThey’re hiring a GPU Optimisation Engineer who understands GPUs at an architectural level. Someone who knows where performance is really lost: memory hierarchy, kernel launch overhead, occupancy limits, scheduling inefficiencies, KV cache behaviour, attention paths. The work sits close to the metal, inside inference execution — not general infra, not model research.\nYou’ll operate across the kernel and runtime layers, profiling large-scale speech and multimodal models end-to-end and removing bottlenecks wherever they appear.\nWhat you’ll work on Profiling GPU bottlenecks across memory bandwidth, kernel fusion, quantisation, and scheduling\n\nWriting and tuning custom CUDA / Triton kernels for performance-critical paths\n\nImproving attention, decoding, and KV cache efficiency in inference runtimes\n\nModifying and extending vLLM-style systems to better suit real-time workloads\n\nOptimising models to fit GPU memory constraints without degrading output quality\n\nBenchmarking across NVIDIA GPUs (with exposure to AMD and other accelerators over time)\n\nPartnering directly with research to turn new model ideas into fast, production-ready inference\n\nThis is hands-on optimisation work across the stack. No layers of bureaucracy. No “platform ownership” theatre. Just deep performance engineering applied to models that are actively evolving.\nWhat tends to work well Strong experience with CUDA and/or Triton\n\nDeep understanding of GPU execution (memory hierarchy, scheduling, occupancy, concurrency)\n\nExperience optimising inference latency and throughput for large generative models\n\nFamiliarity with attention kernels, decoding paths, or LLM-style runtimes\n\nComfort profiling with low-level GPU tooling\n\nThe company is revenue-generating, its models are used by global enterprises, and the SF R&D team is expanding following a recent raise. This is growth hiring, not backfill.\nPackage & location Base salary: up to ~$300,000 (negotiable based on depth)\n\nEquity: Meaningful stock\n\nLocation: San Francisco preferred (relocation and visa sponsorship can be provided)\n\nIf you care about real-time constraints, GPU architecture, and squeezing every last millisecond out of large models, this is worth a conversation.\nAll applicants will receive a response.\n\n#J-18808-Ljbffr","datePosted":"2026-04-09T09:37:09.077Z","dateModified":"2026-04-09T09:37:09.077Z","hiringOrganization":{"@type":"Organization","name":"Trades Workforce Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a28f66f399da7b202170765f"},"url":"https://jobsearcher.com/jobs/a28f66f399da7b202170765f"}}