{"schemaVersion":"jobsearcher.job.v1","id":"8ae23d7d8f81a54e1bb00c46","url":"https://jobsearcher.com/jobs/8ae23d7d8f81a54e1bb00c46","canonicalUrl":"https://jobsearcher.com/jobs/8ae23d7d8f81a54e1bb00c46","title":"Kernel Engineer","description":"Member of Technical Staff - Kernels & GPU Performance\r\nEmployment Type:Full-time\r\nWork Model:On-site\r\nAbout the Company\r\nA fast-growing AI infrastructure company is building a next-generation compute platform designed for high-performance, efficient AI inference across heterogeneous hardware.\r\nThe platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.\r\nThe team works with leading AI organizations on technical challenges spanning frontier models, production infrastructure, and emerging accelerator architectures.\r\nAbout the Role\r\nAs a Member of Technical Staff, you will build and optimize the low-level execution primitives that translate accelerator capability into production inference performance.\r\nRather than optimizing for a single hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks.\r\nYour work will directly impact latency, throughput, hardware utilization, and efficiency across both established and emerging accelerator architectures.\r\nYou will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation, while partnering with compiler, ML systems, runtime, and distributed-systems engineers.\r\nWhat You’ll Do\r\nBuild and optimize kernels for production AI workloads.\r\nImprove latency, throughput, and hardware utilization.\r\nDevelop execution strategies across multiple accelerator architectures.\r\nOptimize memory efficiency, scheduling behavior, and low-level execution characteristics.\r\nAnalyze accelerator execution models and memory hierarchies to identify bottlenecks.\r\nProfile and validate performance across different hardware platforms.\r\nPartner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.\r\nDevelop optimization approaches that account for architectural differences between accelerators.\r\nHelp establish performance-engineering standards and best practices across the execution platform.\r\nInfluence how heterogeneous compute hardware is deployed and utilized in production AI infrastructure.\r\nWhat We’re Looking For\r\nStrong software-engineering fundamentals.\r\nExperience developing performance-critical systems close to hardware.\r\nStrong understanding of low-level execution behavior.\r\nAbility to reason about memory hierarchies, compute utilization, scheduling, and performance tradeoffs.\r\nExperience profiling, debugging, and optimizing systems for latency and throughput.\r\nBachelor’s degree in a relevant technical discipline or equivalent practical experience.\r\nNice to Have\r\nCUDA\r\nTriton\r\nCUTLASS\r\nAccelerator programming models\r\nGPU execution models including warps, wavefronts, blocks, and grids\r\nMemory-access optimization and coalescing\r\nShared-memory optimization\r\nCache optimization\r\nOccupancy tuning\r\nGPU profiling and performance-analysis tools\r\nMulti-GPU execution\r\nDistributed execution\r\nAI inference optimization\r\nHeterogeneous accelerator experience\r\nKeywords:\r\nGPU Kernels, CUDA, Triton, CUTLASS, Kernel Optimization, GPU Performance, Performance Engineering, Accelerator Programming, AI Inference, Inference Optimization, Low-Level Systems, GPU Architecture, Memory Hierarchy, Memory Coalescing, Shared Memory, Cache Optimization, Occupancy, Latency Hiding, Instruction-Level Parallelism, ILP, Warp, Wavefront, Thread Block, Grid, Profiling, Nsight, Performance Analysis, Throughput Optimization, Latency Optimization, Hardware Utilization, Multi-GPU, Distributed Systems, Heterogeneous Compute, Accelerator Runtime, ML Systems, Compiler Runtime.#J-18808-Ljbffr","company":"Acceler8 Talent","rawCompany":"acceler8 talent","city":"Millbrae","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-10-04T02:45:43.190Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Kernel Engineer","description":"Member of Technical Staff - Kernels & GPU Performance\r\nEmployment Type:Full-time\r\nWork Model:On-site\r\nAbout the Company\r\nA fast-growing AI infrastructure company is building a next-generation compute platform designed for high-performance, efficient AI inference across heterogeneous hardware.\r\nThe platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.\r\nThe team works with leading AI organizations on technical challenges spanning frontier models, production infrastructure, and emerging accelerator architectures.\r\nAbout the Role\r\nAs a Member of Technical Staff, you will build and optimize the low-level execution primitives that translate accelerator capability into production inference performance.\r\nRather than optimizing for a single hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks.\r\nYour work will directly impact latency, throughput, hardware utilization, and efficiency across both established and emerging accelerator architectures.\r\nYou will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation, while partnering with compiler, ML systems, runtime, and distributed-systems engineers.\r\nWhat You’ll Do\r\nBuild and optimize kernels for production AI workloads.\r\nImprove latency, throughput, and hardware utilization.\r\nDevelop execution strategies across multiple accelerator architectures.\r\nOptimize memory efficiency, scheduling behavior, and low-level execution characteristics.\r\nAnalyze accelerator execution models and memory hierarchies to identify bottlenecks.\r\nProfile and validate performance across different hardware platforms.\r\nPartner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.\r\nDevelop optimization approaches that account for architectural differences between accelerators.\r\nHelp establish performance-engineering standards and best practices across the execution platform.\r\nInfluence how heterogeneous compute hardware is deployed and utilized in production AI infrastructure.\r\nWhat We’re Looking For\r\nStrong software-engineering fundamentals.\r\nExperience developing performance-critical systems close to hardware.\r\nStrong understanding of low-level execution behavior.\r\nAbility to reason about memory hierarchies, compute utilization, scheduling, and performance tradeoffs.\r\nExperience profiling, debugging, and optimizing systems for latency and throughput.\r\nBachelor’s degree in a relevant technical discipline or equivalent practical experience.\r\nNice to Have\r\nCUDA\r\nTriton\r\nCUTLASS\r\nAccelerator programming models\r\nGPU execution models including warps, wavefronts, blocks, and grids\r\nMemory-access optimization and coalescing\r\nShared-memory optimization\r\nCache optimization\r\nOccupancy tuning\r\nGPU profiling and performance-analysis tools\r\nMulti-GPU execution\r\nDistributed execution\r\nAI inference optimization\r\nHeterogeneous accelerator experience\r\nKeywords:\r\nGPU Kernels, CUDA, Triton, CUTLASS, Kernel Optimization, GPU Performance, Performance Engineering, Accelerator Programming, AI Inference, Inference Optimization, Low-Level Systems, GPU Architecture, Memory Hierarchy, Memory Coalescing, Shared Memory, Cache Optimization, Occupancy, Latency Hiding, Instruction-Level Parallelism, ILP, Warp, Wavefront, Thread Block, Grid, Profiling, Nsight, Performance Analysis, Throughput Optimization, Latency Optimization, Hardware Utilization, Multi-GPU, Distributed Systems, Heterogeneous Compute, Accelerator Runtime, ML Systems, Compiler Runtime.#J-18808-Ljbffr","datePosted":"2026-10-04T02:45:43.190Z","dateModified":"2026-10-04T02:45:43.190Z","hiringOrganization":{"@type":"Organization","name":"Acceler8 Talent","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8ae23d7d8f81a54e1bb00c46"},"url":"https://jobsearcher.com/jobs/8ae23d7d8f81a54e1bb00c46"}}