JOBSEARCHER

ML Engineer, Inference & Optimization

PikaSan Jose, CAL6 LeadSeptember 15th, 2026
Overview In this role you accelerate inference for Pika’s AI-powered video and language models, driving faster and more efficient real-time experiences. You will design and optimize inference pipelines, and apply cutting-edge acceleration techniques to boost model speed at scale. You’ll maximize GPU parallelism and develop high-performance kernels while collaborating with researchers and engineers to deploy videogen and LLMs. This position shapes the next generation of Pika’s creative AI products and requires an ownership mindset in a fast-paced startup. You will work closely with cross-functional teams to deliver measurable impact. Compensation / BenefitsCompetitive salaryEquity in a fast-growing startupComprehensive health benefitsMonthly stipendsCompany retreatsOn-site work in Palo Alto ResponsibilitiesLead and implement advanced inference acceleration techniques (attention optimization, quantization)Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP)Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCLCollaborate with research and engineering teams to deploy state-of-the-art videogen and large language modelsContribute to improvements in training speed, stability and resource utilization as part of deployment lifecycle (bonus)Drive rigorous code reviews and mentor engineers on inference and GPU programming Key requirements5+ years engineering experience in inference acceleration and model deployment at scaleExpertise in inference optimization (quantization, attention acceleration, deep learning compiler stacks)Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other parallelism forms for distributed inferenceFamiliarity with video generation (videogen) models and large language models (LLMs)Strong cross-discipline collaboration and communication skillsOwnership mindset; capable of navigating ambiguity in a startup environmentBonus: experience improving training efficiency, stability, or resource optimization for large modelscross-functional collaborationownership and initiativeclear communicationCUDANCCLquantization