JOBSEARCHER

Staff Engineer, High Performance Data & Algorithm Infrastructure

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

High Performance Data & Algorithm Infrastructure EngineerLocation: San Diego, CAJob Type: Full-Time$175k - $185k, bonus, equityPosition OverviewWe are a startup building performance-critical systems that push large volumes of data through tightly optimized compute pipelines. We are looking for a Senior Staff Software Engineer with deep expertise in high-performance computing (HPC), Linux systems, and GPU-accelerated data pipelines.This is a highly technical, hands-on role focused on extracting maximum performance from modern CPUs, GPUs, memory subsystems, and high-speed networks. You will work close to the hardware and operating system, tuning kernels, BIOS settings, and drivers, while also designing and implementing low-latency data processing pipelines that include real-time signal processing.If you enjoy profiling, tuning, and eliminating bottlenecks across the full stack—from BIOS to CUDA kernels to network offload—this role is for you.Key ResponsibilitiesHigh-Performance System EngineeringDesign, build, and optimize high-throughput, low-latency compute pipelinesProfile and tune performance across CPUs, GPUs, memory, storage, and networkingIdentify and eliminate bottlenecks in data movement and computationWork directly with hardware and OS configuration to achieve deterministic, repeatable performanceLinux Systems & Kernel ExpertiseConfigure and tune Linux systems for high-performance workloadsCustomize and tune Linux kernel parameters (scheduler, NUMA, IRQs, huge pages, IOMMU, etc.)Tune CPU and BIOS parameters (power states, frequency scaling, SMT, NUMA, memory timing)Manage and optimize DMA paths between devices and system memoryMinimize context switches, cache misses, and system jitterGPU & CUDA Programming (Critical)Develop and optimize GPU-accelerated compute pipelines using CUDAOptimize memory transfers between host and GPU (pinned memory, zero-copy, GPUDirect where applicable)Tune kernel launches, memory access patterns, and occupancyConfigure and manage GPU drivers, runtime, and system-level settings for maximum throughputProfile GPU workloads using tools such as Nsight Systems and Nsight ComputeData Movement & NetworkingOptimize high-speed data ingestion and offload to HPC systemsWork with low-latency and high-bandwidth networking technologies (e.g., RDMA, InfiniBand, high-speed Ethernet)Minimize data transfer latencies across network, PCIe, and memory boundariesDesign zero-copy or near-zero-copy data paths where possibleSignal Processing & AlgorithmsImplement and optimize digital signal processing algorithms, including:FFTsDeconvolutionThresholding and detection algorithmsOptimize DSP workloads for CPU vectorization and GPU accelerationBalance numerical accuracy, latency, and throughput constraintsQualificationsEducation:BS/MS Computer Science or EngineeringRequired:Experience & Technical Skills7+ years of professional software engineering experience (or equivalent depth)Strong background in high-performance computing or performance-critical systemsExpert-level Linux experience, including kernel and system tuningDeep experience with GPU computing and CUDA (required)Strong systems programming skills in C/C++ (and/or Rust)Solid understanding of computer architecture:CPU caches, NUMA, memory hierarchiesPCIe and DMAGPU architecturesPerformance & Debugging SkillsExtensive experience profiling and tuning complex systemsComfortable using tools such as perf, ftrace, eBPF, valgrind, Nsight, and similarAbility to reason quantitatively about latency, bandwidth, and throughputDSP & Mathematical FoundationsPractical experience implementing DSP algorithms in production systemsStrong understanding of FFTs, convolution/deconvolution, filtering, and thresholdingAbility to optimize numerical algorithms for real-time or near-real-time constraintsPreferred:Experience with RDMA, GPUDirect RDMA, or other hardware offload technologiesExperience with custom kernel builds or kernel module developmentFamiliarity with real-time or low-latency Linux variantsExperience deploying HPC workloads at scaleBackground in scientific computing, signal processing, or computational physicsWhat Success Looks LikeData pipelines consistently hit performance targets with headroomLatency and throughput are predictable, measurable, and well understoodGPUs and CPUs are efficiently utilized with minimal idle timeSystem-level bottlenecks are identified early and resolved decisivelyWhy This Role is InterestingYou will work on problems where performance truly mattersYou will operate across the full stack, from BIOS and kernel settings to CUDA kernels and DSP algorithmsYour optimizations will have immediate, measurable impactYou will have the freedom to deeply understand and tune the system, not just work around itWhy Join UsInfluence the foundational technologies and strategies of a company poised to shape the future of clinical genomics and healthcare. Work in a dynamic, collaborative environment where innovation and scientific rigor are deeply valued. Join a seasoned and multidisciplinary team tackling high-impact problems at the intersection of science and engineering. Competitive compensation and equity package, comprehensive benefits, and flexibility to support work-life integration. If you enjoy building orchestration layers for real machines and solving hard distributed real-time problems, this role offers significant technical ownership and impact.We are an equal opportunity employer. We thrive on diversity and collaboration.Compensation Range: $175K - $185K