JOBSEARCHER

Senior Software Engineer — Backend Performance & Data Systems

Senior Software Engineer — Backend Performance & Data SystemsLocation: San Francisco, CA — On-siteEmployment Type: Full-timeCompensation: $210,000–$245,000 base salary, plus equity and benefitsNote: No C2C arrangements will be considered.Any attempt to use personal or household contact information for solicitation, candidate submission, or vendor outreach is strictly prohibited and will be reported to LinkedIn.About the CompanyOur client is a well-funded, high-growth technology company building an advanced enterprise data and AI platform. Its systems process large volumes of structured and unstructured data to help major organizations make high-impact business decisions.The engineering team is small, highly technical, and focused on solving complex infrastructure, data-processing, and performance challenges at significant scale.About the RoleWe are seeking a Senior Software Engineer to own the performance-critical paths within a sophisticated data-processing platform.This is an engineers-first, systems-heavy role for someone who profiles before guessing, understands where Python reaches its limits, and is comfortable moving into C, C++, Cython, Rust, or GPU acceleration when a workload demands it.You will improve the speed, throughput, latency, memory efficiency, and overall computational performance of large-scale data pipelines. You will work closely with product, data, and infrastructure engineers while independently owning difficult performance problems from initial profiling through production implementation.This is not a conventional Python application-development role. The right engineer will be comfortable reasoning about what is happening beneath the application layer, including memory allocation, cache behavior, concurrency, data movement, serialization, and hardware utilization.What You’ll DoProfile, benchmark, and eliminate bottlenecks across performance-critical data pipelines.Build high-performance components using C, C++, Rust, or Cython.Optimize Python workloads by identifying interpreter, memory, serialization, and data-movement overhead.Determine when workloads should remain in Python and when they should move into compiled or hardware-accelerated implementations.Apply GPU acceleration and CUDA where they produce measurable performance improvements.Engineer highly parallel systems using threading, multiprocessing, SIMD, vectorization, asynchronous execution, and GPU parallelism.Optimize memory layout, allocation patterns, cache utilization, concurrency, and data-transfer costs.Build automated benchmarks and performance-regression testing into CI/CD pipelines.Process and transform large volumes of structured and unstructured data with low latency and high throughput.Improve the performance of database-intensive and data-processing workloads involving PostgreSQL.Partner with infrastructure, product, and data engineers to select the appropriate execution model and technology for each workload.Independently investigate complex systems behavior and turn findings into durable production improvements.Document performance assumptions, benchmarks, tradeoffs, and architectural decisions.What You BringDeep hands-on systems engineering experience.Strong proficiency in at least one lower-level systems language, preferably:CC++RustAdvanced Python experience, particularly for data-processing and performance-sensitive workloads.Demonstrated experience profiling and benchmarking production software.Strong understanding of:Memory managementConcurrency and parallelismCache behaviorData movementThroughput optimizationLatency optimizationExperience building high-throughput or low-latency data-processing systems.Strong knowledge of data structures and algorithms.Production experience working with PostgreSQL.Ability to independently write, debug, and reason about performance-critical code.Experience identifying the actual source of performance problems rather than relying primarily on additional compute resources.Strong communication skills and the ability to explain technical tradeoffs to other engineers.Highly PreferredCUDA or GPU-accelerated computing experience.Cython experience.SIMD, vectorization, or other hardware-aware optimization experience.Experience with the NVIDIA GPU ecosystem.Apache Arrow, Parquet, or other columnar data formats.Lakehouse architecture or data-platform internals.Numerical, array-based, or scientific computing experience.Experience building automated performance-regression tests.Experience optimizing distributed or parallel data systems.Familiarity with high-performance computing environments.You’ll Thrive Here IfYou reach for a profiler before requesting more hardware.You can explain where processing time, memory, and data movement are being consumed.You enjoy writing lower-level systems code yourself.You are comfortable moving between Python and compiled languages.You prefer measuring performance improvements rather than relying on intuition alone.You are comfortable owning technically ambiguous problems without a predefined implementation plan.You thrive in a fast-moving environment with shifting priorities and meaningful technical ownership.You are a self-directed, curious, low-ego engineer.You can work effectively with a small group of highly experienced engineers.You care about production results rather than optimization for its own sake.This Role May Not Be the Right Fit IfYou prefer working exclusively in high-level Python application code.You want to avoid lower-level systems or performance work.Your primary approach to optimization is adding more compute resources.You have limited hands-on experience profiling or benchmarking production systems.You prefer delegating lower-level implementation rather than writing the code yourself.You rely heavily on AI coding tools without being able to independently reason about and validate the underlying systems.You prefer highly structured environments with fixed priorities and fully established processes.Why JoinWork on technically difficult, high-impact data and AI systems.Own critical architectural and performance decisions.Solve significant throughput, latency, memory, and compute-efficiency challenges.Work across Python, compiled systems languages, databases, and accelerated computing.Collaborate with a small team of highly experienced engineers.See the direct production impact of your engineering decisions.Receive a competitive base salary, meaningful equity, and comprehensive benefits.Work Authorization: Candidates must be currently authorized to work in the United States. This position is not eligible for new or future employer-sponsored work authorization.