Research Engineer
Research EngineerFull time · San Francisco or New York · On site 💰 Compensation: $400K to $500K base salary plus meaningful equity📍 Location: San Francisco or New York, on site At Dynamism, we partner with breakout startups backed by Sequoia, a16z, and Y Combinator to help them hire engineers with founder energy, people who build fast, think like owners, and thrive in zero to one environments.We're hiring for this role at an early stage applied AI company building high fidelity environments, tasks, data, and evaluation infrastructure that help frontier models learn reliable behaviour in real work. The company works closely with frontier AI labs and enterprises, and is backed by Sequoia Capital, Menlo Ventures, BCV, and SV Angel. Its engineering first team brings together former founders, researchers, and deeply technical builders who value speed with rigor, truth over comfort, and precision under pressure. 🎯 The RoleYou will own the infrastructure that research runs on, including a large GPU cluster, the training and inference stacks, the research codebase, and the kernels underneath it. The job is to find what is actually slow or fragile and fix it where the fix belongs, without breaking the science. 🧩 What You'll Own• Capacity, failover, fault tolerance, and reliability across the GPU cluster• Performance profiling across NCCL, KV cache, dataloaders, collectives, and rollout generation• Distributed training using FSDP, tensor parallelism, pipeline parallelism, and asynchronous workflows• CUDA and Triton kernels when the profiler shows that the bottleneck is low level• A fast, maintainable research codebase that can evolve without invalidating experiments ✅ You Might Be a Fit If You• Have hands-on experience with multi-node training across tens of GPUs or more• Can profile a slow training run and land the fix, including at the kernel level• Have built or maintained high-throughput inference using systems such as vLLM or SGLang• Understand RL post-training methods and can diagnose training failures from logs and metrics• Enjoy working closely with researchers in a small team where priorities change with the evidence 🛠 StackPython, PyTorch, CUDA, Triton, FSDP, NCCL, vLLM or SGLang, distributed training, GPU networking Bonus if you've:• Contributed to PyTorch, vLLM, SGLang, Megatron, DeepSpeed, flash-attention, or similar systems• Operated GPU clusters at scale• Worked with InfiniBand, RoCE, GPUDirect, or other high-performance networking• Built RL rollout systems, asynchronous training, or off-policy training infrastructure 📋 Role Details• Compensation: $400K to $500K base salary plus meaningful equity• Location: San Francisco or New York, on site• Full time and on site This role will not suit someone looking for routine, a narrow technical silo, or a management only position. The team moves quickly, expects direct ownership, and does not trade quality for speed.Apply via LinkedIn, or reach us at apply@sfdynamism.com.