{"schemaVersion":"jobsearcher.job.v1","id":"8c2a54063e3f156825c06c54","url":"https://jobsearcher.com/jobs/8c2a54063e3f156825c06c54","canonicalUrl":"https://jobsearcher.com/jobs/8c2a54063e3f156825c06c54","title":"Research Scientist, Performance Engineering","description":"TBC is building next-generation AI systems at the intersection of biological computing, generative models, and large-scale AI infrastructure. As we scale our world-model and neural-optimizer efforts, we are looking for an optimization-focused Research Scientist / ML Engineer to improve the efficiency, latency, throughput, and deployability of large models.\n\nThis role is focused on making frontier models run faster, cheaper, and more reliably — especially LLMs, diffusion models, video generation models, and world-model systems. You will work across inference optimization, training efficiency, model compression, memory management, and GPU-level performance to help turn research systems into scalable, customer-ready products.\n\nWhat You’ll Work On\n\nOptimize inference for LLMs, diffusion models, video models, and world-model systems\n\nImprove serving efficiency through techniques such as KV caching, batching, quantization, distillation, speculative decoding, and memory optimization\n\nBuild and optimize high-throughput inference pipelines for large models running on GPU clusters\n\nProfile model performance across latency, throughput, memory usage, GPU utilization, and cost\n\nImplement custom kernels or low-level optimizations using Triton, CUDA, PyTorch, or related systems\n\nImprove training and fine-tuning efficiency for large generative models, including distributed training, checkpointing, parallelism, and data loading\n\nWork with research teams to identify bottlenecks in model architecture, inference paths, and deployment workflows\n\nTranslate model performance improvements into clear customer-facing benchmarks and technical proof points\n\nEvaluate trade-offs across model quality, latency, cost, memory, and deployability\n\nWhat We’re Looking For\n\nStrong background in machine learning systems, model optimization, or high-performance AI infrastructure\n\nHands‑on experience optimizing LLMs, diffusion models, video generation models, or other large generative systems\n\nExperience with one or more of:\n\nInference optimization\n\nKV caching / attention optimization\n\nTriton or CUDA kernel development\n\nQuantization, pruning, distillation, or model compression\n\nDistributed training / fine-tuning efficiency\n\nGPU profiling and performance debugging\n\nStrong PyTorch experience and comfort working close to the model/runtime boundary\n\nAbility to reason about trade‑offs between quality, latency, throughput, memory, and cost\n\nComfortable working across research code, production systems, and benchmarking infrastructure\n\nExcited to work in an ambiguous, early-stage environment where optimization work directly shapes product feasibility\n\nWhat Success Looks Like\n\nLarge models run faster, cheaper, and more reliably across TBC’s core workloads\n\nInference pipelines show measurable improvements in latency, throughput, memory use, and GPU utilization\n\nTraining and fine‑tuning workflows become more efficient, reproducible, and scalable\n\nOptimization work translates into clear product and customer value, not just internal benchmarks\n\nResearch prototypes become deployable systems that can support demos, evaluations, and early partner use cases\n\nPreferred Qualifications\n\nPhD, MS, or equivalent industry experience in Computer Science, Machine Learning, Systems, Robotics, or related field\n\nPrior work optimizing large-scale generative models in production or research settings\n\nExperience with modern inference/training stacks such as PyTorch, Triton, CUDA, vLLM, TensorRT, DeepSpeed, FSDP, Ray, or similar tooling\n\nExperience working with LLMs, diffusion models, video generation models, or world models\n\n#J-18808-Ljbffr","company":"Biological Computing","rawCompany":"biological computing","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-26T03:18:57.086Z","occupations":[{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541714","title":"Research and Development in Biotechnology (except Nanobiotechnology)","slug":"research-and-development-in-biotechnology-except-nanobiotechnology"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Research Scientist, Performance Engineering","description":"TBC is building next-generation AI systems at the intersection of biological computing, generative models, and large-scale AI infrastructure. As we scale our world-model and neural-optimizer efforts, we are looking for an optimization-focused Research Scientist / ML Engineer to improve the efficiency, latency, throughput, and deployability of large models.\n\nThis role is focused on making frontier models run faster, cheaper, and more reliably — especially LLMs, diffusion models, video generation models, and world-model systems. You will work across inference optimization, training efficiency, model compression, memory management, and GPU-level performance to help turn research systems into scalable, customer-ready products.\n\nWhat You’ll Work On\n\nOptimize inference for LLMs, diffusion models, video models, and world-model systems\n\nImprove serving efficiency through techniques such as KV caching, batching, quantization, distillation, speculative decoding, and memory optimization\n\nBuild and optimize high-throughput inference pipelines for large models running on GPU clusters\n\nProfile model performance across latency, throughput, memory usage, GPU utilization, and cost\n\nImplement custom kernels or low-level optimizations using Triton, CUDA, PyTorch, or related systems\n\nImprove training and fine-tuning efficiency for large generative models, including distributed training, checkpointing, parallelism, and data loading\n\nWork with research teams to identify bottlenecks in model architecture, inference paths, and deployment workflows\n\nTranslate model performance improvements into clear customer-facing benchmarks and technical proof points\n\nEvaluate trade-offs across model quality, latency, cost, memory, and deployability\n\nWhat We’re Looking For\n\nStrong background in machine learning systems, model optimization, or high-performance AI infrastructure\n\nHands‑on experience optimizing LLMs, diffusion models, video generation models, or other large generative systems\n\nExperience with one or more of:\n\nInference optimization\n\nKV caching / attention optimization\n\nTriton or CUDA kernel development\n\nQuantization, pruning, distillation, or model compression\n\nDistributed training / fine-tuning efficiency\n\nGPU profiling and performance debugging\n\nStrong PyTorch experience and comfort working close to the model/runtime boundary\n\nAbility to reason about trade‑offs between quality, latency, throughput, memory, and cost\n\nComfortable working across research code, production systems, and benchmarking infrastructure\n\nExcited to work in an ambiguous, early-stage environment where optimization work directly shapes product feasibility\n\nWhat Success Looks Like\n\nLarge models run faster, cheaper, and more reliably across TBC’s core workloads\n\nInference pipelines show measurable improvements in latency, throughput, memory use, and GPU utilization\n\nTraining and fine‑tuning workflows become more efficient, reproducible, and scalable\n\nOptimization work translates into clear product and customer value, not just internal benchmarks\n\nResearch prototypes become deployable systems that can support demos, evaluations, and early partner use cases\n\nPreferred Qualifications\n\nPhD, MS, or equivalent industry experience in Computer Science, Machine Learning, Systems, Robotics, or related field\n\nPrior work optimizing large-scale generative models in production or research settings\n\nExperience with modern inference/training stacks such as PyTorch, Triton, CUDA, vLLM, TensorRT, DeepSpeed, FSDP, Ray, or similar tooling\n\nExperience working with LLMs, diffusion models, video generation models, or world models\n\n#J-18808-Ljbffr","datePosted":"2026-07-26T03:18:57.086Z","dateModified":"2026-07-26T03:18:57.086Z","hiringOrganization":{"@type":"Organization","name":"Biological Computing","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8c2a54063e3f156825c06c54"},"url":"https://jobsearcher.com/jobs/8c2a54063e3f156825c06c54"}}