{"schemaVersion":"jobsearcher.job.v1","id":"8d9d9e816f64b78e84232f2f","url":"https://jobsearcher.com/jobs/8d9d9e816f64b78e84232f2f","canonicalUrl":"https://jobsearcher.com/jobs/8d9d9e816f64b78e84232f2f","title":"RESEARCHER, EFFICIENT INFERENCE","description":"ABOUT THE COMPANY\nWe're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site\nABOUT THE ROLE\nYou'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run.\nThis is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).\nWHAT YOU'LL DO\nResearch and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic\nDesign and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding\nInvestigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning\nRun controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks\nCo-design methods with the inference engineering team: push results to where they actually run, not stop at the paper\nRead deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack\nPublish when the work warrants it; share findings internally\nPartner with model and training researchers so efficiency choices align with model architecture and post-training decisions\nWHAT WE'RE LOOKING FOR\nStrong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent\n5+ years of hands-on research experience\nDeep familiarity with both training and inference performance characteristics\nFluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving-framework level when methods require it\nTrack record of moving efficiency research from prototype to production\nStrong statistical expertise: you'd notice a flawed comparison before someone else points it out\nStrong written communication\nPublished research at NeurIPS, ICML, ICLR, MLSys, or comparable venues\nNICE TO HAVE\nPhD in ML, systems, or related field\nOpen-source contributions to quantization, speculative-decoding, or\nefficient-inference libraries\nExperience with hardware-aware optimization and accelerator-specific\ntooling\nBackground in numerical methods, low-precision arithmetic, or\napproximate computation\nTHIS ROLE IS PROBABLY NOT FOR YOU IF\nYou want to focus on pretraining large models from scratch (that's a different role)\nYou prefer abstract algorithmic research without hands-on implementation\nYou want a fixed benchmark with stable targets (our targets shift with what our models actually need to do)","company":"Makermaker","rawCompany":"makermaker","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-23T14:52:53.872Z","occupations":[{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541990","title":"All Other Professional, Scientific, and Technical Services","slug":"all-other-professional-scientific-and-technical-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"RESEARCHER, EFFICIENT INFERENCE","description":"ABOUT THE COMPANY\nWe're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site\nABOUT THE ROLE\nYou'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run.\nThis is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).\nWHAT YOU'LL DO\nResearch and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic\nDesign and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding\nInvestigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning\nRun controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks\nCo-design methods with the inference engineering team: push results to where they actually run, not stop at the paper\nRead deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack\nPublish when the work warrants it; share findings internally\nPartner with model and training researchers so efficiency choices align with model architecture and post-training decisions\nWHAT WE'RE LOOKING FOR\nStrong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent\n5+ years of hands-on research experience\nDeep familiarity with both training and inference performance characteristics\nFluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving-framework level when methods require it\nTrack record of moving efficiency research from prototype to production\nStrong statistical expertise: you'd notice a flawed comparison before someone else points it out\nStrong written communication\nPublished research at NeurIPS, ICML, ICLR, MLSys, or comparable venues\nNICE TO HAVE\nPhD in ML, systems, or related field\nOpen-source contributions to quantization, speculative-decoding, or\nefficient-inference libraries\nExperience with hardware-aware optimization and accelerator-specific\ntooling\nBackground in numerical methods, low-precision arithmetic, or\napproximate computation\nTHIS ROLE IS PROBABLY NOT FOR YOU IF\nYou want to focus on pretraining large models from scratch (that's a different role)\nYou prefer abstract algorithmic research without hands-on implementation\nYou want a fixed benchmark with stable targets (our targets shift with what our models actually need to do)","datePosted":"2026-07-23T14:52:53.872Z","dateModified":"2026-07-23T14:52:53.872Z","hiringOrganization":{"@type":"Organization","name":"Makermaker","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8d9d9e816f64b78e84232f2f"},"url":"https://jobsearcher.com/jobs/8d9d9e816f64b78e84232f2f"}}