{"schemaVersion":"jobsearcher.job.v1","id":"5143d40adb27d9a2e4a2c08c","url":"https://jobsearcher.com/jobs/5143d40adb27d9a2e4a2c08c","canonicalUrl":"https://jobsearcher.com/jobs/5143d40adb27d9a2e4a2c08c","title":"Lead AI Engineer","description":"Hybrid – 3 days onsite\r\nAbout the Role\r\nAn early-stage, AI-native product company is hiring a Lead AI Engineer to own the model layer of its core product. This is a senior individual contributor role focused on fine-tuning, model optimization, and custom small model development — not prompt engineering or research-only experimentation. This engineer will be responsible for designing, shipping, and iterating on production-grade LLM systems with real-time user impact.\r\nOwn fine-tuning strategy (LoRA, adapters, distillation, full fine-tuning)\r\nDecide when to fine-tune vs. use system-level or prompt-based approaches\r\nImprove models based on production feedback\r\nBalance quality, latency, and cost\r\nModel Performance & Optimization\r\nOptimize inference speed and throughput\r\nImprove reliability and consistency\r\nDefine evaluation frameworks and benchmarking standards\r\nCustom Small Model Development\r\nDesign and deploy custom small language models (SLMs)\r\nDetermine when smaller models outperform larger ones\r\nMaintain real-time performance for interactive UX workflows\r\nWhat You'll Build\r\nFine-tuned models powering generation workflows\r\nCustom SLMs for narrow, high-precision tasks\r\nReal-time AI features embedded directly into product workflows\r\nMulti-step AI systems supporting contextual user interactions\r\nRequirements\r\n5+ years software engineering experience\r\n2+ years working hands-on with LLMs in production\r\nProven experience fine-tuning and deploying models to real users\r\nStrong applied production track record (not research-only)\r\nPython, TypeScript / Node.js\r\nExperience deploying custom models to production\r\nDeep understanding of inference and performance tradeoffs\r\nIdeal Background / Nice to Have's\r\nWebSockets\r\nRedis\r\nONNX Runtime\r\nLLM evaluation & observability tooling\r\nOrchestration frameworks (e.g., LangChain)\r\nEarly-stage AI product companies (Seed – Series B/C)\r\nAI developer tools, automation, or code-generation platforms\r\nFast-paced startup environments\r\nPlease note - this role can not provide sponsorship of any kind. All candidates must be Green Card Holders or US Citizens.\r\nJ-18808-Ljbffr","company":"Harnham","rawCompany":"harnham","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-04-09T07:58:18.610Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1254.00","title":"Web Developers","slug":"web-developers"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead AI Engineer","description":"Hybrid – 3 days onsite\r\nAbout the Role\r\nAn early-stage, AI-native product company is hiring a Lead AI Engineer to own the model layer of its core product. This is a senior individual contributor role focused on fine-tuning, model optimization, and custom small model development — not prompt engineering or research-only experimentation. This engineer will be responsible for designing, shipping, and iterating on production-grade LLM systems with real-time user impact.\r\nOwn fine-tuning strategy (LoRA, adapters, distillation, full fine-tuning)\r\nDecide when to fine-tune vs. use system-level or prompt-based approaches\r\nImprove models based on production feedback\r\nBalance quality, latency, and cost\r\nModel Performance & Optimization\r\nOptimize inference speed and throughput\r\nImprove reliability and consistency\r\nDefine evaluation frameworks and benchmarking standards\r\nCustom Small Model Development\r\nDesign and deploy custom small language models (SLMs)\r\nDetermine when smaller models outperform larger ones\r\nMaintain real-time performance for interactive UX workflows\r\nWhat You'll Build\r\nFine-tuned models powering generation workflows\r\nCustom SLMs for narrow, high-precision tasks\r\nReal-time AI features embedded directly into product workflows\r\nMulti-step AI systems supporting contextual user interactions\r\nRequirements\r\n5+ years software engineering experience\r\n2+ years working hands-on with LLMs in production\r\nProven experience fine-tuning and deploying models to real users\r\nStrong applied production track record (not research-only)\r\nPython, TypeScript / Node.js\r\nExperience deploying custom models to production\r\nDeep understanding of inference and performance tradeoffs\r\nIdeal Background / Nice to Have's\r\nWebSockets\r\nRedis\r\nONNX Runtime\r\nLLM evaluation & observability tooling\r\nOrchestration frameworks (e.g., LangChain)\r\nEarly-stage AI product companies (Seed – Series B/C)\r\nAI developer tools, automation, or code-generation platforms\r\nFast-paced startup environments\r\nPlease note - this role can not provide sponsorship of any kind. All candidates must be Green Card Holders or US Citizens.\r\nJ-18808-Ljbffr","datePosted":"2026-04-09T07:58:18.610Z","dateModified":"2026-04-09T07:58:18.610Z","hiringOrganization":{"@type":"Organization","name":"Harnham","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"5143d40adb27d9a2e4a2c08c"},"url":"https://jobsearcher.com/jobs/5143d40adb27d9a2e4a2c08c"}}