{"schemaVersion":"jobsearcher.job.v1","id":"66de1ff2650581a2d7fe136c","url":"https://jobsearcher.com/jobs/66de1ff2650581a2d7fe136c","canonicalUrl":"https://jobsearcher.com/jobs/66de1ff2650581a2d7fe136c","title":"Forward Deployed AI Engineer (remote - US)","description":"Forward Deployed AI Engineer(Internal title: AI Field Engineer)Location: Remote (US) Work model: Fully remote, US-based. Regular travel to customer sites. Salary: $176K–$224K base | $220K–$280K OTE (80/20 base/variable) + meaningful equityI'm hiring a Forward Deployed AI Engineer for a generative AI inference platform built for open models — independently benchmarked as the fastest in the industry, and running production workloads behind AI products you've almost certainly used. ~200 people, $300M+ raised, founded in 2021 by engineers from one of the most respected ML infrastructure teams in the industry.This is the technical tip of the spear: you embed with the company's most ambitious customers and turn complex GenAI problems into production systems, fast.The problem you'd help solve:Corporate GenAI dies in the gap between the demo and production. A closed-model API wrapper gets a team to a convincing prototype. Then the workload hits real traffic — latency budgets, cost per token, concurrency, data residency, an ML team that wants to fine-tune on their own data — and the prototype doesn't survive contact.Open models solve that, but only if someone can pick the right model, tune it, serve it on the right hardware shape, prove it with real load tests and evals, and get it running inside the customer's infrastructure and security constraints.You'd own that gap end to end: lead technical discovery, scope and build the POC hands-on-keyboard in the customer's environment, run the benchmarks that prove the architecture, guide fine-tuning strategy (SFT, DPO, RFT), and drive it through to production at scale. Every recurring pain point you find goes straight back into the product roadmap — you're the feedback loop between the field and engineering.This role isn't advisory. You ship the code.You'll likely be a fit if you have:3+ years in customer-facing AI/ML field engineering — FDE, Applied AI Engineer, Solutions Architect, AI Infra or ML Engineer with pre-sales exposure, or a research background moving into customer-facing workDeep hands-on experience with LLM inference and/or training, with working knowledge of open-model frameworks (vLLM, SGLang, TensorRT-LLM). Closed-model / API-wrapper experience alone won't clear the barHands-on fine-tuning workflows — SFT at minimum, DPO/RFT a strong plusShipped production code inside someone else's prod system, not just architecture decksStrong Python, GPU/cloud infrastructure experience (AWS, Azure or GCP), and comfort with KubernetesRun load tests and evals that actually decide an architecture — throughput, TTFT, p95/p99 latency, cost-performanceExecutive presence: a technical deep dive with an ML engineer and an architecture trade-off conversation with a VP, in the same afternoonNice to have:Hyperscaler GenAI or AI infrastructure field background — Bedrock, SageMaker, Vertex AI, Azure AI FoundryOpen-source contributions to the inference stack (vLLM, SGLang, Triton, TensorRT-LLM)A track record of building repeatable assets — reference architectures, deployment playbooks, benchmark harnesses, workshops — that let other teams move without youMultimodal, function calling, or agentic production workloadsDemonstrable commercial impact from technical work: pipeline built, deals closed, customers scaledWhat you won't find here:This won't suit you if you want a clean specialty, clear job boundaries, or someone else to hand the hard part to. Expect ambiguity, rapid context-switching, and ownership from day one. The team screens hard for arrogance — in the hiring manager's words, \"we don't vibe with people that are arrogant.\"The package:$176K–$224K base | $220K–$280K OTE (80/20 split, variable paid quarterly on individual and team performance)Compensation scales with experience — 10+ years may be considered above rangeMeaningful equity on top of OTEFully remote, US-basedRegular on-site travel to customersH-1B transfers and TN visas sponsored; O-1 considered case by case","company":"Confidential","rawCompany":"confidential","isRemote":true,"isActive":false,"createdAt":"2026-07-27T12:05:16.431Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Forward Deployed AI Engineer (remote - US)","description":"Forward Deployed AI Engineer(Internal title: AI Field Engineer)Location: Remote (US) Work model: Fully remote, US-based. Regular travel to customer sites. Salary: $176K–$224K base | $220K–$280K OTE (80/20 base/variable) + meaningful equityI'm hiring a Forward Deployed AI Engineer for a generative AI inference platform built for open models — independently benchmarked as the fastest in the industry, and running production workloads behind AI products you've almost certainly used. ~200 people, $300M+ raised, founded in 2021 by engineers from one of the most respected ML infrastructure teams in the industry.This is the technical tip of the spear: you embed with the company's most ambitious customers and turn complex GenAI problems into production systems, fast.The problem you'd help solve:Corporate GenAI dies in the gap between the demo and production. A closed-model API wrapper gets a team to a convincing prototype. Then the workload hits real traffic — latency budgets, cost per token, concurrency, data residency, an ML team that wants to fine-tune on their own data — and the prototype doesn't survive contact.Open models solve that, but only if someone can pick the right model, tune it, serve it on the right hardware shape, prove it with real load tests and evals, and get it running inside the customer's infrastructure and security constraints.You'd own that gap end to end: lead technical discovery, scope and build the POC hands-on-keyboard in the customer's environment, run the benchmarks that prove the architecture, guide fine-tuning strategy (SFT, DPO, RFT), and drive it through to production at scale. Every recurring pain point you find goes straight back into the product roadmap — you're the feedback loop between the field and engineering.This role isn't advisory. You ship the code.You'll likely be a fit if you have:3+ years in customer-facing AI/ML field engineering — FDE, Applied AI Engineer, Solutions Architect, AI Infra or ML Engineer with pre-sales exposure, or a research background moving into customer-facing workDeep hands-on experience with LLM inference and/or training, with working knowledge of open-model frameworks (vLLM, SGLang, TensorRT-LLM). Closed-model / API-wrapper experience alone won't clear the barHands-on fine-tuning workflows — SFT at minimum, DPO/RFT a strong plusShipped production code inside someone else's prod system, not just architecture decksStrong Python, GPU/cloud infrastructure experience (AWS, Azure or GCP), and comfort with KubernetesRun load tests and evals that actually decide an architecture — throughput, TTFT, p95/p99 latency, cost-performanceExecutive presence: a technical deep dive with an ML engineer and an architecture trade-off conversation with a VP, in the same afternoonNice to have:Hyperscaler GenAI or AI infrastructure field background — Bedrock, SageMaker, Vertex AI, Azure AI FoundryOpen-source contributions to the inference stack (vLLM, SGLang, Triton, TensorRT-LLM)A track record of building repeatable assets — reference architectures, deployment playbooks, benchmark harnesses, workshops — that let other teams move without youMultimodal, function calling, or agentic production workloadsDemonstrable commercial impact from technical work: pipeline built, deals closed, customers scaledWhat you won't find here:This won't suit you if you want a clean specialty, clear job boundaries, or someone else to hand the hard part to. Expect ambiguity, rapid context-switching, and ownership from day one. The team screens hard for arrogance — in the hiring manager's words, \"we don't vibe with people that are arrogant.\"The package:$176K–$224K base | $220K–$280K OTE (80/20 split, variable paid quarterly on individual and team performance)Compensation scales with experience — 10+ years may be considered above rangeMeaningful equity on top of OTEFully remote, US-basedRegular on-site travel to customersH-1B transfers and TN visas sponsored; O-1 considered case by case","datePosted":"2026-07-27T12:05:16.431Z","dateModified":"2026-07-27T12:05:16.431Z","hiringOrganization":{"@type":"Organization","name":"Confidential","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"66de1ff2650581a2d7fe136c"},"url":"https://jobsearcher.com/jobs/66de1ff2650581a2d7fe136c"}}