JOBSEARCHER

Agent Reliability Engineer, GTM Engineering

LangchainMillbrae, CAL6 LeadSeptember 16th, 2026
GTM Agent EngineerAt LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating $125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we're at a stage where we're continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and The Role:You'll own the health, cost, performance, and business impact of the GTM Agent, and build the feedback loops that keep it improving. Because we build the platform we run on, you'll also operate the agent on LangSmith the way we tell customers to, and turn that practice into the reference story enterprises keep asking us for. You'll work across Python 3.11, FastAPI, LangGraph, DeepAgents, LangSmith, Supabase Postgres, BigQuery, Anthropic and OpenAI models, and Slack and Next.js surfaces.What You'll Do:Monitor production health across every graph, catching errors, slow runs, expensive runs, and silent failures before reps report themTriage incoming issues from Slack, tickets, and rep reports, fixing small things directly and routing the rest to the right ownerRun the weekly eval suite, investigate failures, and turn real production bugs into permanent regression testsTrack cost and latency by model, graph, use case, and role, and recommend concrete changes to model choice, reasoning effort, and cachingTrack usage and adoption per rep and per feature, and own the weekly health report the team runs onBuild the business metrics that show leadership what the agent is worth, from reply rates and meetings booked to hours reclaimed and ROIBuild our own monitoring and alerting on LangSmith, and write the \"how we run our own agent\" story for customersWhat You'll Bring:Strong production Python and SQL, comfortable working in traces, logs, and warehouse tablesReal experience running LLM applications, including tracing, evals, and prompt and cache mechanicsSRE or production operations instincts: percentiles, SLOs, and separating noise from real patternHealthy skepticism about metrics; you check what a number actually counts before you publish itClear writing skills, and interest in publishing what you learnHigh agency; you notice what's missing and take initiative to build itNice To Haves:LangGraph or LangSmith experienceExperience building an eval suite from scratchBigQuery or dbtPrior DevRel-adjacent writingEmpathy for sales and go-to-market usersSalary: $150,000 - $190,000Compensation Philosophy: We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.