Software Engineer, Agents (Internal Audit)
About UsFieldguide is establishing a new state of trust for global commerce and capital markets by automating and streamlining the work of assurance and audit practitioners - specifically in cybersecurity, privacy, and financial audits. We build software for the people who enable trust between businesses.We're based in San Francisco, CA, and we're backed by Goldman Sachs Alternatives, Bessemer Venture Partners, 8VC, Floodgate, Y Combinator, and more. Over 50 of the top 100 accounting and consulting firms trust Fieldguide to power mission-critical work.About the RoleYou'll join a genuine 0→1 team on the ground floor of one of the company's biggest new bets. This seat is specifically product-focused: you'll own agent quality, ship agents that do real audit work, and work alongside practitioners.Depending on your experience and what you're looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team. We're hiring across all levels and will calibrate during interviews based on scope and demonstrated experience.What You'll DoMake agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixesTradeoffs such as quality/latency/cost across a long multi-phase runBuild structured-output pipelines that turn model output into real audit artifactsTake ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and whyWork directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within daysExpand agent coverage into new controls and new areas of internal auditWho You Are (All Levels)Product-minded and full-stack: you've shipped LLM-backed features to production against real users, and you measure yourself on whether they got usedYou're fluent in evals and error analysis, and you apply them in service of shipping something practitioners trustYou have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibesEnergized by 0→1 work: you'd rather define the problem than inherit a spec, and you don't stall on ambiguityStrong instincts for human-in-the-loop designA genuine team player across the organization, not just within engineering: you'll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the jobShip fast without leaving a mess: your code is reviewable, tested where it counts, and instrumentedAble to internalize a hard domain fast. You don't need to know SOX today, but you'll understand it well enough to make the right product callsHigher-Level ResponsibilitiesAt the Senior level, you may:Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signsSet the evals and error-analysis practice for the team's agent work, and decide what evidence justifies shipping a change or rolling it backCollaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decidesOwn the harder model and orchestration judgment calls across a long multi-phase runMentor other engineers and raise the bar on 0→1 execution and applied eval rigorAt the Staff level, you may:Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across FieldguideSet and champion engineering standards for agent reliability, reproducibility, and defensibilityPartner with engineering and product leadership to define long-term technical strategy for agentic audit workServe as a trusted advisor to leaders across Engineering, Product, and DesignRepresent Fieldguide externally through writing, speaking, and open-source contributionsExperienceMust-have:Shipped LLM-backed product features to production against real usersApplied AI skillset: evals, error analysis, and model-selection decisions you owned and can explainComfortable full-stack, with enough backend depth to work in agent orchestrationAutonomy working from an ambiguous specA collaborative mode that works across PM, design, and domain expertsNice-to-have:Python, TypeScript, React, Postgres, Hasura, GraphQLTemporal or comparable durable-execution / workflow orchestrationHands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)Structured-output work including schema contracts, generating real artifacts from model outputStartup experience, as a founder or as an early engineerA 0→1 track record: things you started where no scaffolding existedExperience working directly with customers, and comfort being in the room when they use what you builtBackground in internal audit, SOX, accounting, or another regulated domainDocument processing, including PDF and Excel manipulation and annotationNot a fit if:Prompt engineering is your whole skill setYour agent work never carried production trafficYou want to own eval methodology or the evaluation harness itself rather than ship product features (better fit on Foundation Agents)You need a fully specified ticket to startYou'd rather not be in the room with customers and domain expertsWhat Should Excite You0→1 on the biggest bet: You're building the agent and the product from scratch, on the ground floor of where the company is goingRepeatable judgment: Making an agent reach the same defensible conclusion twice, in a domain where ground truth requires expert judgmentReal audit stakes: Your work directly affects what firms put in front of their clients, and what a reviewer is willing to signCustomer proximity: Design-partner firms and an embedded SOX expert use what you ship within days of it landingHuman-in-the-loop design: Deciding where the agent acts and where the auditor decides, on work that genuinely mattersHigh trust, high autonomy: You're given ambiguous problems and trusted to define the planBenefitsCompetitive compensation with equityComprehensive health and wellness benefitsFlexible time off and work schedulesTechnology reimbursements401(k) planTwice-yearly in-person offsites across the U.S.Wellness benefits starting on your first dayOur ValuesFearless - Inspire and break down seemingly impossible wallsFast - Launch fast with excellence; iterate to perfectionLovable - Deliver happiness and 11-star experiencesOwners - Execute and run the business with ownershipWin-win - Create mutual value and earn trust for lifeInclusive - Scale the best ideas with inclusive teams