{"schemaVersion":"jobsearcher.job.v1","id":"84e7bebbfe287df501fbc3d3","url":"https://jobsearcher.com/jobs/84e7bebbfe287df501fbc3d3","canonicalUrl":"https://jobsearcher.com/jobs/84e7bebbfe287df501fbc3d3","title":"AI Engineer","description":"ABOUT THE ROLE\nAI Engineers at Varick own the intelligence layer. You design, build, and optimize the agent systems that run inside enterprise operations — processing thousands of transactions, making classification decisions, routing exceptions, and learning from human feedback.\n\nThis role is for engineers who have been deep in LLMs, agent architectures, and evaluation systems. You’ve built agentic workflows that run in production, not just demos. You understand prompt engineering, retrieval, tool calling, multi‑agent orchestration, and the evaluation infrastructure required to ship AI systems that enterprises trust.\n\nWHAT YOU'LL DO\n\nDesign and build agent architectures for complex enterprise workflows (multi-step reasoning, tool calling, exception handling)\n\nBuild and maintain evaluation systems for agent quality, accuracy, safety, and groundedness\n\nDesign prompt systems, retrieval pipelines, and context engineering strategies for reliable agent behavior\n\nBuild the feedback loops that allow agents to learn from human corrections and improve over time\n\nOptimize inference cost and latency for production workloads\n\nDefine best practices for agent reliability, observability, and governance\n\nStay current with the latest models, frameworks, and research — and ship what matters into production\n\nWHAT WE'RE LOOKING FOR\n\n3+ years of software engineering with at least 1–2 years focused on LLM applications or AI systems in production\n\nHands‑on experience building agentic workflows with tool calling, retrieval, and multi-step reasoning\n\nDeep understanding of prompt engineering, context engineering , and how to get reliable behavior from LLMs\n\nExperience building evaluation and quality systems for AI outputs\n\nStrong Python skills and backend engineering fundamentals\n\nYou’ve shipped AI features to real users and dealt with the messy parts: hallucinations, edge cases, accuracy degradation, cost management\n\nBased in SF.\n\nHELPFUL EXPERIENCE\n\nAgent frameworks: LangGraph, CrewAI, Claude Code/Codex patterns , or custom orchestration\n\nRetrieval systems: vector databases (Qdrant, pgvector, Pinecone), reranking, hybrid search\n\nMCP, tool-calling protocols, and third-party API integrations\n\nFine‑tuning, LoRA , or other model adaptation methods\n\nEvaluation frameworks and continuous quality monitoring\n\nExperience with enterprise AI deployments (compliance, audit trails, governance)\n\nPrior work at AI labs, AI-native startups, or applied ML teams\n\nWHY VARICK\n\nShip to production, not to demos. Every system you build runs inside real enterprise operations. 100% deployment rate.\n\nEarly enough to shape everything. Your work defines the product, the platform, and the company.\n\nCompounding impact. Every client deployment feeds the pattern library and makes the next one faster. You’re building leverage, not doing the same thing twice.\n\nWork with operators, not committees. You talk directly to the people who run the business — CFOs, COOs, ops leads — not procurement layers.\n\n#J-18808-Ljbffr","company":"Gravity Engineering Services","rawCompany":"gravity engineering services","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:37:13.718Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Engineer","description":"ABOUT THE ROLE\nAI Engineers at Varick own the intelligence layer. You design, build, and optimize the agent systems that run inside enterprise operations — processing thousands of transactions, making classification decisions, routing exceptions, and learning from human feedback.\n\nThis role is for engineers who have been deep in LLMs, agent architectures, and evaluation systems. You’ve built agentic workflows that run in production, not just demos. You understand prompt engineering, retrieval, tool calling, multi‑agent orchestration, and the evaluation infrastructure required to ship AI systems that enterprises trust.\n\nWHAT YOU'LL DO\n\nDesign and build agent architectures for complex enterprise workflows (multi-step reasoning, tool calling, exception handling)\n\nBuild and maintain evaluation systems for agent quality, accuracy, safety, and groundedness\n\nDesign prompt systems, retrieval pipelines, and context engineering strategies for reliable agent behavior\n\nBuild the feedback loops that allow agents to learn from human corrections and improve over time\n\nOptimize inference cost and latency for production workloads\n\nDefine best practices for agent reliability, observability, and governance\n\nStay current with the latest models, frameworks, and research — and ship what matters into production\n\nWHAT WE'RE LOOKING FOR\n\n3+ years of software engineering with at least 1–2 years focused on LLM applications or AI systems in production\n\nHands‑on experience building agentic workflows with tool calling, retrieval, and multi-step reasoning\n\nDeep understanding of prompt engineering, context engineering , and how to get reliable behavior from LLMs\n\nExperience building evaluation and quality systems for AI outputs\n\nStrong Python skills and backend engineering fundamentals\n\nYou’ve shipped AI features to real users and dealt with the messy parts: hallucinations, edge cases, accuracy degradation, cost management\n\nBased in SF.\n\nHELPFUL EXPERIENCE\n\nAgent frameworks: LangGraph, CrewAI, Claude Code/Codex patterns , or custom orchestration\n\nRetrieval systems: vector databases (Qdrant, pgvector, Pinecone), reranking, hybrid search\n\nMCP, tool-calling protocols, and third-party API integrations\n\nFine‑tuning, LoRA , or other model adaptation methods\n\nEvaluation frameworks and continuous quality monitoring\n\nExperience with enterprise AI deployments (compliance, audit trails, governance)\n\nPrior work at AI labs, AI-native startups, or applied ML teams\n\nWHY VARICK\n\nShip to production, not to demos. Every system you build runs inside real enterprise operations. 100% deployment rate.\n\nEarly enough to shape everything. Your work defines the product, the platform, and the company.\n\nCompounding impact. Every client deployment feeds the pattern library and makes the next one faster. You’re building leverage, not doing the same thing twice.\n\nWork with operators, not committees. You talk directly to the people who run the business — CFOs, COOs, ops leads — not procurement layers.\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:37:13.718Z","dateModified":"2026-07-16T03:37:13.718Z","hiringOrganization":{"@type":"Organization","name":"Gravity Engineering Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"84e7bebbfe287df501fbc3d3"},"url":"https://jobsearcher.com/jobs/84e7bebbfe287df501fbc3d3"}}