AI – Sr. Application Engineer
Job Title: AI Sr. Application EngineerLocation: Bay area, CADuration: FTEWill work on the intelligence layer for multiple programs owns all model quality, RAG accuracy, prompt engineering, and AI safety across applications Socratic tutor persona, adaptive learning recommendation engine, multi-modal AI (text and voice), RAG evaluation framework, and feedback loop into retrieval6-LLM call chain orchestration (NeMoGuardrails intent classification query rewriting RAG synthesis), , and compatibility check logicProduction-grade AI quality from launch this is not a research or prototyping role; accuracy thresholds, latency requirements, and safety guardrails must pass InfoSec adversarial testing before Release 1ExperienceRequired SkillsTotal IT 10+ Years4 7 years of software engineering with at least 2 years focused on LLM application development in production not research, not demos, not internal tools with 10 usersHas shipped an LLM-powered feature or product to production where real users depend on the accuracy and the engineer owns the quality metricsHas owned an AI safety or guardrails implementation for a customer-facing product not just added an off-the-shelf filter; designed and tested the safety layerHas built RAG evaluation pipelines and used them to make go/no-go release decisions accuracy gating is part of the workflow.Has profiled and optimized a multi-step LLM call chain for latencyLLM Application DevelopmentLLM prompt engineering system prompts, few-shot examples, chain-of-thought, instruction following Expert Must-haveMulti-step LLM chain orchestration LangChain, LlamaIndex, or custom orchestration Expert Must-haveMulti-turn conversation design context window management, conversation summarization, session memory Advanced Must-haveStreaming LLM response handling token-by-token streaming, partial response rendering Advanced Must-haveModel selection and benchmarking matching model size to task; balancing latency, cost, and accuracy Advanced Must-haveRAG Pipeline Design & QualityRAG pipeline design chunking strategy, embedding model selection, retrieval configuration Expert Must-haveVector similarity search tuning index parameters, similarity thresholds, retrieval depth Advanced Must-haveReranking cross-encoder rerankers, relevance scoring Advanced Must-haveRAG evaluation frameworks RAGAS, TruLens, or equivalent; automated eval pipelines Advanced Must-haveHybrid search combining dense vector retrieval with BM25 or keyword search Proficient Nice to haveAI Safety & GuardrailsPrompt injection detection and mitigation Advanced Must-haveJailbreak testing and red-teaming LLM systems Advanced Must-haveContent safety classifier integration Advanced Must-haveHallucination detection and mitigation strategies Advanced Must-haveTopical control enforcing scope boundaries on LLM responses Advanced Must-haveEvaluation & Production QualityAutomated evaluation pipeline design test set curation, metric selection, regression detection Advanced Must-haveA/B evaluation methodology for prompt and model changes Proficient Must-haveLatency profiling for LLM call chains identifying bottlenecks across multi-step pipelines Proficient Must-haveFeedback loop design user signal collection, signal-to-retrieval-weight integration Proficient Must-haveProduction model monitoring accuracy drift detection, quality degradation alerting Proficient Must-haveDevelopmentPython ML/AI application development, async programming Expert Must-haveAPI design for AI services streaming endpoints, error handling, timeout management Advanced Must-haveEmbedding model operations model selection, batch embedding, index updates Advanced Must-haveNice to HaveAdaptive learning systems or personalization engine experienceKnowledge graph integration with RAGMulti-agent orchestration patterns ServiceNow API integrationPrior experience building AI products on NVIDIA infrastructure