Machine Learning Engineer
Overview
As a key member of the ML/GenAI platform, you design and operate scalable ML infrastructure on Databricks, enabling robust experiment tracking, model governance, and production serving. You lead ML Ops platforms, automated pipelines, and feature storage to support reliable model deployment and lifecycle management. You architect RAG systems and vector search solutions for enterprise document access, while overseeing LLM fine-tuning and safety guardrails for compliant, high-quality outputs. You drive CI/CD and orchestration automation to streamline training, testing, deployment, and monitoring across traditional ML and GenAI workloads.
ResponsibilitiesDesign, implement, and maintain scalable ML infrastructure on Databricks (MLflow, model registry, serving endpoints)Oversee ML Ops platform and automated deployment/monitoring pipelines in productionManage model versioning, retraining, and artifact governance using Unity CatalogDevelop and manage Databricks Feature Store for consistent training/inference feature engineeringArchitect Retrieval-Augmented Generation (RAG) systems for document Q&ADeploy and manage vector search solutions (Databricks Vector Search, Pinecone, etc.) for semantic retrievalLead LLM fine-tuning and customization with CIM data while ensuring privacy/complianceCreate document processing pipelines (PDF parsing, chunking, embeddings) for RAG appsImplement prompt engineering practices and LLM evaluation frameworksBuild guardrails for GenAI safety, including hallucination detection and source attributionDesign comprehensive automation across ML workflows (training, testing, validation, deployment)Establish CI/CD pipelines using GitHub Actions, Azure DevOps, etc.Automate data/model workflows using Airflow/Prefect/Databricks Workflows
Key requirementsDatabricksMLflowUnity CatalogDatabricks Feature StoreRAGvector databases