ML Ops Engineer
OVERVIEW:We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.GENERAL DUTIES:Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.REQUIRED QUALIFICATIONS:Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.Feature store consumption (Feast or equivalent) with real focus on training/serving skew.Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.DESIRED QUALIFICATIONS:The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).OpenShift and/or air-gapped model deployment.AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.CLEARANCE:Active U.S. Citizenship with the eligibility to gain a clearance