{"schemaVersion":"jobsearcher.job.v1","id":"2ffa3dac16b845e1fb253055","url":"https://jobsearcher.com/jobs/2ffa3dac16b845e1fb253055","canonicalUrl":"https://jobsearcher.com/jobs/2ffa3dac16b845e1fb253055","title":"ML Ops Engineer","description":"OVERVIEW:We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.GENERAL DUTIES:Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.REQUIRED QUALIFICATIONS:Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.Feature store consumption (Feast or equivalent) with real focus on training/serving skew.Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.DESIRED QUALIFICATIONS:The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).OpenShift and/or air-gapped model deployment.AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.CLEARANCE:Active U.S. Citizenship with the eligibility to gain a clearance","company":"Bana Solutions","rawCompany":"bana solutions","city":"Chantilly","state":"VA","isRemote":false,"isActive":false,"createdAt":"2026-08-06T23:45:21.366Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"ML Ops Engineer","description":"OVERVIEW:We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.GENERAL DUTIES:Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.REQUIRED QUALIFICATIONS:Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.Feature store consumption (Feast or equivalent) with real focus on training/serving skew.Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.DESIRED QUALIFICATIONS:The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).OpenShift and/or air-gapped model deployment.AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.CLEARANCE:Active U.S. Citizenship with the eligibility to gain a clearance","datePosted":"2026-08-06T23:45:21.366Z","dateModified":"2026-08-06T23:45:21.366Z","hiringOrganization":{"@type":"Organization","name":"Bana Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Chantilly","addressRegion":"VA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"2ffa3dac16b845e1fb253055"},"url":"https://jobsearcher.com/jobs/2ffa3dac16b845e1fb253055"}}