{"schemaVersion":"jobsearcher.job.v1","id":"fd6f3c287b04687ce790634f","url":"https://jobsearcher.com/jobs/fd6f3c287b04687ce790634f","canonicalUrl":"https://jobsearcher.com/jobs/fd6f3c287b04687ce790634f","title":"ML Ops Engineer","description":"Position Overview\r\nAs an ML Ops Engineer at Circadia Health, you will own the infrastructure and operational lifecycle of the machine learning systems that power our clinical monitoring platform. You will build and maintain the production ML pipelines, deployment infrastructure, and monitoring systems that enable Circadia's predictive models to identify early signs of clinical deterioration.\r\nReporting to the Principal ML Engineer, you will work across ML, backend, data, and clinical teams to ensure models are reliably trained, versioned, deployed, and monitored in both cloud and edge environments. You will be a key driver in elevating Circadia's ML practice – from reproducibility and experiment tracking to CI/CD for models and operational observability.\r\nThis is a high-ownership role at a lean company where production reliability, rapid iteration, and pragmatic engineering are essential. Your work will directly impact patient outcomes by ensuring our predictive models are always running, always accurate, and always improving.\r\nKey Responsibilities\r\nOwn and extend Circadia's ML pipeline orchestration using Apache Airflow, including training, evaluation, and deployment workflows.\r\nBuild and maintain automated pipelines for model retraining, validation, and promotion across development, staging, and production environments.\r\nImplement pipeline monitoring, alerting, and failure recovery to eliminate silent failures and ensure operational reliability.\r\nDesign pipeline architectures that support rapid experimentation while enforcing production-grade reproducibility.\r\nDeploy and manage ML models on AWS infrastructure (e.g. AWS Batch for batch inference workloads).\r\nSupport deployment of models to edge devices, including Circadia's clinical monitoring hardware, working with firmware and embedded engineering teams as needed.\r\nManage model versioning, promotion, and rollback workflows through the MLflow model registry.\r\nEvaluate and implement strategies for safe model rollouts (e.g. shadow deployments, canary releases) as the platform maturing.\r\nMaintain and improve the MLflow-based experiment tracking and model registry infrastructure.\r\nEstablish conventions for experiment logging, artifact storage, model metadata, and lineage tracking.\r\nEnable ML engineers to move seamlessly from experimentation to production deployment with minimal friction.\r\nImplement and maintain training data versioning and dataset management practices to ensure reproducibility of model training runs.\r\nTrack dataset lineage, labeling provenance, and feature dependencies alongside model versions.\r\nCollaborate with ML engineers and data engineers to formalise dataset release and validation workflows.\r\nBuild monitoring systems for model performance in production, including data drift detection, prediction quality tracking, and alerting on degradation.\r\nImplement operational dashboards for pipeline health, compute utilisation, and deployment status.\r\nCollaborate with data engineering to ensure upstream data quality and pipeline reliability for ML feature inputs.\r\nDevelop incident response procedures and runbooks for ML system failures.\r\nManage and optimise AWS compute resources (Batch, EC2, or similar) used for model training and inference.\r\nDesign infrastructure-as-code solutions for reproducible ML environments.\r\nDrive cost optimisation across ML compute, storage, and data transfer.\r\nSupport Snowflake integrations for feature generation and training data pipelines.\r\nIntroduce and champion ML engineering best practices including CI/CD for models, automated testing for ML pipelines, and reproducible training workflows.\r\nBuild internal tooling and templates that accelerate the ML development-to-production cycle.\r\nDocument operational processes, architecture decisions, and onboarding materials for the ML platform.\r\nParticipate in architecture discussions and technical planning to ensure ML systems scale with Circadia's growth.\r\nEnsure all ML pipelines and infrastructure meet healthcare security and privacy requirements, including HIPAA and SOC 2.\r\nApply best practices for handling Protected Health Information (PHI) in training data, model artifacts, and inference outputs.\r\nMaintain audit trails for model decisions, data access, and deployment history.\r\nRequired Qualifications\r\n4+ years of experience in MLOps, ML Engineering, DevOps, or a closely related infrastructure role.\r\nStrong proficiency in Python for ML pipeline development, tooling, and automation.\r\nHands-on experience with ML pipeline orchestration tools, particularly Apache Airflow.\r\nExperience with model registries and experiment tracking platforms (MLflow preferred).\r\nExperience deploying and operating ML workloads on AWS (Batch, EC2, S3, IAM, CloudWatch).\r\nSolid understanding of the ML lifecycle: training, evaluation, deployment, monitoring, and retraining.\r\nExperience with containerisation (Docker) and infrastructure-as-code.\r\nProficiency with Git and version control workflows.\r\nFamiliarity with SQL and data warehousing platforms (Snowflake preferred).\r\nExperience implementing monitoring, logging, and alerting for production systems.\r\nStrong debugging and incident response skills for complex distributed systems.\r\nPreferred Qualifications\r\nExperience deploying models to edge or embedded devices.\r\nBackground in healthcare, medical devices, or clinical data systems.\r\nFamiliarity with model serving frameworks (e.g., TorchServe, TF Serving, Triton, or custom solutions).\r\nExperience with CI/CD systems for ML (e.g., GitHub Actions, Jenkins, or similar).\r\nExperience with data versioning tools (e.g., DVC, LakeFS, or similar).\r\nExperience supporting data science or ML research teams in a production context.\r\nExposure to HIPAA compliance and healthcare security best practices.\r\nExperience with distributed compute frameworks (e.g. Apache Spark, Dask) for large-scale data processing.\r\nExperience with streaming or real-time inference architectures.\r\nWhat You Bring\r\nYou take ownership of ML infrastructure end-to-end — from training pipelines to production monitoring.\r\nYou care deeply about reliability, reproducibility, and operational excellence in ML systems.\r\nYou have strong opinions (loosely held) on how to build a great ML platform, and you're eager to put them into practice.\r\nYou are comfortable working in a startup environment where you'll wear multiple hats and move fast.\r\nYou communicate clearly across engineering, data science, and clinical teams.\r\nYou're motivated by building technology that directly improves patient care.\r\nWhy Circadia Health\r\nCircadia Health is redefining patient monitoring through contactless sensing and AI-driven clinical insights. As we scale from tens of thousands to hundreds of thousands of monitored patients, our data infrastructure is central to everything we do.\r\nYou'll have the opportunity to:\r\nWork on real-world healthcare problems with measurable patient impact\r\nBuild data systems that power clinical-grade AI and ML\r\nTake ownership in a fast-growing, mission-driven company\r\nCollaborate with a highly skilled, multidisciplinary team\r\nJ-18808-Ljbffr","company":"Circadia Technologies","rawCompany":"circadia technologies","city":"El Segundo","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T01:56:23.583Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"ML Ops Engineer","description":"Position Overview\r\nAs an ML Ops Engineer at Circadia Health, you will own the infrastructure and operational lifecycle of the machine learning systems that power our clinical monitoring platform. You will build and maintain the production ML pipelines, deployment infrastructure, and monitoring systems that enable Circadia's predictive models to identify early signs of clinical deterioration.\r\nReporting to the Principal ML Engineer, you will work across ML, backend, data, and clinical teams to ensure models are reliably trained, versioned, deployed, and monitored in both cloud and edge environments. You will be a key driver in elevating Circadia's ML practice – from reproducibility and experiment tracking to CI/CD for models and operational observability.\r\nThis is a high-ownership role at a lean company where production reliability, rapid iteration, and pragmatic engineering are essential. Your work will directly impact patient outcomes by ensuring our predictive models are always running, always accurate, and always improving.\r\nKey Responsibilities\r\nOwn and extend Circadia's ML pipeline orchestration using Apache Airflow, including training, evaluation, and deployment workflows.\r\nBuild and maintain automated pipelines for model retraining, validation, and promotion across development, staging, and production environments.\r\nImplement pipeline monitoring, alerting, and failure recovery to eliminate silent failures and ensure operational reliability.\r\nDesign pipeline architectures that support rapid experimentation while enforcing production-grade reproducibility.\r\nDeploy and manage ML models on AWS infrastructure (e.g. AWS Batch for batch inference workloads).\r\nSupport deployment of models to edge devices, including Circadia's clinical monitoring hardware, working with firmware and embedded engineering teams as needed.\r\nManage model versioning, promotion, and rollback workflows through the MLflow model registry.\r\nEvaluate and implement strategies for safe model rollouts (e.g. shadow deployments, canary releases) as the platform maturing.\r\nMaintain and improve the MLflow-based experiment tracking and model registry infrastructure.\r\nEstablish conventions for experiment logging, artifact storage, model metadata, and lineage tracking.\r\nEnable ML engineers to move seamlessly from experimentation to production deployment with minimal friction.\r\nImplement and maintain training data versioning and dataset management practices to ensure reproducibility of model training runs.\r\nTrack dataset lineage, labeling provenance, and feature dependencies alongside model versions.\r\nCollaborate with ML engineers and data engineers to formalise dataset release and validation workflows.\r\nBuild monitoring systems for model performance in production, including data drift detection, prediction quality tracking, and alerting on degradation.\r\nImplement operational dashboards for pipeline health, compute utilisation, and deployment status.\r\nCollaborate with data engineering to ensure upstream data quality and pipeline reliability for ML feature inputs.\r\nDevelop incident response procedures and runbooks for ML system failures.\r\nManage and optimise AWS compute resources (Batch, EC2, or similar) used for model training and inference.\r\nDesign infrastructure-as-code solutions for reproducible ML environments.\r\nDrive cost optimisation across ML compute, storage, and data transfer.\r\nSupport Snowflake integrations for feature generation and training data pipelines.\r\nIntroduce and champion ML engineering best practices including CI/CD for models, automated testing for ML pipelines, and reproducible training workflows.\r\nBuild internal tooling and templates that accelerate the ML development-to-production cycle.\r\nDocument operational processes, architecture decisions, and onboarding materials for the ML platform.\r\nParticipate in architecture discussions and technical planning to ensure ML systems scale with Circadia's growth.\r\nEnsure all ML pipelines and infrastructure meet healthcare security and privacy requirements, including HIPAA and SOC 2.\r\nApply best practices for handling Protected Health Information (PHI) in training data, model artifacts, and inference outputs.\r\nMaintain audit trails for model decisions, data access, and deployment history.\r\nRequired Qualifications\r\n4+ years of experience in MLOps, ML Engineering, DevOps, or a closely related infrastructure role.\r\nStrong proficiency in Python for ML pipeline development, tooling, and automation.\r\nHands-on experience with ML pipeline orchestration tools, particularly Apache Airflow.\r\nExperience with model registries and experiment tracking platforms (MLflow preferred).\r\nExperience deploying and operating ML workloads on AWS (Batch, EC2, S3, IAM, CloudWatch).\r\nSolid understanding of the ML lifecycle: training, evaluation, deployment, monitoring, and retraining.\r\nExperience with containerisation (Docker) and infrastructure-as-code.\r\nProficiency with Git and version control workflows.\r\nFamiliarity with SQL and data warehousing platforms (Snowflake preferred).\r\nExperience implementing monitoring, logging, and alerting for production systems.\r\nStrong debugging and incident response skills for complex distributed systems.\r\nPreferred Qualifications\r\nExperience deploying models to edge or embedded devices.\r\nBackground in healthcare, medical devices, or clinical data systems.\r\nFamiliarity with model serving frameworks (e.g., TorchServe, TF Serving, Triton, or custom solutions).\r\nExperience with CI/CD systems for ML (e.g., GitHub Actions, Jenkins, or similar).\r\nExperience with data versioning tools (e.g., DVC, LakeFS, or similar).\r\nExperience supporting data science or ML research teams in a production context.\r\nExposure to HIPAA compliance and healthcare security best practices.\r\nExperience with distributed compute frameworks (e.g. Apache Spark, Dask) for large-scale data processing.\r\nExperience with streaming or real-time inference architectures.\r\nWhat You Bring\r\nYou take ownership of ML infrastructure end-to-end — from training pipelines to production monitoring.\r\nYou care deeply about reliability, reproducibility, and operational excellence in ML systems.\r\nYou have strong opinions (loosely held) on how to build a great ML platform, and you're eager to put them into practice.\r\nYou are comfortable working in a startup environment where you'll wear multiple hats and move fast.\r\nYou communicate clearly across engineering, data science, and clinical teams.\r\nYou're motivated by building technology that directly improves patient care.\r\nWhy Circadia Health\r\nCircadia Health is redefining patient monitoring through contactless sensing and AI-driven clinical insights. As we scale from tens of thousands to hundreds of thousands of monitored patients, our data infrastructure is central to everything we do.\r\nYou'll have the opportunity to:\r\nWork on real-world healthcare problems with measurable patient impact\r\nBuild data systems that power clinical-grade AI and ML\r\nTake ownership in a fast-growing, mission-driven company\r\nCollaborate with a highly skilled, multidisciplinary team\r\nJ-18808-Ljbffr","datePosted":"2026-07-16T01:56:23.583Z","dateModified":"2026-07-16T01:56:23.583Z","hiringOrganization":{"@type":"Organization","name":"Circadia Technologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"El Segundo","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"fd6f3c287b04687ce790634f"},"url":"https://jobsearcher.com/jobs/fd6f3c287b04687ce790634f"}}