JOBSEARCHER

MLOps Engineer (Senior)

BetterdataDenver, COL5 SeniorSeptember 11th, 2026
Who Are We Looking ForWe seek an experienced Machine Learning Operations Engineer (Senior) to transform cutting-edge research into robust, production-ready services for synthetic data generation and to optimize both deep learning and classical ML algorithms (e.g. tree-based models) at enterprise scale (billions of rows). You will build and tune model pipelines end-to-end, ensuring high performance, scalability, and reliability across diverse workloads and dataset sizes.Key Responsibilities:Algorithm Optimization & ScalingOptimize bottlenecks of the deep generative models to accelerate training and generation of generative models (e.g. transformer, diffusion, GANs).Implement distributed training of the models across multi-GPU clusters.Optimize distributed training of traditional ML models (e.g. XGBoost, LightGBM, CatBoost) on billion-row datasets.Design best practices for memory management to maximize resource utilization (compute and memory), enabling faster training at lower cost.Data Handling at ScaleCollaborate with data engineers to design ETL/ELT workflows handling terabyte to petabyte scale tabular and unstructured data.Implement scalable feature engineering pipelines using distributed computing frameworks (e.g. Spark, Dask, or Ray).Automate data validation (e.g. schema checks, anomaly detection) with rule-based and ML-driven frameworks.End to End OrchestrationBuild ML pipelines that transition research prototypes into reliable production-grade workflow.Package models into Docker containers and deploy using Kubernetes.Build automated model and data quality monitoring and validation systems to ensure data integrity throughout the pipeline lifecycle.Design robust error handling mechanisms, with automatic retries and data recovery in case of pipeline failures.Implement logging, monitoring and alerting systems.QualificationsBachelor's or Master's degree in Computer Science, Electrical Engineering, Software Engineering, Data Science or a related quantitative discipline.5+ years of hands-on experience optimizing and scaling machine learning models in production environments.Demonstrated track record of accelerating model training workflows (e.g., transformers, diffusion models, GANs) at multi-GPU scale.Experience in operating ETL/ELT pipelines handling terabytes to petabytes of tabular and unstructured data using distributed computing tools (e.g. Apache Spark, Dask, Ray).Demonstrated ability to translate research prototypes into reliable, production-grade ML pipelines with rigorous testing and validation.Experience in the ML orchestration (e.g. airflow, dagster).Good to haveExperience hosting models to scalable cloud infrastructure (AWS / Azure / GCP).Experience containerisation of the data pipelines & AI models in docker with supporting orchestration tools (e.g. kubernetes).BenefitsFlexible time-off arrangementsFlexible work arrangements - work from office at One North or WFH on some daysEquity eligibility: Competitive equity packages, with grant size evaluated based on the candidate's experience, skills, and impact.How To ApplyDoes this role sound like a good fit to you?We see this first: Submit your application here https://www.betterdata.ai/company/job-applicationWe see this last: If the above does not work, you may email us your CV (pdf format) at jobs@betterdata.ai.Include the title of the role in your subjectIndicate your available start - end dates (DDMMYY - DDMMYY)Send along links/supporting information that best showcase the relevant things you have built and done