JOBSEARCHER

Machine Learning Engineer

About The RoleThis role focuses on bridging the gap between theoretical machine learning models and robust production-grade services. The engineer will build and optimize inference pipelines, design scalable features stores, and orchestrate automated model retraining loops that handle millions of predictions per day.The team works closely with platform engineers and data scientists to build low-latency systems. This position offers the opportunity to architect robust MLOps foundations while deploying cutting-edge architectures to solve critical business problems.Key ResponsibilitiesDesign, optimize, and deploy high-throughput ML pipelines using PyTorch or TensorFlow within Docker/Kubernetes environmentsDevelop and maintain feature engineering workflows using PySpark, SQL, and feature stores like Feast to ensure training-serving consistencyImplement model monitoring, logging, and alerting systems to track data drift, concept drift, and system latency in real-timeBuild automated CI/CD and MLOps pipelines using MLflow, Kubeflow, or Argo Workflows for seamless model promotion and rollbackCollaborate with backend engineering teams to integrate ML models into microservice architectures via high-performance gRPC or REST APIsProfile and optimize model performance at runtime, utilizing techniques such as quantization, distillation, and TensorRTWhat We Are Looking For3-6 years of professional software engineering or machine learning engineering experience, with a track record of deploying models to productionStrong proficiency in Python, including familiarity with pandas, NumPy, scikit-learn, and at least one deep learning framework (PyTorch or TensorFlow)Hands-on experience with cloud infrastructure (AWS or GCP) and managed ML services (SageMaker or Vertex AI)Familiarity with containerization (Docker), orchestration (Kubernetes), and relational/non-relational database designBS or MS in Computer Science, Data Science, Math, or a related quantitative fieldBonus: Experience with vector databases, parameter-efficient fine-tuning (PEFT), or deploying LLMs using frameworks like vLLM