{"schemaVersion":"jobsearcher.job.v1","id":"449fb4c2f4f764064ccf0b39","url":"https://jobsearcher.com/jobs/449fb4c2f4f764064ccf0b39","canonicalUrl":"https://jobsearcher.com/jobs/449fb4c2f4f764064ccf0b39","title":"Data Engineer with AI/ML","description":"This role bridges traditional data engineering (ETL/ELT) and data science, ensuring data is clean, accessible, and optimized for model training, real-time inference, and Generative AI (GenAI) workloads.ResponsibilitiesAI/ML Pipeline Development: Design and maintain robust ETL/ELT pipelines specifically for feeding data into machine learning models.Data Preparation & Feature Engineering: Automate data preprocessing (normalization, encoding, augmentation) and collaborate with data scientists to create feature stores.Infrastructure Optimization: Implement data lakes, warehouses, and vector databases (e.g., Pinecone, Weaviate) optimized for AI workloads.MLOps & Deployment: Implement MLOps practices, including CI/CD, model versioning, monitoring model performance, and automating retrain workflows.Real-time Streaming: Implement real-time data streaming for live inference using tools like Kafka or Flink.Data Governance & Security: Ensure data quality, lineage tracking, and compliance with privacy standards (GDPR, CCPA) in AI contexts. QualificationsEducation: Bachelor’s or Master’s in Computer Science, Data Engineering, or a related field.Experience: Generally 3–7+ years in data engineering or related ML roles. Required SkillsProgramming: High proficiency in Python and SQL; experience with Scala or Java is often preferred.Big Data Technologies: Experience with Apache Spark, Hadoop, Hive, and Kafka.Cloud Platforms: Proficiency in AWS, Azure, or GCP cloud services.ML Frameworks: Familiarity with TensorFlow, PyTorch, or Scikit-learn.Data Tools: Experience with Airflow, dbt, and MLflow.","company":"Sharpatoms","rawCompany":"sharpatoms","city":"Houston","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-05-11T21:10:58.757Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer with AI/ML","description":"This role bridges traditional data engineering (ETL/ELT) and data science, ensuring data is clean, accessible, and optimized for model training, real-time inference, and Generative AI (GenAI) workloads.ResponsibilitiesAI/ML Pipeline Development: Design and maintain robust ETL/ELT pipelines specifically for feeding data into machine learning models.Data Preparation & Feature Engineering: Automate data preprocessing (normalization, encoding, augmentation) and collaborate with data scientists to create feature stores.Infrastructure Optimization: Implement data lakes, warehouses, and vector databases (e.g., Pinecone, Weaviate) optimized for AI workloads.MLOps & Deployment: Implement MLOps practices, including CI/CD, model versioning, monitoring model performance, and automating retrain workflows.Real-time Streaming: Implement real-time data streaming for live inference using tools like Kafka or Flink.Data Governance & Security: Ensure data quality, lineage tracking, and compliance with privacy standards (GDPR, CCPA) in AI contexts. QualificationsEducation: Bachelor’s or Master’s in Computer Science, Data Engineering, or a related field.Experience: Generally 3–7+ years in data engineering or related ML roles. Required SkillsProgramming: High proficiency in Python and SQL; experience with Scala or Java is often preferred.Big Data Technologies: Experience with Apache Spark, Hadoop, Hive, and Kafka.Cloud Platforms: Proficiency in AWS, Azure, or GCP cloud services.ML Frameworks: Familiarity with TensorFlow, PyTorch, or Scikit-learn.Data Tools: Experience with Airflow, dbt, and MLflow.","datePosted":"2026-05-11T21:10:58.757Z","dateModified":"2026-05-11T21:10:58.757Z","hiringOrganization":{"@type":"Organization","name":"Sharpatoms","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Houston","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"449fb4c2f4f764064ccf0b39"},"url":"https://jobsearcher.com/jobs/449fb4c2f4f764064ccf0b39"}}