JOBSEARCHER

AI Data Engineer

TalenthopDenver, COL6 LeadSeptember 11th, 2026
This is a Fully Remote Job1. About Our Client:The organization is a technology consulting and software development company specializing in cloud, AI, data, and enterprise solutions across the United States. It addresses the challenges of delivering scalable and efficient technology solutions to diverse clients, focusing on innovative applications in AI and data engineering.2. About the Opportunity:The AI Data Engineer role is responsible for building and operating large-scale data systems that support modern AI training and evaluation. This position contributes by ensuring high-quality, efficient data pipelines and infrastructure that enhance model training performance and reproducibility. The role is critical for maintaining the data backbone that drives AI development and continual improvement.3. Responsibilities:Design and manage large-scale data pipelines for AI training and evaluation workflows.Build ingestion systems for varied data types including text, image, audio, video, and structured signals.Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.Develop systems for dataset versioning, lineage, and provenance to enable reproducible training.Construct high-throughput data loading systems to optimize GPU utilization during training.Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement processes.Design storage architectures balancing cost, throughput, and latency.Build evaluation dataset pipelines with integrity and contamination controls.Implement data privacy, redaction, and consent enforcement throughout data pipelines.Collaborate with machine learning researchers and engineers to align data systems with model development needs.Monitor data quality, drift, and pipeline health across AI data infrastructure.Optimize cost and performance via compression, format selection, and caching strategies.Document data systems, schemas, and operational procedures.Stay updated with AI data infrastructure research and emerging open-source tools.4. Requirements:Bachelor’s or Master’s degree in Computer Science or related field.10+ years of data engineering experience, including support for machine learning or AI workloads.Proficiency in Python and at least one JVM or systems programming language.Experience with modern data processing frameworks such as Spark, Ray, or Beam.Hands-on experience managing petabyte-scale storage and pipelines.Strong understanding of distributed systems, data modeling, and storage formats.Familiarity with dataset versioning, lineage, and reproducibility for ML workflows.Experience with high-throughput data loading for accelerator-based training.Solid software engineering practices including testing, CI/CD, and code review.Excellent communication and cross-functional collaboration skills.5. Pay Range and Compensation Package:Salary range: $100,000 to $150,000 annually.Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.Note:TalentHop is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.