JOBSEARCHER

Data Engineer with Data Testing Experience

QuantiphiDenver, COL5 SeniorSeptember 25th, 2026
About Quantiphi:Quantiphi is an award-winning, AI-First global digital engineering company that helps the world’s leading Fortune 1000 organizations transform bold ideas into measurable business impact. We go beyond building innovative AI technologies—we solve the problems that matter most to our clients.Since our founding in 2013, Quantiphi has built a proven track record of turning complex challenges into meaningful outcomes across industries.Headquartered in Boston, with more than 4,000 professionals worldwide, we partner with global enterprises to deliver large-scale digital, cloud, and AI-driven transformation. #SolvingWhatMattersWe are an Elite and Premier partner to Google Cloud, AWS, NVIDIA, Snowflake, and other leading technology platforms, and our work has been recognized across the industry, including:21 Google Cloud Partner of the Year awards in the past 10 years3 AWS AI/ML Partner of the Year awards3 NVIDIA Partner of the Year awards3 Snowflake Partner of the Year awardsRated Leaders by Gartner, Forrester, IDC, ISG, Everest Group and other leading analyst firmsQuantiphi delivers First-in-class AI solutions across Life Sciences, Healthcare, Banking, Financial Services, CPG, Manufacturing, Energy, High-Tech, Telecommunications, etc., powered by cutting-edge Generative AI and Agentic AI accelerators.We are also proud to be certified as a Great Place to Work—reflecting our commitment to our people and our culture.For more details, visit: Website or LinkedIn PagePosition Overview : We are looking for an experienced Data Engineer with strong hands-on expertise in Python, PySpark, Apache Airflow, and data engineering. The ideal candidate will have experience working with large-scale data processing environments and, preferably, a background in Healthcare data and systems.This role combines data engineering and data/ETL testing responsibilities. The candidate will be responsible for building and maintaining scalable data pipelines while ensuring data quality, accuracy, completeness, and reliability through comprehensive testing.The role requires strong analytical skills, attention to detail, and the ability to collaborate with data engineers, QA teams, business stakeholders, and offshore/nearshore teams.Key Responsibilities : Design, develop, and maintain scalable data pipelines using Python and PySpark.Develop efficient batch data processing and transformation workflows using Apache Spark/PySpark.Build data ingestion and transformation processes across structured and semi-structured data sources.Develop and optimize ETL/ELT pipelines for large volumes of data.Implement data transformations, cleansing, enrichment, and validation processes.Optimize Spark jobs for performance, scalability, and resource utilization.Work with SQL databases, cloud data platforms, data lakes, and data warehouses.Troubleshoot production data pipelines and resolve data processing issues.Develop, maintain, and monitor Apache Airflow DAGs for data pipeline orchestration.Make data and ETL testing a core part of the data engineering lifecycle.Develop and execute test cases for ETL/ELT pipelines and data transformations.Validate source-to-target data reconciliation and transformation logic.Perform data completeness, accuracy, consistency, uniqueness, and integrity checks.Validate record counts, aggregates, data types, null values, duplicates, and business rules.Identify data quality issues and work with engineering teams to determine root causes.Develop automated data validation and testing frameworks using Python/PySpark/SQL where appropriate.Perform regression testing following pipeline, schema, or transformation changes.Validate end-to-end data flows across ingestion, transformation, storage, and consumption layers.Support production validation and post-deployment data quality checks.Required Skills & Experience : 5+ years of experience in Data Engineering or a related field.Experience with GCP Cloud is must.Strong hands-on experience with Python and PySpark.Strong SQL development and data analysis skills.Experience developing and maintaining ETL/ELT data pipelines.Hands-on experience with Apache Airflow and DAG development.Strong understanding of distributed data processing and Apache Spark.Experience with data validation, reconciliation, and ETL/data testing.Ability to analyze complex data issues and perform root-cause analysis.Experience working with relational and/or cloud-based data warehouses.Familiarity with data lakes, file-based data, APIs, and structured/semi-structured data.Experience with Git and modern software development practices.Strong understanding of data quality, data integrity, and pipeline reliability.Preferred Skills : Healthcare domain experience.Experience with HL7, FHIR, healthcare claims, clinical, provider, member/patient, or pharmacy data.Experience with Snowflake, Databricks, Delta Lake, or similar modern data platforms.Experience with automated data testing frameworks.Experience with CI/CD and DevOps practices for data engineering.Familiarity with data observability and data quality tools.Experience working in Agile/Scrum environments.