{"schemaVersion":"jobsearcher.job.v1","id":"04ba130bca982696e8f83cef","url":"https://jobsearcher.com/jobs/04ba130bca982696e8f83cef","canonicalUrl":"https://jobsearcher.com/jobs/04ba130bca982696e8f83cef","title":"Remote Data Engineer","description":"DescriptionJob Title: Data EngineerPosition Type: Full-Time, RemoteWorking Hours: U.S. client business hours (with flexibility for pipeline monitoring, deployments, and data refresh cycles)About the RoleOur client is seeking a Data Engineer to design, build, and maintain scalable data infrastructure and reliable data pipelines that power analytics, reporting, and operational decision-making across the business.This role requires strong software engineering fundamentals, deep experience with modern data stacks, and a passion for building clean, reliable, and high-performance data systems. The Data Engineer will ensure data flows seamlessly from source systems into warehouses, dashboards, and downstream applications while maintaining high standards for quality, governance, and scalability.The ideal candidate is analytical, detail-oriented, and comfortable working across engineering, analytics, and business teams to deliver trustworthy and actionable data.ResponsibilitiesPipeline Development & Data IntegrationBuild, maintain, and optimize ETL/ELT pipelines using Python, SQL, or ScalaOrchestrate workflows using Airflow, Prefect, Dagster, or similar orchestration toolsIngest structured and unstructured data from APIs, SaaS platforms, databases, files, and streaming systemsDevelop scalable connectors and automated ingestion workflowsData Warehousing & ModelingManage and optimize cloud data warehouses such as Snowflake, BigQuery, or RedshiftDesign scalable schemas using star and snowflake modeling techniquesImplement partitioning, clustering, indexing, and performance optimization strategiesBuild clean, analytics-ready datasets for business intelligence and reporting use casesData Quality, Governance & ReliabilityImplement validation checks, anomaly detection, logging, and monitoring to ensure data integrityEnforce naming conventions, lineage tracking, and documentation standards using tools such as dbt or Great ExpectationsMaintain audit-ready data processes and ensure compliance with GDPR, HIPAA, or industry-specific requirementsMonitor pipeline health and proactively resolve failures or inconsistenciesStreaming & Real-Time Data ProcessingBuild and manage real-time data pipelines using Kafka, Kinesis, Pub/Sub, or similar platformsSupport low-latency ingestion and event-driven architectures for time-sensitive applicationsMonitor streaming infrastructure and optimize throughput and reliabilityCollaboration & Analytics EnablementPartner closely with analysts, data scientists, and business stakeholders to deliver reliable datasetsSupport dashboard and reporting initiatives across Tableau, Looker, or Power BITranslate business requirements into scalable data solutions and modelsMaintain clear technical documentation for pipelines, schemas, and workflowsInfrastructure, DevOps & AutomationContainerize data services using Docker and manage deployments through Kubernetes when applicableAutomate deployments using CI/CD pipelines such as GitHub Actions, Jenkins, or GitLab CIManage cloud infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code toolsContinuously optimize performance, scalability, reliability, and cloud costsWhat Makes You a Perfect FitPassionate about building clean, reliable, and scalable data systemsStrong debugging and problem-solving mindset with high attention to detailBalance of software engineering discipline and analytical thinkingComfortable working cross-functionally with technical and non-technical stakeholdersProactive communicator who takes ownership of data quality and reliabilityRequired Experience & Skills3+ years of experience in Data Engineering, Back-End Engineering, or Data Infrastructure rolesStrong proficiency in Python and SQLExperience with at least one modern data warehouse (Snowflake, Redshift, BigQuery)Hands-on experience with orchestration tools such as Airflow or PrefectStrong understanding of ETL/ELT pipelines, data modeling, and data transformation workflowsFamiliarity with cloud platforms such as AWS, GCP, or AzurePreferred Experience & SkillsExperience with dbt for data modeling and transformation managementStreaming and event-driven data pipeline experience (Kafka, Kinesis, Pub/Sub)Experience with cloud-native data services such as AWS Glue, GCP Dataflow, or Azure Data FactoryFamiliarity with Docker, Kubernetes, Terraform, or CI/CD workflowsBackground in regulated industries such as healthcare, fintech, or enterprise SaaSExperience optimizing warehouse costs and query performance at scaleWhat Does a Typical Day Look Like?A Data Engineer's day revolves around maintaining reliable pipelines, improving data quality, and enabling teams with scalable access to trustworthy data. You will:Monitor pipeline health and troubleshoot failed jobs in Airflow or related orchestration systemsBuild and maintain ingestion pipelines for APIs, SaaS platforms, and operational databasesOptimize SQL queries and warehouse performance to improve efficiency and reduce cloud costsCollaborate with analysts and data scientists to provide curated datasets for reporting and modelingImplement validation checks and monitoring to prevent downstream data quality issuesDocument data models, transformations, and workflows to ensure scalability and maintainabilityIn essence: you ensure the organization has accurate, timely, and reliable data powering operational, analytical, and strategic decisions.Key Metrics for Success (KPIs)Pipeline uptime ≥ 99%Data freshness maintained within agreed SLAsZero critical data quality issues reaching downstream reporting systemsImproved warehouse query performance and cost optimizationTimely delivery of scalable and reliable datasetsPositive feedback from analysts, data scientists, and business stakeholdersInterview ProcessInitial Phone ScreenVideo Interview with Pavago RecruiterTechnical Assessment (e.g., build a small ETL pipeline or optimize a SQL query)Client Interview with Engineering/Data TeamOffer & Background Verification#DataEngineer #ETL #DataPipelines #BigQuery #Snowflake #Redshift #Airflow #Python #SQL #CloudData #AnalyticsEngineering #DataInfrastructure #RemoteWork #DataEngineeringJobs","company":"Pavago","rawCompany":"pavago","city":"Lancaster","state":"CA","isRemote":true,"isActive":false,"createdAt":"2026-09-06T14:44:41.277Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Remote Data Engineer","description":"DescriptionJob Title: Data EngineerPosition Type: Full-Time, RemoteWorking Hours: U.S. client business hours (with flexibility for pipeline monitoring, deployments, and data refresh cycles)About the RoleOur client is seeking a Data Engineer to design, build, and maintain scalable data infrastructure and reliable data pipelines that power analytics, reporting, and operational decision-making across the business.This role requires strong software engineering fundamentals, deep experience with modern data stacks, and a passion for building clean, reliable, and high-performance data systems. The Data Engineer will ensure data flows seamlessly from source systems into warehouses, dashboards, and downstream applications while maintaining high standards for quality, governance, and scalability.The ideal candidate is analytical, detail-oriented, and comfortable working across engineering, analytics, and business teams to deliver trustworthy and actionable data.ResponsibilitiesPipeline Development & Data IntegrationBuild, maintain, and optimize ETL/ELT pipelines using Python, SQL, or ScalaOrchestrate workflows using Airflow, Prefect, Dagster, or similar orchestration toolsIngest structured and unstructured data from APIs, SaaS platforms, databases, files, and streaming systemsDevelop scalable connectors and automated ingestion workflowsData Warehousing & ModelingManage and optimize cloud data warehouses such as Snowflake, BigQuery, or RedshiftDesign scalable schemas using star and snowflake modeling techniquesImplement partitioning, clustering, indexing, and performance optimization strategiesBuild clean, analytics-ready datasets for business intelligence and reporting use casesData Quality, Governance & ReliabilityImplement validation checks, anomaly detection, logging, and monitoring to ensure data integrityEnforce naming conventions, lineage tracking, and documentation standards using tools such as dbt or Great ExpectationsMaintain audit-ready data processes and ensure compliance with GDPR, HIPAA, or industry-specific requirementsMonitor pipeline health and proactively resolve failures or inconsistenciesStreaming & Real-Time Data ProcessingBuild and manage real-time data pipelines using Kafka, Kinesis, Pub/Sub, or similar platformsSupport low-latency ingestion and event-driven architectures for time-sensitive applicationsMonitor streaming infrastructure and optimize throughput and reliabilityCollaboration & Analytics EnablementPartner closely with analysts, data scientists, and business stakeholders to deliver reliable datasetsSupport dashboard and reporting initiatives across Tableau, Looker, or Power BITranslate business requirements into scalable data solutions and modelsMaintain clear technical documentation for pipelines, schemas, and workflowsInfrastructure, DevOps & AutomationContainerize data services using Docker and manage deployments through Kubernetes when applicableAutomate deployments using CI/CD pipelines such as GitHub Actions, Jenkins, or GitLab CIManage cloud infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code toolsContinuously optimize performance, scalability, reliability, and cloud costsWhat Makes You a Perfect FitPassionate about building clean, reliable, and scalable data systemsStrong debugging and problem-solving mindset with high attention to detailBalance of software engineering discipline and analytical thinkingComfortable working cross-functionally with technical and non-technical stakeholdersProactive communicator who takes ownership of data quality and reliabilityRequired Experience & Skills3+ years of experience in Data Engineering, Back-End Engineering, or Data Infrastructure rolesStrong proficiency in Python and SQLExperience with at least one modern data warehouse (Snowflake, Redshift, BigQuery)Hands-on experience with orchestration tools such as Airflow or PrefectStrong understanding of ETL/ELT pipelines, data modeling, and data transformation workflowsFamiliarity with cloud platforms such as AWS, GCP, or AzurePreferred Experience & SkillsExperience with dbt for data modeling and transformation managementStreaming and event-driven data pipeline experience (Kafka, Kinesis, Pub/Sub)Experience with cloud-native data services such as AWS Glue, GCP Dataflow, or Azure Data FactoryFamiliarity with Docker, Kubernetes, Terraform, or CI/CD workflowsBackground in regulated industries such as healthcare, fintech, or enterprise SaaSExperience optimizing warehouse costs and query performance at scaleWhat Does a Typical Day Look Like?A Data Engineer's day revolves around maintaining reliable pipelines, improving data quality, and enabling teams with scalable access to trustworthy data. You will:Monitor pipeline health and troubleshoot failed jobs in Airflow or related orchestration systemsBuild and maintain ingestion pipelines for APIs, SaaS platforms, and operational databasesOptimize SQL queries and warehouse performance to improve efficiency and reduce cloud costsCollaborate with analysts and data scientists to provide curated datasets for reporting and modelingImplement validation checks and monitoring to prevent downstream data quality issuesDocument data models, transformations, and workflows to ensure scalability and maintainabilityIn essence: you ensure the organization has accurate, timely, and reliable data powering operational, analytical, and strategic decisions.Key Metrics for Success (KPIs)Pipeline uptime ≥ 99%Data freshness maintained within agreed SLAsZero critical data quality issues reaching downstream reporting systemsImproved warehouse query performance and cost optimizationTimely delivery of scalable and reliable datasetsPositive feedback from analysts, data scientists, and business stakeholdersInterview ProcessInitial Phone ScreenVideo Interview with Pavago RecruiterTechnical Assessment (e.g., build a small ETL pipeline or optimize a SQL query)Client Interview with Engineering/Data TeamOffer & Background Verification#DataEngineer #ETL #DataPipelines #BigQuery #Snowflake #Redshift #Airflow #Python #SQL #CloudData #AnalyticsEngineering #DataInfrastructure #RemoteWork #DataEngineeringJobs","datePosted":"2026-09-06T14:44:41.277Z","dateModified":"2026-09-06T14:44:41.277Z","hiringOrganization":{"@type":"Organization","name":"Pavago","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Lancaster","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"04ba130bca982696e8f83cef"},"url":"https://jobsearcher.com/jobs/04ba130bca982696e8f83cef"}}