Engineer - Data Engineer (Remote)
Job Title: Data Engineer Position Type: Full-Time, RemoteS. client business hours (with flexibility for pipeline monitoring, deployments, and data refresh cycles)About the Role Our client is seeking a Data Engineer to design, build, and maintain scalable data infrastructure and reliable data pipelines that power analytics, reporting, and operational decision-making across the business.This role requires strong software engineering fundamentals, deep experience with modern data stacks, and a passion for building clean, reliable, and high-performance data systems. The Data Engineer will ensure data flows seamlessly from source systems into warehouses, dashboards, and downstream applications while maintaining high standards for quality, governance, and scalability.The ideal candidate is analytical, detail-oriented, and comfortable working across engineering, analytics, and business teams to deliver trustworthy and actionable data.Responsibilities Pipeline Development & Data Integration • Build, maintain, and optimize ETL/ELT pipelines using Python, SQL, or ScalaIngest structured and unstructured data from APIs, SaaS platforms, databases, files, and streaming systemsData Warehousing & Modeling • Manage and optimize cloud data warehouses such as Snowflake, BigQuery, or RedshiftImplement partitioning, clustering, indexing, and performance optimization strategiesBuild clean, analytics-ready datasets for business intelligence and reporting use casesData Quality, Governance & Reliability • Implement validation checks, anomaly detection, logging, and monitoring to ensure data integrityMaintain audit-ready data processes and ensure compliance with GDPR, HIPAA, or industry-specific requirementsMonitor pipeline health and proactively resolve failures or inconsistenciesStreaming & Real-Time Data Processing • Build and manage real-time data pipelines using Kafka, Kinesis, Pub/Sub, or similar platformsSupport low-latency ingestion and event-driven architectures for time-sensitive applicationsCollaboration & Analytics Enablement • Partner closely with analysts, data scientists, and business stakeholders to deliver reliable datasetsSupport dashboard and reporting initiatives across Tableau, Looker, or Power BITranslate business requirements into scalable data solutions and modelsMaintain clear technical documentation for pipelines, schemas, and workflowsInfrastructure, DevOps & Automation • Containerize data services using Docker and manage deployments through Kubernetes when applicableAutomate deployments using CI/CD pipelines such as GitHub Actions, Jenkins, or GitLab CIManage cloud infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code toolsContinuously optimize performance, scalability, reliability, and cloud costsWhat Makes You a Perfect Fit • Passionate about building clean, reliable, and scalable data systemsBalance of software engineering discipline and analytical thinkingProactive communicator who takes ownership of data quality and reliabilityRequired Experience & Skills • 3+ years of experience in Data Engineering, Back-End Engineering, or Data Infrastructure rolesStrong proficiency in Python and SQLExperience with at least one modern data warehouse (Snowflake, Redshift, BigQuery)Strong understanding of ETL/ELT pipelines, data modeling, and data transformation workflowsFamiliarity with cloud platforms such as AWS, GCP, or AzurePreferred Experience & Skills • Experience with dbt for data modeling and transformation managementStreaming and event-driven data pipeline experience (Kafka, Kinesis, Pub/Sub)Experience with cloud-native data services such as AWS Glue, GCP Dataflow, or Azure Data FactoryExperience optimizing warehouse costs and query performance at scaleA Data Engineer's day revolves around maintaining reliable pipelines, improving data quality, and enabling teams with scalable access to trustworthy data. Monitor pipeline health and troubleshoot failed jobs in Airflow or related orchestration systemsBuild and maintain ingestion pipelines for APIs, SaaS platforms, and operational databasesOptimize SQL queries and warehouse performance to improve efficiency and reduce cloud costsCollaborate with analysts and data scientists to provide curated datasets for reporting and modelingImplement validation checks and monitoring to prevent downstream data quality issuesDocument data models, transformations, and workflows to ensure scalability and maintainabilityIn essence: you ensure the organization has accurate, timely, and reliable data powering operational, analytical, and strategic decisions.Key Metrics for Success (KPIs) • Pipeline uptime ≥ 99%Data freshness maintained within agreed SLAsZero critical data quality issues reaching downstream reporting systemsImproved warehouse query performance and cost optimizationPositive feedback from analysts, data scientists, and business stakeholdersInterview Process • Initial Phone ScreenVideo Interview with Pavago Recruiterbuild a small ETL pipeline or optimize a SQL query)Client Interview with Engineering/Data TeamDataEngineer #ETL #DataPipelines #BigQuery #Snowflake #Redshift #Airflow #Python #SQL #CloudData #AnalyticsEngineering #DataInfrastructure #RemoteWork #DataEngineeringJobs