AWS Data Engineer
Overview
In this role you will design and maintain data integration pipelines on AWS, enabling timely data availability for analytics. You will work with cross-functional teams to translate requirements into scalable data workflows, optimize Redshift performance, and ensure data quality. You will leverage Glue, EMR, and Airflow to process large datasets and support BI insights. This is a hands-on, on-site position focused on building robust data platforms in a cloud-native environment.
ResponsibilitiesDesign and implement data integration workflows using AWS Glue, EMR, MWAA (Airflow), Lambda, and RedshiftDevelop proficiency in PySpark, Apache Spark, and Python for large-scale data processingEnsure accurate extraction, transformation, and loading of data into target systemsValidate and cleanse data to maintain high data quality and implement monitoring and error handlingOptimize data workflows for SLA adherence, scalability, and cost-efficiency on AWSIdentify performance bottlenecks, tune queries, and improve Redshift performanceRegularly review and refine integration processes to boost efficiencyCollaborate with data analysts and business stakeholders to meet data requirements
Key requirements7+ years in data engineering, database design, ETL, and data warehousing7+ years with AWS tools and technologies (S3, EMR, Glue, Athena, Redshift, RDS, Spectrum, Airflow)7+ years with CI/CD toolsStrong knowledge of data storage and processing technologies (Hadoop, Spark)7+ years programming in PythonAWS GlueEMRMWAA (Airflow)LambdaRedshiftS3