Senior Databricks Data Engineer
Senior Databricks Data EngineerJob SummaryWe are seeking an experienced Senior Databricks Data Engineer with 7+ years of hands-on experience designing, developing, and optimizing scalable data engineering and analytics solutions. The ideal candidate will have strong expertise in Databricks, PySpark, Spark SQL, and Apache Spark, with proven experience building enterprise-grade ETL pipelines across AWS and cloud data platforms.The candidate will be responsible for developing high-performance data pipelines, implementing data quality and validation frameworks, optimizing large-scale data processing workloads, and supporting business-critical analytics, reporting, and regulatory use cases.Key ResponsibilitiesDesign, develop, and maintain scalable end-to-end data engineering pipelines using Databricks, PySpark, Spark SQL, and Apache Spark.Build and optimize ETL/ELT pipelines integrating AWS S3, Snowflake, AWS Glue, and Apache Airflow.Process and transform terabyte-scale structured and semi-structured datasets for enterprise reporting, analytics, and regulatory compliance.Develop reusable and modular PySpark frameworks for data ingestion, cleansing, transformation, validation, enrichment, and implementation of complex business rules.Implement robust data quality, reconciliation, validation, and exception-handling frameworks to ensure data accuracy and reliability.Optimize Databricks and Spark workloads through partitioning, caching, efficient joins, file optimization, query tuning, and resource optimization.Develop production-ready data pipelines with appropriate monitoring, logging, error handling, and operational controls.Work with Snowflake for data warehousing, data integration, transformation, and analytical workloads.Integrate data from multiple AWS and enterprise data sources into centralized analytics and reporting platforms.Collaborate with data analysts, architects, business stakeholders, and application teams to translate business requirements into scalable data solutions.Support use cases including customer segmentation, campaign analytics, AML/regulatory reporting, production data processing, and advanced data validation.Troubleshoot and resolve pipeline failures, data discrepancies, performance issues, and production incidents.Establish and follow best practices for data engineering, coding standards, performance optimization, scalability, and maintainability.Participate in design reviews, code reviews, testing, deployment, and production support activities.Required Skills & Qualifications7+ years of experience in data engineering, with significant hands-on experience with Databricks.Strong expertise in PySpark, Apache Spark, and Spark SQL.Strong experience developing ETL/ELT data pipelines and large-scale data processing solutions.Hands-on experience with AWS S3, AWS Glue, and Apache Airflow.Strong experience working with Snowflake and cloud-based data platforms.Experience processing large-scale/terabyte-level datasets.Strong understanding of data ingestion, transformation, cleansing, validation, enrichment, and data quality concepts.Experience implementing reusable PySpark/data engineering frameworks.Strong understanding of Spark performance tuning and optimization.Experience working with both structured and semi-structured data.Strong troubleshooting and production support experience.Experience working in enterprise data environments with reporting, analytics, compliance, or regulatory requirements.Strong analytical, problem-solving, and communication skills.Preferred QualificationsExperience with Databricks Delta Lake / Delta tables and modern lakehouse architectures.Experience with Databricks Workflows, Unity Catalog, and production job orchestration.Knowledge of AWS cloud architecture and services.Experience with CI/CD, Git, and DevOps practices for data engineering.Experience with AML, regulatory reporting, financial services, or customer analytics use cases.Familiarity with data governance, metadata management, and enterprise data quality practices.Key TechnologiesDatabricks | PySpark | Apache Spark | Spark SQL | AWS S3 | AWS Glue | Snowflake | Apache Airflow | ETL/ELT | Delta Lake | Python | SQL | Data Quality | Data Validation | Performance Tuning | AWS