JOBSEARCHER

Sr. Data Engineer

Role : Sr. Data Engineer Location: Pasadena, CAWork Arrangement: Hybrid – 3 Days/Week OnsiteJob SummaryWe are looking for an experienced Data Engineer with strong expertise in Databricks, PySpark, and Python to design, develop, and maintain scalable data engineering solutions. The ideal candidate will have hands-on experience building ETL/ELT pipelines, data processing frameworks, and data lake/lakehouse solutions using Databricks and cloud technologies.Key ResponsibilitiesDesign, develop, and maintain scalable data pipelines using Databricks, PySpark, and Python.Develop ETL/ELT workflows to ingest, transform, cleanse, and integrate data from multiple sources.Build and optimize data processing jobs using PySpark and Spark SQL.Work extensively with Databricks Lakehouse, Delta Lake, notebooks, workflows, and clusters.Develop reusable Python modules and frameworks for data processing and automation.Implement data quality checks, validation, error handling, and monitoring within data pipelines.Optimize Spark jobs, including partitioning, caching, joins, and performance tuning.Work with Delta Lake for data storage, transformation, versioning, and incremental processing.Integrate data from relational databases, APIs, files, cloud storage, and other enterprise data sources.Collaborate with Data Architects, Data Scientists, BI Developers, and business stakeholders to understand data requirements.Implement CI/CD and source-control practices for data engineering code.Troubleshoot production data pipeline failures and perform root cause analysis (RCA).Ensure data security, governance, lineage, and compliance requirements are followed.Participate in design discussions, code reviews, testing, deployment, and production support.Required SkillsStrong hands-on experience with DatabricksStrong PySpark / Apache Spark experienceStrong Python programming skillsExperience developing ETL/ELT pipelinesStrong SQL skillsExperience with Delta LakeExperience with data lake/lakehouse architectureExperience with Spark performance tuning and optimizationExperience working with large-volume datasetsStrong understanding of data modeling and data engineering conceptsExperience with Git and CI/CDExperience with cloud platforms such as AWS, Azure, or GCP