{"schemaVersion":"jobsearcher.job.v1","id":"5c8ae15a5ed0238c33120087","url":"https://jobsearcher.com/jobs/5c8ae15a5ed0238c33120087","canonicalUrl":"https://jobsearcher.com/jobs/5c8ae15a5ed0238c33120087","title":"Sr. Data Engineer","description":"Role : Sr. Data EngineerLocation: Pasadena, CAWork Arrangement: HybridJob SummaryWe are looking for an experienced Data Engineer with strong expertise in Databricks, PySpark, and Python to design, develop, and maintain scalable data engineering solutions. The ideal candidate will have hands‑on experience building ETL/ELT pipelines, data processing frameworks, and data lake/lakehouse solutions using Databricks and cloud technologies.\r\nKey ResponsibilitiesDesign, develop, and maintain scalable data pipelines using Databricks, PySpark, and Python.\r\nDevelop ETL/ELT workflows to ingest, transform, cleanse, and integrate data from multiple sources.\r\nBuild and optimize data processing jobs using PySpark and Spark SQL.\r\nWork extensively with Databricks Lakehouse, Delta Lake, notebooks, workflows, and clusters.\r\nDevelop reusable Python modules and frameworks for data processing and automation.\r\nImplement data quality checks, validation, error handling, and monitoring within data pipelines.\r\nOptimize Spark jobs, including partitioning, caching, joins, and performance tuning.\r\nWork with Delta Lake for data storage, transformation, versioning, and incremental processing.\r\nIntegrate data from relational databases, APIs, files, cloud storage, and other enterprise data sources.\r\nCollaborate with Data Architects, Data Scientists, BI Developers, and business stakeholders to understand data requirements.\r\nImplement CI/CD and source-control practices for data engineering code.\r\nTroubleshoot production data pipeline failures and perform root cause analysis (RCA).\r\nEnsure data security, governance, lineage, and compliance requirements are followed.\r\nParticipate in design discussions, code reviews, testing, deployment, and production support.\r\nRequired SkillsStrong hands‑on experience with Databricks\r\nStrong PySpark / Apache Spark experience\r\nStrong Python programming skills\r\nExperience developing ETL/ELT pipelines\r\nStrong SQL skills\r\nExperience with Delta Lake\r\nExperience with data lake/lakehouse architecture\r\nExperience with Spark performance tuning and optimization\r\nExperience working with large-volume datasets\r\nStrong understanding of data modeling and data engineering concepts\r\nExperience with Git and CI/CD\r\nExperience with cloud platforms such as AWS, Azure, or GCP\r\nPreferred SkillsDatabricks certification\r\nExperience with Azure Data Factory / AWS Glue / Airflow\r\nExperience with Azure Data Lake / Amazon S3\r\nExperience with Unity Catalog\r\nExperience with Kafka or other streaming technologies\r\nExperience with Terraform\r\nExperience with data governance and data quality frameworks\r\nExperience with Power BI, Tableau, or other BI platforms\r\nTypical Technology StackDatabricks | PySpark | Python | Spark SQL | Delta Lake | SQL | AWS/Azure | Data Lake | Git | CI/CD | Airflow/ADF/Glue#J-18808-Ljbffr","company":"Techgene Solutions","rawCompany":"techgene solutions","city":"Pasadena","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-10-05T01:33:04.443Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. Data Engineer","description":"Role : Sr. Data EngineerLocation: Pasadena, CAWork Arrangement: HybridJob SummaryWe are looking for an experienced Data Engineer with strong expertise in Databricks, PySpark, and Python to design, develop, and maintain scalable data engineering solutions. The ideal candidate will have hands‑on experience building ETL/ELT pipelines, data processing frameworks, and data lake/lakehouse solutions using Databricks and cloud technologies.\r\nKey ResponsibilitiesDesign, develop, and maintain scalable data pipelines using Databricks, PySpark, and Python.\r\nDevelop ETL/ELT workflows to ingest, transform, cleanse, and integrate data from multiple sources.\r\nBuild and optimize data processing jobs using PySpark and Spark SQL.\r\nWork extensively with Databricks Lakehouse, Delta Lake, notebooks, workflows, and clusters.\r\nDevelop reusable Python modules and frameworks for data processing and automation.\r\nImplement data quality checks, validation, error handling, and monitoring within data pipelines.\r\nOptimize Spark jobs, including partitioning, caching, joins, and performance tuning.\r\nWork with Delta Lake for data storage, transformation, versioning, and incremental processing.\r\nIntegrate data from relational databases, APIs, files, cloud storage, and other enterprise data sources.\r\nCollaborate with Data Architects, Data Scientists, BI Developers, and business stakeholders to understand data requirements.\r\nImplement CI/CD and source-control practices for data engineering code.\r\nTroubleshoot production data pipeline failures and perform root cause analysis (RCA).\r\nEnsure data security, governance, lineage, and compliance requirements are followed.\r\nParticipate in design discussions, code reviews, testing, deployment, and production support.\r\nRequired SkillsStrong hands‑on experience with Databricks\r\nStrong PySpark / Apache Spark experience\r\nStrong Python programming skills\r\nExperience developing ETL/ELT pipelines\r\nStrong SQL skills\r\nExperience with Delta Lake\r\nExperience with data lake/lakehouse architecture\r\nExperience with Spark performance tuning and optimization\r\nExperience working with large-volume datasets\r\nStrong understanding of data modeling and data engineering concepts\r\nExperience with Git and CI/CD\r\nExperience with cloud platforms such as AWS, Azure, or GCP\r\nPreferred SkillsDatabricks certification\r\nExperience with Azure Data Factory / AWS Glue / Airflow\r\nExperience with Azure Data Lake / Amazon S3\r\nExperience with Unity Catalog\r\nExperience with Kafka or other streaming technologies\r\nExperience with Terraform\r\nExperience with data governance and data quality frameworks\r\nExperience with Power BI, Tableau, or other BI platforms\r\nTypical Technology StackDatabricks | PySpark | Python | Spark SQL | Delta Lake | SQL | AWS/Azure | Data Lake | Git | CI/CD | Airflow/ADF/Glue#J-18808-Ljbffr","datePosted":"2026-10-05T01:33:04.443Z","dateModified":"2026-10-05T01:33:04.443Z","hiringOrganization":{"@type":"Organization","name":"Techgene Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Pasadena","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"5c8ae15a5ed0238c33120087"},"url":"https://jobsearcher.com/jobs/5c8ae15a5ed0238c33120087"}}