Senior Data Engineer
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
Role: Sr. Data Engineer – DatabricksExperience: 8 - 15 yearsLocation: San Ramon, CA Job Description:We are seeking a highly skilled Data Engineer with deep Databricks expertise to design, build, and optimize scalable, cloud-based data platforms and analytics solutions. The ideal candidate will have strong hands-on experience with distributed data processing, modern data architectures, and performance optimization across large-scale datasets.Key ResponsibilitiesDesign, develop, and maintain end-to-end data pipelines using Databricks (Spark) for batch and streaming workloads.Build and optimize ETL/ELT frameworks leveraging Delta Lake, medallion architecture (Bronze/Silver/Gold), and scalable data models.Implement data ingestion from diverse sources including relational databases, APIs, event streams, and cloud storage.Develop robust data transformations using PySpark/Scala, ensuring performance, reliability, and scalability.Optimize Spark jobs through partitioning, caching, broadcast joins, and cluster tuning.Implement data quality checks, validations, and monitoring to ensure accuracy and consistency.Work closely with analytics, data science, and business teams to enable self-service analytics and reporting.Support data governance, security, and access control using cloud-native and Databricks capabilities.Troubleshoot pipeline failures, optimize costs, and improve overall platform efficiency.Required & Key SkillsDatabricks Certified Data Engineer (Mandatory)Advanced SQL skills and experience in data modeling (dimensional & analytical models)Hands-on experience building large-scale ETL/ELT pipelinesDeep experience with Databricks features: Delta Lake, Unity Catalog, Workflows, and NotebooksStrong experience with cloud platforms – AWS, Azure, or GCP (storage, compute, networking)Experience with data orchestration tools such as Airflow, Azure Data Factory, or similarKnowledge of streaming frameworks (Spark Structured Streaming, Kafka/Event Hubs – preferred)Experience in performance tuning, cost optimization, and data reliability engineeringStrong understanding of data security, governance, and compliance best practicesExcellent communication, analytical, and problem-solving skills