Lead Data Platform enginner
Job Summary
We are seeking an experienced *Lead Data Platform Engineer* to lead the design, development, and optimization of an enterprise-scale data platform. The ideal candidate will drive technical strategy, establish engineering best practices, and collaborate with cross-functional teams to deliver scalable, secure, and high-performing data solutions across the *Hadoop ecosystem and Azure, including Databricks*.
Key Responsibilities
* Lead the architecture, development, and enhancement of large-scale data platforms supporting analytics, machine learning, and operational workloads.
* Design and optimize big data solutions using Hadoop ecosystem technologies including *HDFS, Hive, Spark, YARN, and related components*.
* Develop and manage data pipelines and transformations using *Azure Data Lake Storage, Azure Data Factory, Azure Synapse, and Azure Databricks*.
* Implement and enforce robust *data governance, security, and data quality frameworks* across all data layers.
* Partner with data engineering, analytics, product, and infrastructure teams to translate business requirements into scalable technical solutions.
* Drive *performance tuning, capacity planning, and cost optimization* across on-premises and cloud-based data platforms.
* Mentor and technically guide data engineers while establishing engineering standards, reusable patterns, and best practices.
* Oversee *CI/CD, deployment, and monitoring* for data workflows.
* Evaluate emerging technologies and contribute to long-term *data platform modernization and technology strategy*.
Required Qualifications
* *8+ years* of experience in data engineering or data platform roles, including *3+ years in a technical lead or architect capacity*.
* Strong hands-on experience with the *Hadoop ecosystem*, including HDFS, Hive, Spark, Oozie, Ranger, Airflow, and related technologies.
* Deep expertise in *Azure Data Services*, including:
* Azure Data Lake Storage
* Azure Data Factory
* Azure Synapse
* Azure Functions
* Azure Key Vault
* Advanced hands-on experience with *Databricks*, including:
* Apache Spark optimization
* Delta Lake
* Unity Catalog
* Strong proficiency in *Python and SQL* and distributed data processing frameworks.
* Experience with *DevOps, CI/CD pipelines, and Infrastructure as Code*, such as Terraform or ARM.
* Strong understanding of *data modeling, storage formats, and data governance*, including Parquet, ORC, and Delta.
* Proven ability to lead technical teams, communicate effectively, and influence architecture decisions.
Preferred Qualifications
* Experience migrating *on-premises Hadoop workloads to cloud platforms*, preferably Azure Databricks.
* Knowledge of *real-time data processing technologies*, including:
* Kafka
* Azure Event Hubs
* Spark Streaming
Pay: $55.00 - $65.00 per hour
Work Location: In person