Lead Software Engineer - Data Engineer
Overview
As Lead Software Engineer in the CIB Regulatory Reporting Team, you will drive data-driven solution design and production-grade delivery at scale. You will contribute technically across multiple domains, focusing on data engineering and Spark-based ETL/ELT to support secure, reliable market-leading technology. You’ll promote AI-assisted engineering practices, uphold secure coding and automation standards, and lead architectural discussions with internal and external stakeholders. This role offers impact through improving data pipelines, lakehouse operations, and overall platform robustness.
Compensation / Benefitscompetitive total rewardshealthcare coverageretirement savings planbackup childcaretuition reimbursementmental health support
ResponsibilitiesDesign and develop innovative software solutions with emphasis on data engineering and Spark-based ETL/ELTProduce high-quality Python/PySpark production code and review peers' Spark jobs and data logicDrive adoption of enterprise AI-assisted development practices and establish validation standardsLeverage SDLC tools to improve automation, testing, and delivery velocityIdentify automation opportunities to boost reliability and operational stability of data pipelinesLead architectural evaluation sessions with vendors and internal teams on systems and data architectureFoster Communities of Practice to disseminate knowledge on Spark performance, Iceberg, and data opsCultivate an inclusive, diverse team culture and collaborative engineering environment
Key requirements5+ years of applied experience in data engineering or software engineeringStrong Python with hands-on PySpark experienceAdvanced Spark SQL skills and solid SQL fundamentalsExperience delivering large-scale data pipelines with design, development, testing, and operationsExperience with AI-assisted software development tools and safe adoption practicesUnderstanding of secure and compliant engineering workflows including data sensitivity and resiliencyExperience with AWS data services (S3, AWS Glue Data Catalog) or similar cloud data platformsSpark workloads on EMR or Databricks with tuning and monitoringIceberg production experience (table design, partitioning, schema evolution, compaction) and related opsProficiency in automation, CI/CD, and repeatable deployments for data pipelinesleadership and mentoringcross-functional collaborationstrong communicationPythonPySparkSpark SQL