Data Bricks Migration and Support engineer
Role: Data Bricks Migration and Support engineer
8 - 15 years of experience
Job Description
Must Have Technical/Functional Skills
Successfully executed a data migration or modernization to Data Bricks, preferably IBM Data Stage to Data Bricks on AWS
Should have Experience in handling Large Migrations to Data Bricks.
Should have good analytical skills to compare the legacy and modern data platform end to end right from source to target.
Good understanding of DataBricks implementation of Medallion layer architecture.
Independently Lead and Managed large Data Bricks migrations.
CI/CD Integration: Implement version control (e.g., Git) and automated deployment processes for Databricks assets
Technical and architectural skills required are below.
Core Data Engineering Languages
Experience in Advanced SQL for building modular analytics workflows, utilizing advanced Common Table Expressions (CTEs), and writing high-performance queries inside Data Bricks SQL Analytics.
Experience in Python or Scala to build, optimize, and debug complex data transformation scripts, custom functions, and machine learning pipelines.
Big Data & Architecture Core
Experience in Apache Spark Ecosystem for understanding cluster execution flow, memory allocation, driver/worker nodes, and handling data frames.
Experience in Delta Lake Architecture to understand ACID transactions on object storage, data skipping, partition strategies, and automated data compaction.
Databricks Platform Expertise
Experience in Delta Live Tables (DLT) & Workflows for constructing and orchestrating production-ready, declarative streaming, and batch ETL pipelines.
Experience in Unity Catalog for setting up data governance, column/row-level access control, and tracking end-to-end data lineage across workspaces.
Experience in Auto Loader for implementing modern, incremental data ingestion patterns from cloud blob storage into the lakehouse.
Code Translation & Refactoring
Pipeline Conversion: Translate visual DataStage Parallel Jobs and Sequences into Python/PySpark scripts or Data bricks Notebooks
Legacy Refactoring: Modernize legacy logic rather than applying "lift and shift" anti-patterns; adapt workflows to think in distributed DataFrames rather than DataStage stages.
Logic Mapping: Map DataStage components—such as Aggregators, Joiners, Transformers, and Sort stages—to equivalent Spark operations
Validation & Reconciliation: Build automated reconciliation frameworks to compare row counts, checksums, and aggregate sums between legacy DataStage outputs and new Databricks output
Data Cleansing: Identify and resolve data type discrepancies, null-handling differences, and encoding issues during the extraction and loading phases
Platform Orc hestration & Governance
Orchestration: Replace DataStage sequence jobs with Databricks workflows ( or external orchestrators like Azure Data Factory/Airflow) to schedule and manage dependencies
Data Governance: Enforce data lineage, security, and cataloging using Unity Catalog to ensure compliance in the new Lakehouse environment.
Cloud Providers (AWS): Understanding underlying cloud object storage , identity access management (IAM), and network security configurations.
DevOps & Bundles: Familiarity with Databricks Asset Bundles (DABs) and CI/CD tools to automate the deployment of workspaces and pipeline assets.
Legacy Assessment & Migration Mechanics
Code Conversion & Translation: The ability to parse legacy code structures and refactor them into Databricks-native code.
AI-Assisted Migration: Skills in using AI coding assistants and open framework agent tools to analyze application interdependencies, automate schema mapping, and accelerate lift-and-shift workloads
Code Conversion & Translation: The ability to parse legacy code structures from ETL pipelines, Informatica, data Stage preferred
Experience working in Agile teams and understanding of data governance frameworks.
J-18808-Ljbffr