Data Bricks Migration and Support engineer
Data Bricks Migration and Support EngineerLocation: Seattle, WA / Bellevue, WA / Everett, WA / Renton, WA / Richardson, TX / Plano, TX / Dallas, TX / St. Louis, MO / Charleston, SC / Arlington, VA (Onsite) FulltimeMust Have Technical/Functional SkillsSuccessfully executed a data migration or modernization to Data Bricks, preferably IBM Data Stage to Data Bricks on AWSExperience in handling large migrations to Data BricksGood analytical skills to compare the legacy and modern data platform end to end right from source to targetGood understanding of DataBricks implementation of Medallion layer architectureIndependently lead and managed large Data Bricks migrationsCI/CD Integration: Implement version control (e.g., Git) and automated deployment processes for Databricks assetsTechnical and Architectural Skills RequiredCore Data Engineering LanguagesExperience in Advanced SQL for building modular analytics workflows, utilizing advanced Common Table Expressions (CTEs), and writing high-performance queries inside Data Bricks SQL AnalyticsExperience in Python or Scala to build, optimize, and debug complex data transformation scripts, custom functions, and machine learning pipelinesBig Data & Architecture CoreExperience in Apache Spark Ecosystem for understanding cluster execution flow, memory allocation, driver/worker nodes, and handling data framesExperience in Delta Lake Architecture to understand ACID transactions on object storage, data skipping, partition strategies, and automated data compactionDatabricks Platform ExpertiseExperience in Delta Live Tables (DLT) & Workflows for constructing and orchestrating production-ready, declarative streaming, and batch ETL pipelinesExperience in Unity Catalog for setting up data governance, column/row-level access control, and tracking end-to-end data lineage across workspacesExperience in Auto Loader for implementing modern, incremental data ingestion patterns from cloud blob storage into the lakehouseCode Translation & RefactoringPipeline Conversion: Translate visual DataStage Parallel Jobs and Sequences into Python/PySpark scripts or Data bricks NotebooksLegacy Refactoring: Modernize legacy logic rather than applying "lift and shift" anti-patterns; adapt workflows to think in distributed DataFrames rather than DataStage stagesLogic Mapping: Map DataStage components—such as Aggregators, Joiners, Transformers, and Sort stages—to equivalent Spark operationsTesting & ReconciliationValidation & Reconciliation: Build automated reconciliation frameworks to compare row counts, checksums, and aggregate sums between legacy DataStage outputs and new Databricks outputData Cleansing: Identify and resolve data type discrepancies, null-handling differences, and encoding issues during the extraction and loading phasesPlatform Orchestration & GovernanceOrchestration: Replace DataStage sequence jobs with Databricks workflows (or external orchestrators like Azure Data Factory/Airflow) to schedule and manage dependenciesData Governance: Enforce data lineage, security, and cataloging using Unity Catalog to ensure compliance in the new Lakehouse environmentGood To HaveCloud Infrastructure & CI/CDCloud Providers (AWS): Understanding underlying cloud object storage, identity access management (IAM), and network security configurationsDevOps & Bundles: Familiarity with Databricks Asset Bundles (DABs) and CI/CD tools to automate the deployment of workspaces and pipeline assetsLegacy Assessment & Migration MechanicsCode Conversion & Translation: The ability to parse legacy code structures and refactor them into Databricks-native codeAI-Assisted Migration: Skills in using AI coding assistants and open framework agent tools to analyze application interdependencies, automate schema mapping, and accelerate lift-and-shift workloadsExperience working in Agile teams and understanding of data governance frameworksResponsibilitiesSupport post-migration environment from IBM DataStage to DatabricksIncident & Lifecycle ManagementCI/CD Deployment: Support code deployments across Development, Test, and Production environments using Databricks Repos and REST APIsMonitoring & Alerting: Set up monitoring via Databricks System Tables and observability tools to catch job failures, data anomalies, or latency spikes earlyPipeline Maintenance & OrchestrationWorkflow Management: Transition from DataStage job sequences to native data bricks workflows for scheduling, dependency tracking, and alertsETL Refactoring: Troubleshoot and fix issues in generated PySpark or Spark SQL code that replaced legacy DataStage Transformer or Lookup stagesStreaming & Batch Integration: Support ongoing data ingestion using data bricks autoloader to process files continuously from cloud storagePerformance Tuning & Cost OptimizationCompute Management: Monitor and configure serverless or classic clusters to prevent over-provisioningQuery Optimization: Analyze Spark execution plans