{"schemaVersion":"jobsearcher.job.v1","id":"d5b0265f4d7cd7308b696b75","url":"https://jobsearcher.com/jobs/d5b0265f4d7cd7308b696b75","canonicalUrl":"https://jobsearcher.com/jobs/d5b0265f4d7cd7308b696b75","title":"Databricks Data Engineer","description":"Responsibilities:\nThe Databricks Data Engineer Core AI & Data practice helps organizations modernize data platforms, strengthen enterprise data foundations, and scale analytics and artificial intelligence capabilities across the business. The team works with clients to architect, engineer, and deploy cloud-based data solutions that improve decision-making, enable innovation, and support large-scale transformation. As a hands-on Databricks Data Engineer with deep expertise in Azure/AWS Databricks (AKS/EKS as a backbone) and MLOps, this role will have the opportunity to migrate and translate legacy SSIS ETL logic into scalable, cloud-native data pipelines in Databricks. This role will partner with data engineers, data scientists, and product manager to design features, train/evaluate models, and deploy them to production using MLflow, Databricks and Workflows—with rigorous observability, governance (Unity Catalog), and CI/CD automation.\nData Pipeline Engineering\nDesign, build, and maintain high-performance, scalable ETL/ELT pipelines using Azure Databricks, Delta Lake, and PySpark.\nConvert and modernize existing SSIS package logic into cloud-native Databricks pipelines using PySpark notebooks, Delta Live Tables (DLT), and Databricks Workflows.\nImplement reliable batch and streaming pipelines with robust data quality and validation frameworks.\nOptimize pipeline performance using Photon, efficient file formats, partitioning, Z-ordering, and caching strategies.\nLakehouse Platform Development\nDevelop and manage datasets within Delta Lake, ensuring ACID reliability, schema evolution, versioning, and time travel.\nArchitect feature-rich data layers including:\nBronze (raw ingestion)\nSilver (validated, conformed)\nGold (analytics-ready and ML-ready)\nImplement data governance using Unity Catalog for fine-grained access control, lineage, auditability, and metadata management.\nMLOps & ML-Enabled Data Pipelines\nPartner with data scientists and data engineers to create feature pipelines, model training pipelines, and production scoring pipelines.\nDeploy and operationalize models using MLflow, Databricks Model Registry, and Databricks Workflows.\nUse Databricks built-in AI SQL functions such as ai_query, ai_forecast, ai_analyze_sentiment to generate actionable insight from large amount of unstructured or structured raw data\nImplement monitoring for:\nPipeline failures\nData/feature drift\nModel performance degradation\nOperational SLAs/SLIs/SLOs\nBuild automated CI/CD workflows using GitHub Actions or Azure DevOps for notebook deployment, pipeline testing, and environment promotion.\nData Platform, Data Security & Data Governance\nCollaborate with data engineers to design reliable data products on Delta Lake; leverage Delta Live Tables (DLT) for declarative pipelines when applicable.\nEnforce Unity Catalog for lineage, permissions, and audit; manage secrets, tokens, and keys securely (e.g., Databricks secrets, Key Vault/Secrets Manager).\nCollaboration & Leadership\nWork closely with cross-functional teams: data engineering, data scientist, product manager, and business stakeholders.\nServe as a Databricks SME—championing best practices, code standards, governance, and reusable frameworks.\nDocument architecture, workflows, data models, runbooks, and operational procedures.\nQualifications\nMinimum of 5 years of experience in Databricks, PySpark notebooks, Python, DevOps, software development, and data engineering.\nCertified Databricks Data Engineer Associate or Professional is a plus.\nSkills & Competencies\nProficient in designing, building, deploying, and maintaining high-performance, scalable ETL/ELT pipelines using Azure Databricks, Delta Lake, and PySpark Notebook.\nProficient in building, deploying, and operating production ML models such as supervised, unsupervised, and anomaly detection, including techniques for imbalanced datasets\nExperience working with EKS/AKS cluster and containerized patform.\nProficient with ML engineering and MLOps, including model versioning, CI/CD for ML, monitoring, drift detection, and automated retraining\nProficiency in Python including Pandas and PySpark Data frames\nExpert level of SQL skills including Stored Procedure, experience with SSIS, SSRS, Power BI is a plus.\nProficient with cloud data engineering platforms, such as Azure, Databricks, Spark, or SQL, and batch and streaming pipelines\nFamiliar with Databricks AI Built-In Functions such as AI_Query, AI_Gen, AI_Classify, AI_Forecast, AI_Analyze_Sentiment, able to use them to extract actionable insights from large amount of unstructured or structured raw data\nExperience with Python and ML frameworks, such as PyTorch or TensorFlow\nExperience in improving data quality, lineage, and observability in enterprise data environments and operationalizing rules and model-driven scoring for prioritization, routing, or case selection\nExperience with predictive analytics, machine learning, and artificial intelligence desired.\nBachelor’s degree in Computer Science/ Management Information Systems/ Engineering or related field.\nPay: $120,000.00 - $180,000.00 per year\nBenefits:\n401(k)\nDental insurance\nHealth insurance\nVision insurance\nWork Location: Hybrid remote in Aldie, VA 20105","company":"Cyberdash Cryptometrics","rawCompany":"cyberdash cryptometrics","city":"Chantilly","state":"VA","isRemote":false,"isActive":false,"createdAt":"2026-08-03T14:08:47.637Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Databricks Data Engineer","description":"Responsibilities:\nThe Databricks Data Engineer Core AI & Data practice helps organizations modernize data platforms, strengthen enterprise data foundations, and scale analytics and artificial intelligence capabilities across the business. The team works with clients to architect, engineer, and deploy cloud-based data solutions that improve decision-making, enable innovation, and support large-scale transformation. As a hands-on Databricks Data Engineer with deep expertise in Azure/AWS Databricks (AKS/EKS as a backbone) and MLOps, this role will have the opportunity to migrate and translate legacy SSIS ETL logic into scalable, cloud-native data pipelines in Databricks. This role will partner with data engineers, data scientists, and product manager to design features, train/evaluate models, and deploy them to production using MLflow, Databricks and Workflows—with rigorous observability, governance (Unity Catalog), and CI/CD automation.\nData Pipeline Engineering\nDesign, build, and maintain high-performance, scalable ETL/ELT pipelines using Azure Databricks, Delta Lake, and PySpark.\nConvert and modernize existing SSIS package logic into cloud-native Databricks pipelines using PySpark notebooks, Delta Live Tables (DLT), and Databricks Workflows.\nImplement reliable batch and streaming pipelines with robust data quality and validation frameworks.\nOptimize pipeline performance using Photon, efficient file formats, partitioning, Z-ordering, and caching strategies.\nLakehouse Platform Development\nDevelop and manage datasets within Delta Lake, ensuring ACID reliability, schema evolution, versioning, and time travel.\nArchitect feature-rich data layers including:\nBronze (raw ingestion)\nSilver (validated, conformed)\nGold (analytics-ready and ML-ready)\nImplement data governance using Unity Catalog for fine-grained access control, lineage, auditability, and metadata management.\nMLOps & ML-Enabled Data Pipelines\nPartner with data scientists and data engineers to create feature pipelines, model training pipelines, and production scoring pipelines.\nDeploy and operationalize models using MLflow, Databricks Model Registry, and Databricks Workflows.\nUse Databricks built-in AI SQL functions such as ai_query, ai_forecast, ai_analyze_sentiment to generate actionable insight from large amount of unstructured or structured raw data\nImplement monitoring for:\nPipeline failures\nData/feature drift\nModel performance degradation\nOperational SLAs/SLIs/SLOs\nBuild automated CI/CD workflows using GitHub Actions or Azure DevOps for notebook deployment, pipeline testing, and environment promotion.\nData Platform, Data Security & Data Governance\nCollaborate with data engineers to design reliable data products on Delta Lake; leverage Delta Live Tables (DLT) for declarative pipelines when applicable.\nEnforce Unity Catalog for lineage, permissions, and audit; manage secrets, tokens, and keys securely (e.g., Databricks secrets, Key Vault/Secrets Manager).\nCollaboration & Leadership\nWork closely with cross-functional teams: data engineering, data scientist, product manager, and business stakeholders.\nServe as a Databricks SME—championing best practices, code standards, governance, and reusable frameworks.\nDocument architecture, workflows, data models, runbooks, and operational procedures.\nQualifications\nMinimum of 5 years of experience in Databricks, PySpark notebooks, Python, DevOps, software development, and data engineering.\nCertified Databricks Data Engineer Associate or Professional is a plus.\nSkills & Competencies\nProficient in designing, building, deploying, and maintaining high-performance, scalable ETL/ELT pipelines using Azure Databricks, Delta Lake, and PySpark Notebook.\nProficient in building, deploying, and operating production ML models such as supervised, unsupervised, and anomaly detection, including techniques for imbalanced datasets\nExperience working with EKS/AKS cluster and containerized patform.\nProficient with ML engineering and MLOps, including model versioning, CI/CD for ML, monitoring, drift detection, and automated retraining\nProficiency in Python including Pandas and PySpark Data frames\nExpert level of SQL skills including Stored Procedure, experience with SSIS, SSRS, Power BI is a plus.\nProficient with cloud data engineering platforms, such as Azure, Databricks, Spark, or SQL, and batch and streaming pipelines\nFamiliar with Databricks AI Built-In Functions such as AI_Query, AI_Gen, AI_Classify, AI_Forecast, AI_Analyze_Sentiment, able to use them to extract actionable insights from large amount of unstructured or structured raw data\nExperience with Python and ML frameworks, such as PyTorch or TensorFlow\nExperience in improving data quality, lineage, and observability in enterprise data environments and operationalizing rules and model-driven scoring for prioritization, routing, or case selection\nExperience with predictive analytics, machine learning, and artificial intelligence desired.\nBachelor’s degree in Computer Science/ Management Information Systems/ Engineering or related field.\nPay: $120,000.00 - $180,000.00 per year\nBenefits:\n401(k)\nDental insurance\nHealth insurance\nVision insurance\nWork Location: Hybrid remote in Aldie, VA 20105","datePosted":"2026-08-03T14:08:47.637Z","dateModified":"2026-08-03T14:08:47.637Z","hiringOrganization":{"@type":"Organization","name":"Cyberdash Cryptometrics","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Chantilly","addressRegion":"VA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d5b0265f4d7cd7308b696b75"},"url":"https://jobsearcher.com/jobs/d5b0265f4d7cd7308b696b75"}}