{"schemaVersion":"jobsearcher.job.v1","id":"5220314ca674e57f8d0c0786","url":"https://jobsearcher.com/jobs/5220314ca674e57f8d0c0786","canonicalUrl":"https://jobsearcher.com/jobs/5220314ca674e57f8d0c0786","title":"Databricks Data Engineer","description":"Remote\nContract (7 months 25 days)\nPublished 17 hours ago\nAWS certifications\ndata governance\naws cloud\ndata modeling\nperformance optimization\npyspark\nTroubleshooting & Debugging\nSQL\nDevOps & CI/CD\nETL/ELT pipelines\nWe are seeking a Databricks Engineer to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS.\nThe role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns.\nThe architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets.\n\nKey Responsibilities:\nArchitecture & Design:\nDefine end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains\nDesign medallion architecture (Bronze / Silver / Gold) aligned with OneData standards\nEstablish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations\nDefine scalable ingestion patterns for batch and incremental loads\nDrive performance, scalability, and cost optimization best practices\n\nData Engineering & Implementation:\nBuild and guide development of PySpark-based data pipelines in Databricks\nImplement Delta Lake features:\nACID transactions\nSchema evolution & enforcement\nTime travel & versioning\nDesign and optimize large-scale joins, aggregations, and window functions\nImplement CDC and incremental processing using watermarking and change detection\nEnsure idempotent, restartable, and fault-tolerant pipelines\n\nAWS & Platform Integration:\nArchitect solutions using AWS services:\nAmazon S3 (data lake storage)\nIAM (security & access control)\nAWS Glue / Glue Catalog\nCloudWatch (monitoring & logging)\nOptimize Databricks cluster configurations (job vs all-purpose clusters)\nImplement secrets management and secure connectivity patterns\n\nData Quality, Governance & Reliability:\nDefine and implement data quality checks and validations\nEnsure data lineage and metadata capture\nImplement error handling, auditing, and reconciliation frameworks\nSupport data governance and access control requirements\n\nDevOps & Operational Excellence:\nImplement CI/CD pipelines for Databricks notebooks and jobs\nEnforce code versioning, reviews, and deployment standards\nDesign monitoring, alerting, and SLA tracking for pipelines\nSupport production stabilization and performance tuning\n\nCollaboration & Leadership:\nAct as technical lead / mentor for Databricks data engineers\nCollaborate with:\nSource system teams (Sales, CRM, ERP)\nData consumers (analytics, downstream apps)\nCloud/platform teams\nTranslate business requirements into robust technical designs\n\nRequired Skills & Experience:\nCore Technical Skills:\n8+ years in Data Engineering / Data Architecture roles\n4+ years hands-on experience with Databricks\nStrong expertise in PySpark & Spark SQL\nStrong expertise in dbt\nDeep experience with Delta Lake\nStrong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch)\n\nData Engineering Expertise:\nSales data domain experience (orders, revenue, pricing, customers, products)\nStrong understanding of:\nFact & dimension modeling\nSlowly Changing Dimensions (SCD Type 1 / 2)\nLarge-scale data processing patterns\nExperience handling high-volume, high-velocity datasets\n\nPlatform & Operational Skills:\nDatabricks job orchestration and scheduling\nCluster sizing and performance tuning\nCI/CD for data platforms\nStrong troubleshooting and debugging skills\n\nNice-to-Have:\nExperience with enterprise OneData / Data Mesh programs\nExposure to real-time or near-real-time ingestion patterns\nExperience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data)\nAWS certifications or Databricks certifications\nThe pay range that the employer in good faith reasonably expects to pay for this position is $60.10/hour - $93.90/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.\nTundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.","company":"Capgemini","rawCompany":"capgemini","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-12T13:27:45.620Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Databricks Data Engineer","description":"Remote\nContract (7 months 25 days)\nPublished 17 hours ago\nAWS certifications\ndata governance\naws cloud\ndata modeling\nperformance optimization\npyspark\nTroubleshooting & Debugging\nSQL\nDevOps & CI/CD\nETL/ELT pipelines\nWe are seeking a Databricks Engineer to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS.\nThe role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns.\nThe architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets.\n\nKey Responsibilities:\nArchitecture & Design:\nDefine end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains\nDesign medallion architecture (Bronze / Silver / Gold) aligned with OneData standards\nEstablish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations\nDefine scalable ingestion patterns for batch and incremental loads\nDrive performance, scalability, and cost optimization best practices\n\nData Engineering & Implementation:\nBuild and guide development of PySpark-based data pipelines in Databricks\nImplement Delta Lake features:\nACID transactions\nSchema evolution & enforcement\nTime travel & versioning\nDesign and optimize large-scale joins, aggregations, and window functions\nImplement CDC and incremental processing using watermarking and change detection\nEnsure idempotent, restartable, and fault-tolerant pipelines\n\nAWS & Platform Integration:\nArchitect solutions using AWS services:\nAmazon S3 (data lake storage)\nIAM (security & access control)\nAWS Glue / Glue Catalog\nCloudWatch (monitoring & logging)\nOptimize Databricks cluster configurations (job vs all-purpose clusters)\nImplement secrets management and secure connectivity patterns\n\nData Quality, Governance & Reliability:\nDefine and implement data quality checks and validations\nEnsure data lineage and metadata capture\nImplement error handling, auditing, and reconciliation frameworks\nSupport data governance and access control requirements\n\nDevOps & Operational Excellence:\nImplement CI/CD pipelines for Databricks notebooks and jobs\nEnforce code versioning, reviews, and deployment standards\nDesign monitoring, alerting, and SLA tracking for pipelines\nSupport production stabilization and performance tuning\n\nCollaboration & Leadership:\nAct as technical lead / mentor for Databricks data engineers\nCollaborate with:\nSource system teams (Sales, CRM, ERP)\nData consumers (analytics, downstream apps)\nCloud/platform teams\nTranslate business requirements into robust technical designs\n\nRequired Skills & Experience:\nCore Technical Skills:\n8+ years in Data Engineering / Data Architecture roles\n4+ years hands-on experience with Databricks\nStrong expertise in PySpark & Spark SQL\nStrong expertise in dbt\nDeep experience with Delta Lake\nStrong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch)\n\nData Engineering Expertise:\nSales data domain experience (orders, revenue, pricing, customers, products)\nStrong understanding of:\nFact & dimension modeling\nSlowly Changing Dimensions (SCD Type 1 / 2)\nLarge-scale data processing patterns\nExperience handling high-volume, high-velocity datasets\n\nPlatform & Operational Skills:\nDatabricks job orchestration and scheduling\nCluster sizing and performance tuning\nCI/CD for data platforms\nStrong troubleshooting and debugging skills\n\nNice-to-Have:\nExperience with enterprise OneData / Data Mesh programs\nExposure to real-time or near-real-time ingestion patterns\nExperience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data)\nAWS certifications or Databricks certifications\nThe pay range that the employer in good faith reasonably expects to pay for this position is $60.10/hour - $93.90/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.\nTundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.","datePosted":"2026-08-12T13:27:45.620Z","dateModified":"2026-08-12T13:27:45.620Z","hiringOrganization":{"@type":"Organization","name":"Capgemini","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"5220314ca674e57f8d0c0786"},"url":"https://jobsearcher.com/jobs/5220314ca674e57f8d0c0786"}}