Databricks Engineer (Mostly Remote)
This position is a 6 month contract to hire, paying $70 - 75/hrThis role is remote, with occasional on site meetings in Richmond, VA. As a Senior Data Engineer on the Data Solutions and Engineering team, you will drive the transformation of core data capabilities using Databricks. You will focus on building enterprise frameworks, optimizing pipeline infrastructure, and implementing medallion architecture solutions. In this role, you will lead end-to-end SDLC processes—from initial requirements to build, deployment, and support—while integrating advanced capabilities like AI-driven document processing to modernize legacy data structures.ResponsibilitiesFramework & Capability Building: Enhance out-of-the-box Databricks features by building scalable frameworks, dashboards, and Delta table infrastructure (e.g., SLA data freshness tracking, data observability, pipeline monitoring).Data Architecture & Engineering: Design and implement medallion architecture (Bronze, Silver, Gold layers) using ER and dimensional modeling to transform unstructured and legacy data into high-value assets.AI & Document Processing: Leverage features like Intelligent Document Processing and AI extraction tools to parse images, PDFs, and unstructured records into structured analytical models.Full SDLC Ownership: Collaborate with business users, architects, and technical leads to gather requirements, estimate work efforts, write code, run unit/SIT testing, and support UAT.Enablement & Best Practices: Create clear documentation and training materials to help cross-functional and vendor teams (onshore/offshore) adopt your custom frameworks and Databricks best practices.Governance & Security: Implement and maintain robust data security policies, Unity Catalog governance, and industry-standard compliance controls.QualificationsCore Databricks Experience: 3+ years of full-time experience using Databricks (in Azure or AWS) leveraging medallion architecture, batch/streaming orchestration, and compute/storage optimization.Languages & Querying: Expert-level proficiency in SQL and intermediate-to-advanced proficiency in PySpark.Data Modeling: Demonstrated experience creating standardized dimension/fact tables, handling SCD Type 2 time tracking, and configuring ER models.Pipeline Engineering: Proven track record building reusable coding frameworks for data standardization (e.g., missing values, date alignment, reference data management, data repair/imputation).Tooling & DevOps: Strong background with Git/GitLab CI/CD, Lakeflow, data quality rules, Spark UI, and job scheduling.Agile Proficiency: Comfortable operating in an Agile environment (defining technical tasks, writing user stories, and providing effort estimates).