Databricks Engineer
JPG Infotech LLC is seeking a *Hands-On Senior Databricks Engineer & Architect* to lead data modernization and engineering initiatives on a mission-critical Federal contract. In this role, you will own the end-to-end Databricks ecosystem—from infrastructure deployment and security configuration to designing scalable Lakehouse architectures and executing complex data migrations from legacy environments.
You will bridge the gap between high-level systems architecture and hands-on PySpark/SQL development, working closely with federal stakeholders to transform complex data requirements into secure, high-performance analytical pipelines.
Key Responsibilities1. Databricks Installation, Configuration & Governance
* *Platform Setup:* Architect, install, and configure enterprise Databricks workspaces within a secure Federal Cloud environment (AWS GovCloud, Azure Government, or GCP).
* *Cluster & Compute Management:* Configure compute clusters, auto-scaling policies, and job orchestration environments to optimize performance and cloud consumption.
* *Governance & Security:* Implement *Unity Catalog* for centralized data governance, metadata management, fine-grained access control (RBAC/ABAC), and data lineage tracking.
* *Compliance:* Ensure all Databricks deployments align with federal security frameworks (e.g., FedRAMP, FISMA, NIST SP 800-53).
2. Data Architecture & Lakehouse Design
* *Lakehouse Architecture:* Design and deploy modern *Medallion Architectures* (Bronze, Silver, Gold layers) using Delta Lake.
* *Data Modeling:* Create conceptual, logical, and physical data models for transactional, dimensional, and analytical data stores.
* *System Integration:* Establish seamless API and pipeline integrations across multi-cloud services, enterprise data warehouses, and downstream reporting tools.
3. Data Migration & Pipeline Engineering
* *Legacy Modernization:* Lead the migration of on-premises data warehouses, relational databases, and legacy ETL workloads into Databricks Delta Lake.
* *Pipeline Development:* Build, test, and deploy robust ETL/ELT pipelines using *PySpark, Python, and SQL* for both batch and real-time streaming data.
* *Orchestration:* Configure automated scheduling, dependency management, and monitoring using Lakeflow / Databricks Workflows or Apache Airflow.
4. Data Analysis & Performance Tuning
* *Query & Cluster Optimization:* Diagnose and troubleshoot bottlenecks, optimize Spark query execution plans, and implement caching and partitioning strategies.
* *Data Quality & Validation:* Implement automated data-reconciliation, completeness, and schema-enforcement checks to ensure auditability and compliance.
* *Stakeholder Collaboration:* Partner with federal program managers, analysts, and engineering teams to translate operational requirements into actionable technical solutions.
Required QualificationsSecurity & Location
* *Citizenship:* *Must be a U.S. Citizen* (dual citizenship or non-citizen status cannot be accommodated due to federal contract mandate).
* *Clearance:* Ability to pass a federal background investigation and obtain/maintain a *Public Trust* or *Secret clearance*.
* *Location:* Preference for candidates residing in the *Washington, D.C. Metropolitan Area (DC/MD/VA)* with the flexibility to work on-site a few days per week as required by project milestones.
Experience & Technical Skills
* *Experience:* *5+ years* of hands-on data engineering, data architecture, or big data platform experience.
* *Databricks Expertise:* Demonstrated hands-on experience installing, configuring, and maintaining production *Databricks*environments.
* *Programming Languages:* Deep proficiency writing production-grade code in *Python (PySpark)* and *SQL* (Scala is a plus).
* *Core Technologies:*
* Strong expertise with *Apache Spark™* distributed processing and runtime internals.
* Extensive experience building Lakehouses with *Delta Lake* and *Unity Catalog*.
* Familiarity with major cloud infrastructure platforms (*AWS, Azure, or GCP*).
* *Engineering Practices:* Experience with Git-based CI/CD workflows, automated testing, and Infrastructure as Code (e.g., Terraform) for data platform deployment.
Preferred Qualifications
* *Certifications:* _Databricks Certified Data Engineer Professional_ or _Databricks Certified Data Engineer Associate_.
* *Public Sector Experience:* Prior experience working on federal, state, or defense IT contracts.
* *Advanced Databricks Tools:* Experience with Delta Live Tables (DLT) / Lakeflow Declarative Pipelines, Databricks SQL Warehouses, or MLOps integrations.
* *Data Integration:* Familiarity with data contract standards, source-to-target mapping, and open-data government reporting requirements.
Why Join JPG Infotech LLC?
At *JPG Infotech LLC*, we deliver cutting-edge AI, cloud, and digital transformation solutions to federal and enterprise clients. You will be part of an agile, high-impact technical team working on meaningful government missions with access to the latest data and AI technologies.
Pay: $117,766.11 - $142,557.03 per year
Benefits:
* 401(k)
* 401(k) matching
* Health insurance
* Paid time off
* Referral program
* Retirement plan
Work Location: Hybrid remote in Ballston, VA