Data Engineer - AWS/Databricks - Mid Level
Overview
In this role, you will design and deliver AWS cloud-scale data platforms for federal clients, leveraging Spark, Delta Lake, and Databricks to modernize data pipelines. You’ll shape enterprise data solutions, collaborate with cross-functional teams, and drive efficient, scalable data processing and governance. The position offers opportunities to work on mission-critical data modernization initiatives and to advance technical leadership in a federal context.
Compensation / Benefitstraining and certifications support up to $3,000 annuallydegree-seeking program funding up to $3,000competitive compensationcomprehensive benefitswork-life balanceinclusive culture
ResponsibilitiesBuild and maintain PySpark-based data pipelines in Databricks notebooks for ingestion, transformation, and enrichment of structured and semi-structured dataDesign and optimize Delta Lake tables for ACID compliance, partition pruning, schema enforcement, and performance across large datasetsDevelop ETL/ELT workflows integrating multiple sources into a centralized, query-optimized data warehouseUtilize Spark SQL/DataFrame APIs to implement business rules and warehouse modeling logicCollaborate on cloud-native data solutions on AWS using S3, Glue, RDS, and IAM for secure storage and access controlImprove pipeline performance via partitioning, caching, join strategies, and tuningDeploy and version data assets with Git-integrated workflows; automate deployment with CI/CD tools (GitLab, Jenkins)Monitor pipelines, jobs, and clusters with Databricks tools and AWS CloudWatch; optimize cost-performanceConduct technical discovery and mapping of legacy sources; design end-to-end data flowsImplement governance: metadata tagging, data quality validation, audit logging, lineage trackingSupport ad hoc data access, develop reusable assets, and maintain shared notebooks for operational reporting
Key requirements4+ years in data engineering/Agile analytics4+ years building software for structured and unstructured data2+ years building scalable ETL/ELT for analytics2+ years cloud data engineering in AWS and DatabricksExperience with data quality, validation frameworks, and storage optimizationBA or BS degreeUS Citizenship with ability to obtain/maintain US Suitabilitycollaborationproblem-solvingability to communicate with cross-functional teamsPySparkDelta LakeDatabricks