Sensitive-Data Engineer
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
DescriptionActive Top Secret/SCI Clearance with Polygraph (REQUIRED)Are you passionate about harnessing data to solve some of the nation’s most critical challenges? Do you thrive on innovation, collaboration, and building resilient solutions in complex environments?Join a high-impact team at the forefront of national security, where your work directly supports mission success. We're seeking a Data Engineer with a rare mix of curiosity, craftsmanship, and commitment to excellence. In this role, you’ll design and optimize secure, scalable data pipelines while working alongside elite engineers, mission partners, and data experts to unlock actionable insights from diverse datasets.RequirementsEngineer robust, secure, and scalable data pipelines using Apache Spark, Apache Hudi, AWS EMR, and KubernetesMaintain data provenance and access controls to ensure full lineage and auditability of mission-critical datasetsClean, transform, and condition data using tools such as dbt, Apache NiFi, or PandasBuild and orchestrate repeatable ETL workflows using Apache Airflow, Dagster, or PrefectDevelop API connectors for ingesting structured and unstructured data sourcesCollaborate with data stewards, architects, and mission teams to align on data standards, quality, and integrityProvide advanced database administration for Oracle, PostgreSQL, MongoDB, Elasticsearch, and othersIngest and analyze streaming data using tools like Apache Kafka, AWS Kinesis, or Apache FlinkPerform real-time and batch processing on large datasets in secure cloud environments (e.g., AWS GovCloud, C2S)Implement and monitor data quality and validation checks using tools such as Great Expectations or DeequWork across agile teams using DevSecOps practices to build resilient full-stack solutions with Python, Java, or ScalaRequired SkillsExperience building and maintaining data pipelines using Apache Spark, Airflow, NiFi, or dbtProficiency in Python, SQL, and one or more of: Java, ScalaStrong understanding of cloud services (especially AWS and GovCloud), including S3, EC2, Lambda, EMR, Glue, Redshift, or SnowflakeHands-on experience with streaming frameworks such as Apache Kafka, Kafka Connect, or FlinkFamiliarity with data lakehouse formats (e.g., Apache Hudi, Delta Lake, or Iceberg)Experience with NoSQL and RDBMS technologies such as MongoDB, DynamoDB, PostgreSQL, or MySQLAbility to implement and maintain data validation frameworks (e.g., Great Expectations, Deequ)Comfortable working in Linux/Unix environments, using bash scripting, Git, and CI/CD toolsKnowledge of containerization and orchestration tools like Docker and KubernetesCollaborative mindset with experience working in Agile/Scrum environments using Jira, Confluence, and Git-based workflows