Data Engineer – Analytic Platform & Data Pipelines
Role Overview3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.Key ResponsibilitiesData Pipeline Development & IngestionDesign and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feedsDevelop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemasBuild and maintain automated ingestion connectors for web, document, geospatial, and tabular data sourcesManage data orchestration and scheduling using tools such as Airflow, Dagster, or PrefectData Modeling & StorageDesign and maintain data models, schemas, and storage layers across relational, NoSQL, and object storesBuild and maintain data lakes/lakehouses and curated, analysis-ready data martsOptimize partitioning, indexing, and query performance for large datasetsSupport entity resolution and data linking in coordination with the knowledge graph and modeling teamsData Quality, Governance & LineageImplement data validation, quality checks, and monitoring across pipelinesEstablish data lineage, cataloging, and metadata managementEnforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI dataDocument datasets, schemas, and pipeline logic for downstream consumersSecurity & ComplianceEnsure pipelines and data stores meet security requirements for operation in sensitive environmentsImplement encryption, access control, and secure data-handling practicesSupport Authority to Operate (ATO) processes and compliance frameworksRequired QualificationsTechnical Expertise4+ years of data engineering experience building and operating production data pipelinesStrong programming skills in Python and SQL (Scala or Java a plus)Experience with distributed data processing frameworks (Spark, Dask, or similar)Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)Experience with cloud data platforms and services (AWS, Azure, or GCP)Data & InfrastructureExperience designing data models, warehouses, and lakehouse architecturesFamiliarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)Understanding of data quality, lineage, and governance practicesExperience with containerization (Docker) and CI/CD for data workflowsDomain KnowledgeExperience working with large-scale, heterogeneous, or open-source datasetsUnderstanding of data provenance and source-attribution requirementsPreferred QualificationsActive security clearance or ability to obtain oneExperience in government, defense, or intelligence contracting environmentsFamiliarity with PAI/CAI (publicly and commercially available information) data sourcesExperience with geospatial data processing (PostGIS, GDAL, or similar)Knowledge of graph data structures and preparing data for knowledge graphsExperience with streaming platforms (Kafka, Kinesis)Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)Technical EnvironmentLanguages: Python, SQL (Scala/Java a plus)Processing: Spark, Airflow/Dagster/Prefect, streaming frameworksStorage: PostgreSQL, Elasticsearch, object storage / data lake, ParquetInfrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)Security: Encryption at rest and in transit, RBAC, secure data handlingThis role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on.Salary: $135000 - $185000 per year