{"schemaVersion":"jobsearcher.job.v1","id":"092f7495b339c2cdbfe1e21b","url":"https://jobsearcher.com/jobs/092f7495b339c2cdbfe1e21b","canonicalUrl":"https://jobsearcher.com/jobs/092f7495b339c2cdbfe1e21b","title":"Data Engineer","description":"Location: 100% Remote\n\nYears' Experience: 5+ years Professional Experience\n\nEducation: Bachelor's Degree in IT related field\n\nClearance: Applicants must be able to obtain and maintain a secret security clearance. United States Citizenship is required as part of the eligibility criteria to be able to obtain this type of security clearance.\n\nRequired Certifications:\n\nCompTIA Security +\nKey Skills:\n\n5+ years of IT experience focusing on enterprise data architecture and management to include data flow charts, diagrams, and other technical documentation.\nExperience with Databricks, Structured Streaming, Delta Lake concepts, and Delta Live Tables required.\nPython development experience required.\nExperience with ETL and ELT tools such as SSIS, Pentaho, and/or Data Migration Services, and the ability to incorporate Python as required.\nAdvanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization).\nProficiency using Git for version control, including repository management, branching, merging, and pull requests.\nActive CompTIA Security+ certification preferred. If selected, must be able to obtain a CompTIA Security+ certification prior to beginning supporting the program.\nResponsibilities\n\nPlan, create, and maintain data architectures, ensuring alignment with business requirements.\nObtain data, formulate dataset processes, and store optimized data.\nIdentify problems and inefficiencies and apply solutions.\nDetermine tasks where manual participation can be eliminated with automation.\nIdentify and optimize data bottlenecks, leveraging automation where possible.\nCreate and manage data lifecycle policies (retention, backups/restore, etc).\nIn-depth knowledge for creating, maintaining, and managing ETL/ELT pipelines.\nCreate, maintain, and manage data transformations.\nMaintain/update documentation.\nCreate, maintain, and manage data pipeline schedules.\nMonitor data pipelines.\nCreate, maintain, and manage data quality gates (Great Expectations) to ensure high data quality.\nSupport AI/ML teams with optimizing feature engineering code.\nExpertise in Spark/Python/Databricks, Data Lake and SQL.\nCreate, maintain, and manage Spark Structured Steaming jobs, including using the newer Delta Live Tables and/or DBT.\nResearch existing data in the data lake to determine best sources for data.\nCreate, manage, and maintain ksqlDB and Kafka Streams queries/code\nData driven testing for data quality.\nMaintain and update Python-based data processing scripts executed on AWS Lambdas.\nUnit tests for all the Spark, Python data processing and Lambda codes.\nMaintain PCIS Reporting Database data lake with optimizations and maintenance (performance tuning, etc).\nStreamlining data processing experience including formalizing concepts of how to handle lake data, defining windows, and how window definitions impact data freshness.\nQualifications\n\n5+ years of IT experience focusing on enterprise data architecture and management.\nMust have an active Secret security clearance.\nBachelor degree required.\nCompTIA Security+ certification preferred. If selected, must be able to obtain a CompTIA Security+ certification prior to begin supporting the program.\nExperience in Conceptual/Logical/Physical Data Modeling & expertise in Relational and Dimensional Data Modeling.\nExperience with Databricks and Python Development, Structured Streaming, Delta Lake concepts, and Delta Live Tables required.\nAdditional experience with Spark, Spark SQL, Spark DataFrames and DataSets, and PySpark.\nData Lake concepts such as time travel and schema evolution and optimization.\nStructured Streaming and Delta Live Tables with Databricks a bonus.\nKnowledge of Python (Python 3.X) for CI/CD pipelines required.\nFamiliarity with Pytest and Unittest a bonus.\nExperience leading and architecting enterprise-wide initiatives specifically system integration, data migration, transformation, data warehouse build, data mart build, and data lakes implementation / support.\nAdvanced level understanding of streaming data pipelines and how they differ from batch systems.\nFormalize concepts of how to handle late data, defining windows, and data freshness.\nAdvanced understanding of ETL and ELT and ETL/ELT tools such as SSIS, Pentaho, Data Migration Service etc.\nUnderstanding of concepts and implementation strategies for different incremental data loads such as tumbling window, sliding window, high watermark, etc.\nFamiliarity and/or expertise with Great Expectations or other data quality/data validation frameworks a bonus.\nUnderstanding of streaming data pipelines and batch systems.\nFamiliarity with concepts such as late data, defining windows, and how window definitions impact data freshness.\nAdvanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization).\nIndexing and partitioning strategy experience.\nDebug, troubleshoot, design and implement solutions to complex technical issues.\nExperience with large-scale, high-performance enterprise big data application deployment and solution.\nUnderstanding how to create DAGs to define workflows.\nFamiliarity with CI/CD pipelines, containerization, and pipeline orchestration tools such as Airflow, Prefect, etc a bonus but not required.\nArchitecture experience in AWS environment a bonus.\nFamiliarity working with Kinesis and/or Lambda specifically with how to push and pull data, how to use AWS tools to view data in Kinesis streams, and for processing massive data at scale a bonus.\nExperience with Docker, Jenkins, and CloudWatch.\nAbility to write and maintain Jenkinsfiles for supporting CI/CD pipelines.\nExperience working with AWS Lambdas for configuration and optimization.\nExperience working with DynamoDB to query and write data.\nExperience with S3.\nExperience working with JSON and defining JSON Schemas a bonus.\nExperience setting up and management Confluent/Kafka topics and ensuring performance using Kafka a bonus.\nFamiliarity with Schema Registry, message formats such as Avro, ORC, etc.\nUnderstanding how to manage ksqlDB SQL files and migrations and Kafka Streams.\nAbility to thrive in a team-based environment.\nExperience briefing the benefits and constraints of technology solutions to technology partners, stakeholders, team members, and senior level of management.\nProficiency using Git for version control, including repository management, branching, merging, and pull requests.\nRepository setup and management.\nBranching strategies (feature, develop, main).\nMerging and resolving conflicts.\nCreating and reviewing pull requests.\nCommit best practices (clear messages, atomic commits).\nTagging and release management.\nAbout Sparibis\n\nSparibis LLC is a professional solution firm that Clients rely on to access the best talent to drive their business success.\n\nSparibis is an equal opportunity employer that values diversity at all levels. All individuals, regardless of personal characteristics, are encouraged to apply.","company":"Sparibis","rawCompany":"sparibis","city":"Washington","state":"DC","isRemote":false,"isActive":false,"createdAt":"2026-08-06T13:09:01.546Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"Location: 100% Remote\n\nYears' Experience: 5+ years Professional Experience\n\nEducation: Bachelor's Degree in IT related field\n\nClearance: Applicants must be able to obtain and maintain a secret security clearance. United States Citizenship is required as part of the eligibility criteria to be able to obtain this type of security clearance.\n\nRequired Certifications:\n\nCompTIA Security +\nKey Skills:\n\n5+ years of IT experience focusing on enterprise data architecture and management to include data flow charts, diagrams, and other technical documentation.\nExperience with Databricks, Structured Streaming, Delta Lake concepts, and Delta Live Tables required.\nPython development experience required.\nExperience with ETL and ELT tools such as SSIS, Pentaho, and/or Data Migration Services, and the ability to incorporate Python as required.\nAdvanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization).\nProficiency using Git for version control, including repository management, branching, merging, and pull requests.\nActive CompTIA Security+ certification preferred. If selected, must be able to obtain a CompTIA Security+ certification prior to beginning supporting the program.\nResponsibilities\n\nPlan, create, and maintain data architectures, ensuring alignment with business requirements.\nObtain data, formulate dataset processes, and store optimized data.\nIdentify problems and inefficiencies and apply solutions.\nDetermine tasks where manual participation can be eliminated with automation.\nIdentify and optimize data bottlenecks, leveraging automation where possible.\nCreate and manage data lifecycle policies (retention, backups/restore, etc).\nIn-depth knowledge for creating, maintaining, and managing ETL/ELT pipelines.\nCreate, maintain, and manage data transformations.\nMaintain/update documentation.\nCreate, maintain, and manage data pipeline schedules.\nMonitor data pipelines.\nCreate, maintain, and manage data quality gates (Great Expectations) to ensure high data quality.\nSupport AI/ML teams with optimizing feature engineering code.\nExpertise in Spark/Python/Databricks, Data Lake and SQL.\nCreate, maintain, and manage Spark Structured Steaming jobs, including using the newer Delta Live Tables and/or DBT.\nResearch existing data in the data lake to determine best sources for data.\nCreate, manage, and maintain ksqlDB and Kafka Streams queries/code\nData driven testing for data quality.\nMaintain and update Python-based data processing scripts executed on AWS Lambdas.\nUnit tests for all the Spark, Python data processing and Lambda codes.\nMaintain PCIS Reporting Database data lake with optimizations and maintenance (performance tuning, etc).\nStreamlining data processing experience including formalizing concepts of how to handle lake data, defining windows, and how window definitions impact data freshness.\nQualifications\n\n5+ years of IT experience focusing on enterprise data architecture and management.\nMust have an active Secret security clearance.\nBachelor degree required.\nCompTIA Security+ certification preferred. If selected, must be able to obtain a CompTIA Security+ certification prior to begin supporting the program.\nExperience in Conceptual/Logical/Physical Data Modeling & expertise in Relational and Dimensional Data Modeling.\nExperience with Databricks and Python Development, Structured Streaming, Delta Lake concepts, and Delta Live Tables required.\nAdditional experience with Spark, Spark SQL, Spark DataFrames and DataSets, and PySpark.\nData Lake concepts such as time travel and schema evolution and optimization.\nStructured Streaming and Delta Live Tables with Databricks a bonus.\nKnowledge of Python (Python 3.X) for CI/CD pipelines required.\nFamiliarity with Pytest and Unittest a bonus.\nExperience leading and architecting enterprise-wide initiatives specifically system integration, data migration, transformation, data warehouse build, data mart build, and data lakes implementation / support.\nAdvanced level understanding of streaming data pipelines and how they differ from batch systems.\nFormalize concepts of how to handle late data, defining windows, and data freshness.\nAdvanced understanding of ETL and ELT and ETL/ELT tools such as SSIS, Pentaho, Data Migration Service etc.\nUnderstanding of concepts and implementation strategies for different incremental data loads such as tumbling window, sliding window, high watermark, etc.\nFamiliarity and/or expertise with Great Expectations or other data quality/data validation frameworks a bonus.\nUnderstanding of streaming data pipelines and batch systems.\nFamiliarity with concepts such as late data, defining windows, and how window definitions impact data freshness.\nAdvanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization).\nIndexing and partitioning strategy experience.\nDebug, troubleshoot, design and implement solutions to complex technical issues.\nExperience with large-scale, high-performance enterprise big data application deployment and solution.\nUnderstanding how to create DAGs to define workflows.\nFamiliarity with CI/CD pipelines, containerization, and pipeline orchestration tools such as Airflow, Prefect, etc a bonus but not required.\nArchitecture experience in AWS environment a bonus.\nFamiliarity working with Kinesis and/or Lambda specifically with how to push and pull data, how to use AWS tools to view data in Kinesis streams, and for processing massive data at scale a bonus.\nExperience with Docker, Jenkins, and CloudWatch.\nAbility to write and maintain Jenkinsfiles for supporting CI/CD pipelines.\nExperience working with AWS Lambdas for configuration and optimization.\nExperience working with DynamoDB to query and write data.\nExperience with S3.\nExperience working with JSON and defining JSON Schemas a bonus.\nExperience setting up and management Confluent/Kafka topics and ensuring performance using Kafka a bonus.\nFamiliarity with Schema Registry, message formats such as Avro, ORC, etc.\nUnderstanding how to manage ksqlDB SQL files and migrations and Kafka Streams.\nAbility to thrive in a team-based environment.\nExperience briefing the benefits and constraints of technology solutions to technology partners, stakeholders, team members, and senior level of management.\nProficiency using Git for version control, including repository management, branching, merging, and pull requests.\nRepository setup and management.\nBranching strategies (feature, develop, main).\nMerging and resolving conflicts.\nCreating and reviewing pull requests.\nCommit best practices (clear messages, atomic commits).\nTagging and release management.\nAbout Sparibis\n\nSparibis LLC is a professional solution firm that Clients rely on to access the best talent to drive their business success.\n\nSparibis is an equal opportunity employer that values diversity at all levels. All individuals, regardless of personal characteristics, are encouraged to apply.","datePosted":"2026-08-06T13:09:01.546Z","dateModified":"2026-08-06T13:09:01.546Z","hiringOrganization":{"@type":"Organization","name":"Sparibis","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Washington","addressRegion":"DC","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"092f7495b339c2cdbfe1e21b"},"url":"https://jobsearcher.com/jobs/092f7495b339c2cdbfe1e21b"}}