{"schemaVersion":"jobsearcher.job.v1","id":"aabf4e6578d5142d55ec747c","url":"https://jobsearcher.com/jobs/aabf4e6578d5142d55ec747c","canonicalUrl":"https://jobsearcher.com/jobs/aabf4e6578d5142d55ec747c","title":"Sr. Data Integration Engineer","description":"Overview\nIn this role you will design and build resilient data pipelines using Python and AWS to enable secure, scalable data flows for government clients. You will work within Sparksoft’s Innovation Centers to advance end-to-end data solutions and data governance. You’ll tackle data ingestion, quality, and performance challenges, partnering with cross-functional teams to deliver impactful insights. This opportunity lets you shape modern data architectures while contributing to a mission-driven, collaborative culture.\n\nCompensation / BenefitsCompetitive compensation401(k) with employer contributionsFlexible paid time offHybrid work modelComprehensive health coverage (medical, dental, vision, life, disability)Professional development opportunities\nResponsibilitiesDesign and develop batch and event-driven data pipelines using Python and AWS GlueBuild AWS Glue jobs, crawlers, workflows, triggers, and Data Catalog integrationsIngest and process structured and semi-structured data from files, APIs, databases, and streaming sources (JSON/NDJSON)Develop Python components for parsing, validation, normalization, enrichment, deduplication, aggregation, and data quality checksUtilize AWS services (S3, Lambda, EventBridge, Step Functions, SQS, SNS, Kinesis, Athena, Redshift, Lake Formation, Secrets Manager, KMS, CloudWatch, IAM) as neededCreate automated tests (unit, integration, regression, data reconciliation) and embed data quality controlsImplement monitoring, logging, alerting, traceability, restartability, error handling, and recovery for production pipelinesApply security and privacy practices (least privilege, encryption, secure secret management, audit logging)Automate infrastructure and deployments via IaC and CI/CDOptimize performance, reliability, scalability, and cost; profile and tune pipelinesInvestigate production issues, document root causes and preventive actionsCreate and maintain technical documentation (mappings, pipeline designs, contracts, runbooks, lineage, test evidence)Analyze requirements into stories; refine with Product Owners and prioritize backlogCoordinate with external teams on timelines and risk managementMaintain ongoing communication with customers and stakeholders to capture business requirementsReview test scenarios and ensure coverage of impact pointsTrack stories/epics to closure and manage customer requirements from inception to deliveryAssist with user acceptance testing (UAT)Analyze data to understand business problems and opportunitiesIdentify risks and impacts of proposed solutionsProvide ongoing support and maintenance for implemented solutionsAssist product owner and development team to achieve customer satisfaction\nKey requirements7+ years of relevant experienceStrong Python development skills (modular design, testing, debugging, packaging, dependency management, performance optimization)Hands-on experience with AWS Glue and integration with S3 and Glue Data CatalogExperience with multiple AWS data, integration, security, and monitoring servicesExperience ingesting/parsing/validating/transforming/troubleshooting JSON and NDJSON (nested structures, schema drift, large files)Data modeling, schema design, data partitioning, metadata, lineage, and data quality practicesGit-based version control, automated testing, CI/CD, and infrastructure-as-code approachesExcellent analytical, problem-solving, documentation, and communication skills; proactive and customer-focusedStrong verbal/written communication, attention to detail, and follow-up能力Experience creating detailed reports and presenting to technical and non-technical audiencesProficiency with JIRA and Confluence for managing requirementsAbility to obtain and maintain Public Trust clearanceResidence in the United States for 3 of the past 5 yearsanalytical thinkingproactive mindsetcustomer-focusedPythonAWS GlueJSON/NDJSON","company":"Sparksoft","rawCompany":"sparksoft","city":"Wilton","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-22T03:40:30.419Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. Data Integration Engineer","description":"Overview\nIn this role you will design and build resilient data pipelines using Python and AWS to enable secure, scalable data flows for government clients. You will work within Sparksoft’s Innovation Centers to advance end-to-end data solutions and data governance. You’ll tackle data ingestion, quality, and performance challenges, partnering with cross-functional teams to deliver impactful insights. This opportunity lets you shape modern data architectures while contributing to a mission-driven, collaborative culture.\n\nCompensation / BenefitsCompetitive compensation401(k) with employer contributionsFlexible paid time offHybrid work modelComprehensive health coverage (medical, dental, vision, life, disability)Professional development opportunities\nResponsibilitiesDesign and develop batch and event-driven data pipelines using Python and AWS GlueBuild AWS Glue jobs, crawlers, workflows, triggers, and Data Catalog integrationsIngest and process structured and semi-structured data from files, APIs, databases, and streaming sources (JSON/NDJSON)Develop Python components for parsing, validation, normalization, enrichment, deduplication, aggregation, and data quality checksUtilize AWS services (S3, Lambda, EventBridge, Step Functions, SQS, SNS, Kinesis, Athena, Redshift, Lake Formation, Secrets Manager, KMS, CloudWatch, IAM) as neededCreate automated tests (unit, integration, regression, data reconciliation) and embed data quality controlsImplement monitoring, logging, alerting, traceability, restartability, error handling, and recovery for production pipelinesApply security and privacy practices (least privilege, encryption, secure secret management, audit logging)Automate infrastructure and deployments via IaC and CI/CDOptimize performance, reliability, scalability, and cost; profile and tune pipelinesInvestigate production issues, document root causes and preventive actionsCreate and maintain technical documentation (mappings, pipeline designs, contracts, runbooks, lineage, test evidence)Analyze requirements into stories; refine with Product Owners and prioritize backlogCoordinate with external teams on timelines and risk managementMaintain ongoing communication with customers and stakeholders to capture business requirementsReview test scenarios and ensure coverage of impact pointsTrack stories/epics to closure and manage customer requirements from inception to deliveryAssist with user acceptance testing (UAT)Analyze data to understand business problems and opportunitiesIdentify risks and impacts of proposed solutionsProvide ongoing support and maintenance for implemented solutionsAssist product owner and development team to achieve customer satisfaction\nKey requirements7+ years of relevant experienceStrong Python development skills (modular design, testing, debugging, packaging, dependency management, performance optimization)Hands-on experience with AWS Glue and integration with S3 and Glue Data CatalogExperience with multiple AWS data, integration, security, and monitoring servicesExperience ingesting/parsing/validating/transforming/troubleshooting JSON and NDJSON (nested structures, schema drift, large files)Data modeling, schema design, data partitioning, metadata, lineage, and data quality practicesGit-based version control, automated testing, CI/CD, and infrastructure-as-code approachesExcellent analytical, problem-solving, documentation, and communication skills; proactive and customer-focusedStrong verbal/written communication, attention to detail, and follow-up能力Experience creating detailed reports and presenting to technical and non-technical audiencesProficiency with JIRA and Confluence for managing requirementsAbility to obtain and maintain Public Trust clearanceResidence in the United States for 3 of the past 5 yearsanalytical thinkingproactive mindsetcustomer-focusedPythonAWS GlueJSON/NDJSON","datePosted":"2026-09-22T03:40:30.419Z","dateModified":"2026-09-22T03:40:30.419Z","hiringOrganization":{"@type":"Organization","name":"Sparksoft","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Wilton","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"aabf4e6578d5142d55ec747c"},"url":"https://jobsearcher.com/jobs/aabf4e6578d5142d55ec747c"}}