{"schemaVersion":"jobsearcher.job.v1","id":"f7ac39bbbb613f15166bf32b","url":"https://jobsearcher.com/jobs/f7ac39bbbb613f15166bf32b","canonicalUrl":"https://jobsearcher.com/jobs/f7ac39bbbb613f15166bf32b","title":"Lead Data Engineer","description":"Capital Technology Group provides expert consulting services software development, digital transformation, human-centered design, data analytics and visualization, and cybersecurity.\n\nOur multidisciplinary teams use agile methodologies to rapidly and incrementally deliver value in close collaboration with our clients. For over a decade, we have been trusted by both federal and commercial clients to solve complex, mission-critical business challenges. The quality of our work has been recognized by our partners and peers through our inclusion in the Digital Services Coalition, a group of forward- thinking firms recognized for excellence in delivering IT services.\n\nClient Requirements: applicants MUST BE US Citizens and be able to obtain Public Trust clearance\n\nThe CTG Experience\n\nAt Capital Technology Group (CTG), our teams are passionate about modernizing how the federal government delivers software. We partner with federal agencies to build secure, scalable, and mission-driven solutions that make a meaningful impact on millions of people. Recognized by The Washington Post as a Top Workplace in 2025 and 2026. CTG fosters a culture rooted in our core values. Our values guide how we work together and support one another, creating an environment where employees feel trusted, empowered, and encouraged to grow both personally and professionally.\n\nAbout the Role\n\nCTG is seeking a Lead Data Engineer to design, build, and maintain scalable, efficient data pipelines and systems following modern data engineering best practices. The Lead Data Engineer will partner with other Data Engineers to evaluate and prototype new tools and technologies, assess associated risks and benefits, and deliver exceptional value to our clients.\n\nYou Will Get To\nDesign, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), dbt, SQL (PostgreSQL), and AWS Glue.\nDevelop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch.\nBuild high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro.\nDesign and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.\nIntegrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases.\nBuild AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vectors and OpenSearch vector indexes.\nDevelop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures.\nImprove the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization.\nSupport mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST SP 800-53 security controls.\nLead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, Rundeck, and shell-based workflows to cloud-native AWS services.\nMentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices.\nCollaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.\nWho You Are\nA strategic data engineer who enjoys designing sophisticated systems and solving challenging problems\nStrong experience in modern cloud-based solution design\nComfortable balancing business needs with technical constraints and long-term strategy\nA strong communicator\nCollaborative, proactive, and comfortable navigating ambiguity\nQualifications\nBachelor's degree in Computer Science, Engineering, or a related technical field\n15+ years of professional experience in data engineering, data architecture, or related fields\nStrong hands-on experience with:\nApache Spark (PySpark) required, Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling.\nAWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.\nDeveloping scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.\nWorking with modern data lake technologies and data formats such as Parquet, ORC, and Avro.\nDesigning and optimizing solutions using relational and NoSQL databases including PostgreSQL, Redshift, Oracle, GraphDB, and other NoSQL platforms.\nBuilding reliable, high-performance data platforms through performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.\nJava development and modern CI/CD practices using Harness.\nStrong analytical and problem-solving skills\nExperience working in Agile, iterative software development environments\nAbility to quickly learn and apply new technologies and domain knowledge\nExcellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences\nNice to Have\nExperience supporting analytics, data engineering, or modernization initiatives for financial regulators, capital markets, or other highly regulated environments is a plus.\nExperience with Apache Iceberg and modern data lakehouse architectures.\nExperience working with unstructured data processing, including document/text processing, embeddings, vector search, and LLM-based data solutions.\nExposure to integrating LLMs and generative AI capabilities into enterprise data pipelines and platforms.\nExperience designing data architectures that support both structured and unstructured data at scale.\nSalary\n\nWe are committed to offering a competitive salary for this position, with an estimated range of $150k to $200k annually. Please note that this range is intended to provide a general idea of what to expect; however, the final offer may vary based on experience, skills, and other factors. The stated range is not a guarantee and is subject to change.\n\nFull Time Employee Benefits\nRemote Work (Hybrid roles will be specified in the job post)\nCompetitive Compensation Package\nMedical, Dental, and Vision\nLife Insurance, Short/Long Term Disability\nEmployee Assistance Program\n401(k) with 4% matching\nLiberal PTO vacation policy\nGenerous Annual Continuing Education\nAnnual Wellness Budget\nBonus Incentive Programs (Employee referrals and performance-based rewards)\n\nThanks for your interest in Capital Technology Group!\n\nCapital Technology Group is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, or any other characteristic protected by law.","company":"Capitaltechnologygroup","rawCompany":"capitaltechnologygroup","city":"Washington","state":"DC","isRemote":false,"isActive":false,"createdAt":"2026-09-25T12:02:52.286Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"},{"code":"541330","title":"Engineering Services","slug":"engineering-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead Data Engineer","description":"Capital Technology Group provides expert consulting services software development, digital transformation, human-centered design, data analytics and visualization, and cybersecurity.\n\nOur multidisciplinary teams use agile methodologies to rapidly and incrementally deliver value in close collaboration with our clients. For over a decade, we have been trusted by both federal and commercial clients to solve complex, mission-critical business challenges. The quality of our work has been recognized by our partners and peers through our inclusion in the Digital Services Coalition, a group of forward- thinking firms recognized for excellence in delivering IT services.\n\nClient Requirements: applicants MUST BE US Citizens and be able to obtain Public Trust clearance\n\nThe CTG Experience\n\nAt Capital Technology Group (CTG), our teams are passionate about modernizing how the federal government delivers software. We partner with federal agencies to build secure, scalable, and mission-driven solutions that make a meaningful impact on millions of people. Recognized by The Washington Post as a Top Workplace in 2025 and 2026. CTG fosters a culture rooted in our core values. Our values guide how we work together and support one another, creating an environment where employees feel trusted, empowered, and encouraged to grow both personally and professionally.\n\nAbout the Role\n\nCTG is seeking a Lead Data Engineer to design, build, and maintain scalable, efficient data pipelines and systems following modern data engineering best practices. The Lead Data Engineer will partner with other Data Engineers to evaluate and prototype new tools and technologies, assess associated risks and benefits, and deliver exceptional value to our clients.\n\nYou Will Get To\nDesign, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), dbt, SQL (PostgreSQL), and AWS Glue.\nDevelop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch.\nBuild high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro.\nDesign and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.\nIntegrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases.\nBuild AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vectors and OpenSearch vector indexes.\nDevelop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures.\nImprove the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization.\nSupport mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST SP 800-53 security controls.\nLead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, Rundeck, and shell-based workflows to cloud-native AWS services.\nMentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices.\nCollaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.\nWho You Are\nA strategic data engineer who enjoys designing sophisticated systems and solving challenging problems\nStrong experience in modern cloud-based solution design\nComfortable balancing business needs with technical constraints and long-term strategy\nA strong communicator\nCollaborative, proactive, and comfortable navigating ambiguity\nQualifications\nBachelor's degree in Computer Science, Engineering, or a related technical field\n15+ years of professional experience in data engineering, data architecture, or related fields\nStrong hands-on experience with:\nApache Spark (PySpark) required, Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling.\nAWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.\nDeveloping scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.\nWorking with modern data lake technologies and data formats such as Parquet, ORC, and Avro.\nDesigning and optimizing solutions using relational and NoSQL databases including PostgreSQL, Redshift, Oracle, GraphDB, and other NoSQL platforms.\nBuilding reliable, high-performance data platforms through performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.\nJava development and modern CI/CD practices using Harness.\nStrong analytical and problem-solving skills\nExperience working in Agile, iterative software development environments\nAbility to quickly learn and apply new technologies and domain knowledge\nExcellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences\nNice to Have\nExperience supporting analytics, data engineering, or modernization initiatives for financial regulators, capital markets, or other highly regulated environments is a plus.\nExperience with Apache Iceberg and modern data lakehouse architectures.\nExperience working with unstructured data processing, including document/text processing, embeddings, vector search, and LLM-based data solutions.\nExposure to integrating LLMs and generative AI capabilities into enterprise data pipelines and platforms.\nExperience designing data architectures that support both structured and unstructured data at scale.\nSalary\n\nWe are committed to offering a competitive salary for this position, with an estimated range of $150k to $200k annually. Please note that this range is intended to provide a general idea of what to expect; however, the final offer may vary based on experience, skills, and other factors. The stated range is not a guarantee and is subject to change.\n\nFull Time Employee Benefits\nRemote Work (Hybrid roles will be specified in the job post)\nCompetitive Compensation Package\nMedical, Dental, and Vision\nLife Insurance, Short/Long Term Disability\nEmployee Assistance Program\n401(k) with 4% matching\nLiberal PTO vacation policy\nGenerous Annual Continuing Education\nAnnual Wellness Budget\nBonus Incentive Programs (Employee referrals and performance-based rewards)\n\nThanks for your interest in Capital Technology Group!\n\nCapital Technology Group is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, or any other characteristic protected by law.","datePosted":"2026-09-25T12:02:52.286Z","dateModified":"2026-09-25T12:02:52.286Z","hiringOrganization":{"@type":"Organization","name":"Capitaltechnologygroup","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Washington","addressRegion":"DC","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f7ac39bbbb613f15166bf32b"},"url":"https://jobsearcher.com/jobs/f7ac39bbbb613f15166bf32b"}}