{"schemaVersion":"jobsearcher.job.v1","id":"697c74f5b6a09d5c87a840c8","url":"https://jobsearcher.com/jobs/697c74f5b6a09d5c87a840c8","canonicalUrl":"https://jobsearcher.com/jobs/697c74f5b6a09d5c87a840c8","title":"Apache Spark Developer","description":"Apache Spark Developer\r\nThis role focuses on developing high-performance Spark applications within a containerized, Kubernetes-based environment, supporting mission analytics, data exploitation, and AI/ML integration. The ideal candidate thrives in distributed data environments, understands performance tuning deeply, and can operate effectively in secure, air-gapped systems.\r\nThis role is on-site/flexible hours in Herndon, VA; Springfield, VA; St. Louis, MO; or Aurora, CO.\r\nClearance Required for this role\r\nMinimum Secret; TS/SCI eligibility with willingness/ability to obtain CI polygraph preferred.\r\nCore Technology Stack\r\nData / Processing\r\nApache Spark (PySpark, Scala)\r\nDelta Lake, Parquet\r\nStructured Streaming\r\nInfrastructure\r\nKubernetes (execution environment)\r\nDocker\r\nStorage / Cloud (Abstracted)\r\nS3 / object storage\r\nAWS / GCP / Azure (environment-dependent)\r\nDevOps (Exposure Level)\r\nGit, Jenkins (CI/CD)\r\nLanguages\r\nPython (PySpark)\r\nScala (preferred)\r\nBash / scripting\r\nKey Responsibilities\r\nDesign, develop, and maintain Apache Spark pipelines (batch and streaming) using PySpark and/or Scala.\r\nProcess and transform large-scale datasets using modern data lake architectures (Delta Lake, Parquet).\r\nOptimize Spark jobs for performance, including partitioning strategies, shuffle optimization, memory tuning, and file sizing and storage efficiency.\r\nImplement Structured Streaming pipelines for near real-time data processing.\r\nDevelop and deploy Spark applications within containerized environments (Docker).\r\nExecute workloads in Kubernetes clusters, supporting scalable and distributed processing.\r\nIntegrate Spark pipelines with downstream systems, including analytics platforms (SQL, notebooks) and AI/ML workflows and feature engineering pipelines.\r\nSupport data ingestion and storage in object-based systems such as S3-compatible storage.\r\nTroubleshoot data pipeline failures and ensure reliability in mission-critical environments.\r\nOperate within secure, air-gapped environments: manage dependencies without internet access and work within controlled network and security constraints.\r\nRequired Qualifications\r\nTS/SCI eligibility with ability/willingness to obtain/maintain counterintelligence polygraph.\r\nBachelor's degree plus 5 years' experience in data engineering or Spark development (will entertain additional years' experience in lieu of degree).\r\nStrong hands-on experience with Apache Spark (mandatory), Python (PySpark), and data processing at scale.\r\nExperience working with Parquet and/or Delta Lake, and distributed data systems.\r\nFamiliarity with Docker/containerization and Kubernetes (basic to intermediate experience).\r\nExperience with object storage systems such as S3 or equivalent.\r\nStrong troubleshooting and performance tuning skills.\r\nProficiency in Bash or scripting.\r\nPreferred Qualifications\r\nExperience with Scala for Spark development.\r\nExperience with Structured Streaming in production environments.\r\nFamiliarity with Iceberg or lakehouse architectures.\r\nExperience with CI/CD pipelines (Jenkins, Git).\r\nExposure to Terraform or Infrastructure as Code.\r\nExperience supporting AI/ML data pipelines.\r\nPrior experience supporting NGA, IC, or DoD programs.\r\nBenefits\r\nGenerous PTO plus 11 Federal Holidays.\r\nRetirement Planning – 401(k) fully vested with match.\r\nTuition Assistance Program – annual contributions to help you pay down your loans.\r\nAnnual Health and Wellness Allowance – buy an Apple Watch, a treadmill, or hit the gym on us.\r\nCareer Development – annual funds to spend on education and training.\r\nVolunteer Time Off – annually, all employees can spend 8 hours directly supporting a charity of choice.\r\nCharitable Match – ABSC matches an employee's donation to a qualifying charity.\r\nReferral Program – we pay for internal and external referrals.\r\nLOV Awards – earn bonus awards throughout the year from our Living Our Values awards program.\r\nEqual Opportunity Employer, including veterans and individuals with disabilities.\r\nJ-18808-Ljbffr","company":"Abscus","rawCompany":"abscus","city":"Herndon","state":"VA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T02:08:43.086Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1251.00","title":"Computer Programmers","slug":"computer-programmers"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Apache Spark Developer","description":"Apache Spark Developer\r\nThis role focuses on developing high-performance Spark applications within a containerized, Kubernetes-based environment, supporting mission analytics, data exploitation, and AI/ML integration. The ideal candidate thrives in distributed data environments, understands performance tuning deeply, and can operate effectively in secure, air-gapped systems.\r\nThis role is on-site/flexible hours in Herndon, VA; Springfield, VA; St. Louis, MO; or Aurora, CO.\r\nClearance Required for this role\r\nMinimum Secret; TS/SCI eligibility with willingness/ability to obtain CI polygraph preferred.\r\nCore Technology Stack\r\nData / Processing\r\nApache Spark (PySpark, Scala)\r\nDelta Lake, Parquet\r\nStructured Streaming\r\nInfrastructure\r\nKubernetes (execution environment)\r\nDocker\r\nStorage / Cloud (Abstracted)\r\nS3 / object storage\r\nAWS / GCP / Azure (environment-dependent)\r\nDevOps (Exposure Level)\r\nGit, Jenkins (CI/CD)\r\nLanguages\r\nPython (PySpark)\r\nScala (preferred)\r\nBash / scripting\r\nKey Responsibilities\r\nDesign, develop, and maintain Apache Spark pipelines (batch and streaming) using PySpark and/or Scala.\r\nProcess and transform large-scale datasets using modern data lake architectures (Delta Lake, Parquet).\r\nOptimize Spark jobs for performance, including partitioning strategies, shuffle optimization, memory tuning, and file sizing and storage efficiency.\r\nImplement Structured Streaming pipelines for near real-time data processing.\r\nDevelop and deploy Spark applications within containerized environments (Docker).\r\nExecute workloads in Kubernetes clusters, supporting scalable and distributed processing.\r\nIntegrate Spark pipelines with downstream systems, including analytics platforms (SQL, notebooks) and AI/ML workflows and feature engineering pipelines.\r\nSupport data ingestion and storage in object-based systems such as S3-compatible storage.\r\nTroubleshoot data pipeline failures and ensure reliability in mission-critical environments.\r\nOperate within secure, air-gapped environments: manage dependencies without internet access and work within controlled network and security constraints.\r\nRequired Qualifications\r\nTS/SCI eligibility with ability/willingness to obtain/maintain counterintelligence polygraph.\r\nBachelor's degree plus 5 years' experience in data engineering or Spark development (will entertain additional years' experience in lieu of degree).\r\nStrong hands-on experience with Apache Spark (mandatory), Python (PySpark), and data processing at scale.\r\nExperience working with Parquet and/or Delta Lake, and distributed data systems.\r\nFamiliarity with Docker/containerization and Kubernetes (basic to intermediate experience).\r\nExperience with object storage systems such as S3 or equivalent.\r\nStrong troubleshooting and performance tuning skills.\r\nProficiency in Bash or scripting.\r\nPreferred Qualifications\r\nExperience with Scala for Spark development.\r\nExperience with Structured Streaming in production environments.\r\nFamiliarity with Iceberg or lakehouse architectures.\r\nExperience with CI/CD pipelines (Jenkins, Git).\r\nExposure to Terraform or Infrastructure as Code.\r\nExperience supporting AI/ML data pipelines.\r\nPrior experience supporting NGA, IC, or DoD programs.\r\nBenefits\r\nGenerous PTO plus 11 Federal Holidays.\r\nRetirement Planning – 401(k) fully vested with match.\r\nTuition Assistance Program – annual contributions to help you pay down your loans.\r\nAnnual Health and Wellness Allowance – buy an Apple Watch, a treadmill, or hit the gym on us.\r\nCareer Development – annual funds to spend on education and training.\r\nVolunteer Time Off – annually, all employees can spend 8 hours directly supporting a charity of choice.\r\nCharitable Match – ABSC matches an employee's donation to a qualifying charity.\r\nReferral Program – we pay for internal and external referrals.\r\nLOV Awards – earn bonus awards throughout the year from our Living Our Values awards program.\r\nEqual Opportunity Employer, including veterans and individuals with disabilities.\r\nJ-18808-Ljbffr","datePosted":"2026-07-16T02:08:43.086Z","dateModified":"2026-07-16T02:08:43.086Z","hiringOrganization":{"@type":"Organization","name":"Abscus","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Herndon","addressRegion":"VA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"697c74f5b6a09d5c87a840c8"},"url":"https://jobsearcher.com/jobs/697c74f5b6a09d5c87a840c8"}}