JOBSEARCHER

Sr Data Engineer

ArtechO'Fallon, MOL6 LeadSeptember 17th, 2026
Request ID: 107895-1Title: Sr Data EngineerLocation: Ofallon, MODuration: 6 monthsPay Range: $40 - $45/Hour on W2/C2C (All inclusive)JOB DESCRIPTION:We are looking for a highly skilled Senior Data Engineer with deep expertise in Apache Spark, Scala, and PySpark to build and operate large scale batch and streaming data processing systems. The role has a strong emphasis on real time streaming architectures using Kafka and Spark Structured Streaming, alongside ingestion and orchestration with Apache NiFi and scalable storage using Apache Ozone and Ceph. This position is ideal for engineers who enjoy solving complex performance, scalability, latency, and reliability challenges in production data platforms.Key Responsibilities:Design, develop, and maintain large scale Spark applications using Scala and PySparkBuild and operate streaming heavy data pipelines using Kafka and Spark Structured StreamingImplement stateful streaming patterns including windowing, watermarking, late data handling, and checkpointingDevelop robust event replay and reprocessing workflows using Kafka offsets and partitionsBuild ingestion and routing flows using Apache NiFi, including Kafka based ingestion patternsImplement end to end ETL/ELT pipelines with strong emphasis on low latency, fault tolerance, and scalabilityOptimize Spark jobs through partitioning strategies, memory tuning, shuffle optimization, and efficient data formatsIntegrate Spark workloads with distributed object storage systems such as Apache Ozone and CephEnsure data quality, consistency, and auditability through validation, reconciliation, and metadata captureCollaborate with platform, infrastructure, and operations teams on production readiness and capacity planningSupport production systems, including monitoring, incident analysis, and root cause resolutionContribute to reusable frameworks, coding standards, and engineering best practicesParticipate in architecture reviews, code reviews, and technical documentationRequired Qualifications:Bachelor's degree in computer science, Engineering, or equivalent practical experienceStrong hands on experience with Apache Spark in production environmentsAdvanced proficiency in Scala and PySparkSolid understanding of distributed systems and data processing at scaleStrong experience with Kafka based streaming architecturesHands on experience with Spark Structured StreamingExperience building batch and real time pipelinesHands on experience with Apache NiFi for data ingestion and flow managementStrong SQL skills and experience working with structured and semi structured dataExperience working with object storage or distributed storage platformsProficiency with Linux, shell scripting, and Git based version controlPreferred QualificationsExperience with Apache Ozone and/or Ceph as storage backends for analytics workloadsExperience implementing exactly once / at least once streaming semanticsStrong background in Spark performance tuning (CPU, memory, I/O, shuffle)Experience supporting mission critical production systems with strict SLAsFamiliarity with CI/CD pipelines and automated testing for data applicationsExperience designing observability for streaming systems (lag, throughput, backpressure)Technical SkillsLanguages: Scala, Python (PySpark), SQLBig Data: Apache Spark (Core, SQL, Structured Streaming)Streaming: KafkaIngestion / Orchestration: Apache NiFiStorage: Apache Ozone, Ceph, object storage conceptsOS & Tooling: Linux, Git, CI/CD, monitoring and logging tools"EXPERIENCE:7-10 yearsCompany Benefits & CultureOpportunity to work with a dynamic team in a fast-paced environmentExposure to cutting-edge technologies and methodologiesSupportive and collaborative work cultureAppreciate your quick response and please feel free to reach me out for any query you may have.Thanks