JOBSEARCHER

Big Data Engineer

Title: Big Data EngineerLocation: Rockville, MD or McLean, VA (Hybrid)Contract: 6+ Months ContractOnly Local candidates who can take Assessment and only who are in DC/VA/MD who can got for F2F interviewOverviewMust haves: spark, Hadoop, scala, hivescripting is a must- python or perlmust be expert level in Complex SQL- window functioning, complex multiple joins, cloud experience is mandatory-S3, glue, emr, athenaAI- How to use AI for prompt engineeringGithubCopiliotChjatgoptQNeed someone who is well versed in agile, test automations, CICD practicesFinancial Experience Is PreferredROLE FIT5+ years building enterprise-scale data solutions using Spark, Hadoop, Hive, and ScalaStrong scripting skills (Python or Perl) and expert-level complex SQL (window functions, multi-joins)AWS cloud experience required (S3, EMR, Glue, Athena)Experience with Agile delivery, CI/CD pipelines, automated testing, and GitHub workflowsFinancial services or regulated industry experience preferredOBJECTIVESDesign and maintain scalable, reliable big data pipelinesOptimize Spark/Hadoop workloads for performance, scalability, and cost efficiencyImplement automated testing and data quality validationEnable analytics and data science teams with high-quality, accessible datasetsLeverage AI-assisted tools (Copilot, ChatGPT, Q Developer) to improve development productivityPROBLEM-SOLVINGDiagnose and resolve Spark performance bottlenecks and data pipeline failuresOptimize complex SQL transformations and large-scale joinsTroubleshoot data quality, latency, and reliability issues in productionImprove AWS workload efficiency through tuning and resource optimizationAutomate repetitive engineering tasks using AI-assisted development toolsExperience ValidationDelivered end-to-end pipelines using Spark and Hadoop ecosystem toolsOptimized SQL and pipeline performance with measurable improvementsDeployed and supported AWS data workloads (EMR, Glue, Athena, S3)Implemented CI/CD and automated testing for data pipelinesUsed AI coding assistants and GitHub workflows in team-based developmentSkills To ProbeSpark internals, performance tuning, and partitioning strategiesAdvanced SQL techniques and Hive/Trino optimizationAWS data architecture and cost/performance tuning practicesPrompt engineering and safe use of AI coding assistantsCollaboration and communication in fast-paced, cross-functional environments'