Data Engineer
Data EngineerLocation: McLean, VA (Tuesday and Wednesday onsite)Duration: 6+ Months contract (with possible conversion)Job Description:Cleanse, manipulate and analyze large datasets (Structured and Unstructured data – XMLs, JSONs, PDFs) using Hadoop platform.Develop Python, Py Spark, Spark scripts to filter/cleanse/map/aggregate data.Manage and implement data processes (Data Quality reports)Develop data profiling, deduping logic, matching logic for analysisProgramming Languages experience in Python, Hive Query Language, Py Spark and Spark for data ingestionProgramming experience in RDBMS and Bigdata using Hive, Hadoop platformPresent ideas and recommendations on Hadoop and other technologies best use to managementGood understanding of Functional programming and Object-oriented programmingDetail oriented. Excellent communication skills (verbal and written)Must be able to manage multiple priorities and meet deadlinesRequired Skills:5+ years of experience in processing large volumes and variety of data (Structured and unstructured data, writing code for parallel processing, XMLS, JSONs, PDFs) - MandatoryProgramming experience in Python, Spark, SQL, HQL for data processing and analysisStrong SQL experience3+ years of experience – using Hadoop platform and performing analysis. Familiarity with Hadoop cluster environment and configurations for resource management for analysis workDesired Skills:2+ years of experience with containerization and orchestration. Hands on experience with GIT, AWS, Kubernetes, Docker etc.