AWS Data & Solution Engineer
Role Name: AWS Data & Solution EngineerLocation: NYC, NYEmployment Type: Full-TimeMode – Hybrid (2-3 days from Office)Job SummaryAre you committed to building data solutions, and being a Data Management expert? Are you passionate about data? Would you like to work on solutions with tangible impact to our clients?We are looking for an Sr AWS Data & Solutions Engineer with primary skills on Python & PySpark development who will be able to design and build solutions for one of our Fortune 500 Client programs, which aims towards building an Enterprise Data Lake on AWS Cloud platform, build Data pipelines by developing several AWS Data Integration, Engineering & Analytics resources. You will be responsible for building API services using FastAPI or Flask frameworks.Key ResponsibilitiesDesign, build and unit test applications on Spark framework on Python.Build Python and PySpark based applications based on data in both Relational databases (e.g. Oracle), NoSQL databases (e.g. DynamoDB, MongoDB) and filesystems (e.g. S3, HDFS)Build AWS Lambda functions on Python runtime leveraging awswrangler, pandas, json, requestsBuild PySpark based data pipeline jobs on AWS Glue ETL or EMR ClustersBuild Python based event-driven integration with Kafka Topics, leveraging Confluent libsLeveraged Apache Iceberg to manage schema evolution and ACID-compliant CDC merges within the data lakeDesign and Build API services using FastAPI, understand the swagger metadata files and implement OAuth2/JWT authentication for protected endpointsBuild the process orchestration pipelines using AWS Step Functions and Eventbridge rules.Optimize performance for data access requirements by choosing the appropriate native Hadoop file formats (Avro, Parquet, ORC etc) and compression codec respectively.Deploy applications on Docker and Kubernetes containersLeverage copilot/GPT for agentic coding of above tech stackOptimize performance of Spark applications in Hadoop using configurations around Spark Context, Spark-SQL, Data Frame, and Pair RDD'sSetup the Glue crawlers to catalog OracleDB tables, MongoDB collections and S3 objectsAbility to monitor, troubleshoot and debug failures using AWS CloudWatch and DatadogAbility to solve complex data-driven scenarios and triage towards defects and production issuesParticipate in code release and production deployment.Create documentation for user adoption, deployments, runbook, and support client users for enablement or for any issues encountered.Perform code reviews with the team and enable them to develop code for complex scenariosParticipate in the agile development process, and document and communicate issues and bugs relative to data standards in scrum meetingsWork collaboratively with onsite and offshore team.Voice the opinions to multiple teams and thus driving the entire initiative with strong leadership Education & ExperienceBachelor’s Degree or equivalent in computer science or related and minimum 10+ years of experienceCertified on one of - Solution Architect, Data Engineer or Data Analytics Specialty by AWSRequire hand-on experience on Python and PySpark programmingRequire hands-on experience on AWS S3, Glue ETL & Catalog, Lamba Functions, EventBridge, Step Functions, AthenaRequire hands-on experience on Kafka integrationsRequire hands-on experience working on different file formats i.e. avro, parquet, orc, json, xmlRequire hands-on experience on Python pandas, requests, boto3 moduleRequire hands-on experience in writing complex SQL queriesRequire hands-on experience using REST APIs using FastAPI or FlaskRequire hands-on experience building Agentic AI workflowsPreferred expertise on Snowflake, AWS Redshift & DynamoDBAbility to use AWS services, predict application issues and design proactive resolutionsRequire Technical Coordination skills to drive requirements and technical designRequires aptitude to help build skillset within organizationKnowledge, Skills & AbilitiesData pipelines using Python and PySpark on AWS Glue, EMR and lambda functions.Develop and secure RESTful APIs (FastAPI) on Docker/EKS containers and implement OAuth2/JWT authentication for protected endpointsHands-on experience with Apache Iceberg tables for cdc and latest snapshotsEvent based pipelines for consuming/publishing to/from Apache Kafka/MSKLead and communicate complex technical designs and leverage copilot/GPT for agentic coding of above tech stack