Location: NYC, NY Employment Type: Full-Time Mode: Hybrid (2-3 days from Office)
Job Summary
Are you committed to building data solutions, and being a Data Management expert? Are you passionate about data? Would you like to work on solutions with tangible impact to our clients?
We are looking for an Sr AWS Data & Solutions Engineer with primary skills on Python & PySpark development who will be able to design and build solutions for one of our Fortune 500 Client programs, which aims towards building an Enterprise Data Lake on AWS Cloud platform, build Data pipelines by developing several AWS Data Integration, Engineering & Analytics resources. You will be responsible for building API services using FastAPI or Flask frameworks.
Key Responsibilities
Design, build and unit test applications on Spark framework on Python.
Build Python and PySpark based applications based on data in both Relational databases (e.g. Oracle), NoSQL databases (e.g. DynamoDB, MongoDB) and filesystems (e.g. S3, HDFS)
Build PySpark based data pipeline jobs on AWS Glue ETL or EMR Clusters
Build Python based event-driven integration with Kafka Topics, leveraging Confluent libs
Leveraged Apache Iceberg to manage schema evolution and ACID-compliant CDC merges within the data lake
Design and Build API services using FastAPI, understand the swagger metadata files and implement OAuth2/JWT authentication for protected endpoints
Build the process orchestration pipelines using AWS Step Functions and Eventbridge rules.
Optimize performance for data access requirements by choosing the appropriate native Hadoop file formats (Avro, Parquet, ORC etc) and compression codec respectively.
Deploy applications on Docker and Kubernetes containers
Leverage copilot/GPT for agentic coding of above tech stack
Optimize performance of Spark applications in Hadoop using configurations around Spark Context, Spark-SQL, Data Frame, and Pair RDD's
Setup the Glue crawlers to catalog OracleDB tables, MongoDB collections and S3 objects
Ability to monitor, troubleshoot and debug failures using AWS CloudWatch and Datadog
Ability to solve complex data-driven scenarios and triage towards defects and production issues
Participate in code release and production deployment.
Create documentation for user adoption, deployments, runbook, and support client users for enablement or for any issues encountered.
Perform code reviews with the team and enable them to develop code for complex scenarios
Participate in the agile development process, and document and communicate issues and bugs relative to data standards in scrum meetings
Work collaboratively with onsite and offshore team.
Voice the opinions to multiple teams and thus driving the entire initiative with strong leadership
Education & Experience
Bachelor's Degree or equivalent in computer science or related and minimum 10+ years of experience
Certified on one of - Solution Architect, Data Engineer or Data Analytics Specialty by AWS
Require hand-on experience on Python and PySpark programming