Data Engineer Cloud Data Pipelines, Integration & Analytics Enablement
Johns Creek, GA (hybrid)
4+ Months Contract
Summary
The Data Engineer will support data engineering initiatives by building reliable cloud-based data pipelines, scalable data structures, and analytics-ready datasets. This role is responsible for integrating data from multiple internal sources, automating data ingestion and transformation, creating reusable data models, and enabling downstream analytics, dashboards, reporting, and GenAI-enabled applications. The role partners closely with data scientists, analysts, business stakeholders, and platform teams to ensure data is accessible, well-structured, documented, and optimized for decision-making within an AWS cloud environment.
Responsibilities
- Design, build, and maintain cloud-based data pipelines for structured, semi-structured, and unstructured data sources.
- Develop automated data ingestion, transformation, and refresh workflows using AWS services including S3, Glue, Athena, Lambda, Step Functions, DynamoDB, relational databases, and Python.
- Create curated datasets, reusable schemas, metadata tables, and data models that support analytics, reporting, dashboards, and application development.
- Design data models that establish reliable relationships across business entities using identifiers, reference tables, and relational structures.
- Develop SQL queries, views, and data access layers for recurring analytical and reporting requirements.
- Partner with data scientists and analysts to prepare trusted datasets for analytics, machine learning, GenAI workflows, dashboards, and prototype applications.
- Implement data quality checks, validation rules, exception handling, logging, monitoring, and operational controls for data pipelines.
- Document data sources, transformations, refresh schedules, metadata, assumptions, and known limitations.
- Support the migration of manual and file-based processes to scalable, automated cloud data pipelines.
- Collaborate with platform, infrastructure, and security teams to ensure compliance with enterprise standards for data access, governance, and operational reliability.
Qualifications- Bachelor's degree in Computer Science, Data Engineering, Information Systems, Software Engineering, Engineering, Applied Mathematics, or a related technical field.
- 5 8 years of experience in data engineering, analytics engineering, cloud data platforms, ETL/ELT development, database design, or data integration.
- Hands-on experience building cloud data pipelines and data lake solutions using AWS services such as S3, Glue, Athena, Lambda, Step Functions, DynamoDB, and relational databases.
- Strong proficiency in PySpark, Python, and SQL for data extraction, transformation, validation, automation, and loading.
- Experience working with structured and semi-structured data sources including CSV, Excel, JSON, APIs, databases, and file-based data.
- Experience designing reusable data models, metadata structures, reference tables, and relational schemas.
- Experience creating analytics-ready datasets that support reporting, dashboards, and application development.
- Knowledge of data quality, pipeline monitoring, logging, validation, and operational best practices.
- Strong collaboration and communication skills with cross-functional technical and business teams.
Preferred Qualifications- Master's degree in Data Engineering, Cloud Architecture, Analytics Engineering, Enterprise Data Platforms, or a related field.
- Experience designing scalable data architecture patterns and reusable enterprise data models.
- Experience supporting GenAI-enabled applications and analytics workflows.
- Experience improving operational reliability and scaling prototype data pipelines into production-ready data products.
Metasys Technologies is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identify, national origin, veteran or disability status.