Technical Stack & Requirements:
• Core Competencies: Python, Apache Spark (PySpark), AWS Glue ETL, Data Lake concepts (Medallion Architecture), Databricks, Sagemaker.
• Infrastructure: CloudFormation/Terraform, CI/CD with GitHub Actions.
• Core Responsibilities: Orchestrate data extraction from diverse legacy and modern sources, engineer high-performance scalable pipelines, and provide expert-level production support to ensure system reliability.
Qualifications:
• Bachelor’s degree in Computer Science or related field.
• Strong hands-on experience with Apache Spark and Glue.
• Proven ability to work with Python-based data pipelines.
• Experience with Git and GitHub Actions is mandatory.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or related discipline.
• 5-7+ years of enterprise Data Engineering experience.
• Strong Python development experience.
• Advanced SQL programming skills.
• Extensive experience with Apache Spark and PySpark.
• Strong experience with Databricks application.
• Hands-on experience with AWS data services including: o S3 o AWS Glue o Athena o Lambda o Step Functions o EventBridge
• Experience building enterprise ETL and ELT pipelines.
• Strong understanding of distributed data processing.
• Experience working with large-scale cloud data platforms.
• Strong knowledge of dimensional data modeling.
• Experience implementing CI/CD pipelines.
• Familiarity with Agile/Scrum methodologies.
• Excellent communication and stakeholder management skills. Preferred Qualifications
• Financial Services or Asset Management industry experience.
• Knowledge of Apache Iceberg, Delta Lake, or Hudi.
• Experience with Airflow.
• Knowledge of Kafka or Kinesis streaming.
• Experience supporting AI/ML and advanced analytics platforms.
• AWS Professional or Specialty Certifications.
| Location | Malvern, PA |
| Industry | Other/Not Classified |
| Company Size | 50 to 99 employees |
| Year Founded | 1997 |
| Website | https://www.vsoftconsulting.com/ |