Lead PySpark Developer

  • $68 Per Hour
Want to know if you’re a fit?
Upload your resume and let our AI show you.

Skills

  • Amazon Web Services (AWS)unmatched
  • Apache Sparkunmatched
  • Best Practicesunmatched
  • Big Dataunmatched
  • Cachingunmatched
  • Cloud Computingunmatched
  • Code Reviewsunmatched
  • Computer Skillsunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Data Modelingunmatched
  • Data Scienceunmatched
  • Database Extract Transform and Load (ETL)unmatched
  • DevOpsunmatched
  • Distributed Computingunmatched
  • Dockerunmatched
  • IBM DB2unmatched
  • Information/Data Security (InfoSec)unmatched
  • Leadershipunmatched
  • Mentoringunmatched
  • NoSQLunmatched
  • Performance Tuning/Optimizationunmatched
  • PostgreSQLunmatched
  • Problem Solving Skillsunmatched
  • Python Programming/Scripting Languageunmatched
  • SQL Databasesunmatched
  • Snowflake Schemaunmatched
  • Source Code/Configuration Management (SCM)unmatched
  • Structured Dataunmatched
  • Team Lead/Managerunmatched
  • Unit Testunmatched
  • Unstructured Dataunmatched

Description

Requirements:
  • 7+ years of experience in Amazon Web Services (AWS) cloud computing.
  • 10+ years of experience in big data and distributed computing.
  • Strong hands-on experience with PySpark, Apache Spark, and Python.
  • Strong hands-on experience with SQL and NoSQL databases (DB2, PostgreSQL, Snowflake, etc.).
  • Proficiency in data modeling and ETL workflows.
  • Proficiency with workflow schedulers like Airflow.
  • Hands-on experience with AWS cloud-based data platforms.
  • Experience in DevOps, CI/CD pipelines, and containerization (Docker, Kubernetes) is a plus.
  • Strong problem-solving skills and ability to lead a team.
  • Experience with DBT and AWS Astronomer is a plus.
Responsibilities:
  • Lead the design, development, and deployment of PySpark-based big data solutions.
  • Architect and optimize ETL pipelines for structured and unstructured data.
  • Collaborate with clients, data engineers, data scientists, and business teams to provide scalable solutions.
  • Optimize Spark performance through partitioning, caching, and tuning.
  • Implement best practices in data engineering (CI/CD, version control, unit testing).
  • Work with cloud platforms like AWS.
  • Ensure data security, governance, and compliance.
  • Mentor junior developers and review code for best practices and efficiency.
MUST HAVE:
  • 7+ years of experience in AWS cloud computing.
  • 10+ years of experience in big data and distributed computing.
  • Experience with PySpark, Apache Spark, and Python.
  • Experience with SQL and NoSQL databases (DB2, PostgreSQL, Snowflake, etc.).
  • Hands-on experience with AWS cloud-based data platforms.
  • Experience in DevOps, CI/CD pipelines, and containerization (Docker, Kubernetes) is a plus.

Numbers & Facts

LocationOwings Mills, MD
Salary$68 Per Hour

Similar Jobs