Job Title - Hadoop Hive Python Developer
Experience - 8+ years
Location - Charlotte, NC
Hire Type - Fulltime (Onsite)
Primary skills: Hadoop, Hive, Python, PySpark, Apache Kafka, Hadoop Ecosystem, Hive, Databricks Lakehouse Architecture, Delta Lake, Bronze/Silver/Gold Data Modeling, Big Data ETL Pipeline Development, SQL, Real-time Data Ingestion Frameworks, Data Governance & Cataloging, CI/CD Tools Git, Jenkins, Bitbucket, Workflow Orchestration, and Cloud & On-Prem Big Data Platforms.
Key Responsibilities:
Design, develop, and optimize PySpark-based ETL pipelines running on on prem Hadoop clusters and cloud environments.
Build high volume ingestion frameworks using Kafka for real-time and near-real-time trading and market data.
Develop, tune, and manage Hadoop ecosystem components-HDFS, YARN, MapReduce, Tez, Oozie/Airflow.
Build high-performance, optimized Hive data models for regulatory reporting, trade lifecycle, and market risk processing.