Request ID: 107895-1
Title: Sr Data Engineer
Location: Ofallon, MO
Duration: 6 months
Pay Range: $40 - $45/Hour on W2/C2C (All inclusive)
JOB DESCRIPTION:
We are looking for a highly skilled Senior Data Engineer with deep expertise in Apache Spark, Scala, and PySpark to build and operate large scale batch and streaming data processing systems. The role has a strong emphasis on real time streaming architectures using Kafka and Spark Structured Streaming, alongside ingestion and orchestration with Apache NiFi and scalable storage using Apache Ozone and Ceph. This position is ideal for engineers who enjoy solving complex performance, scalability, latency, and reliability challenges in production data platforms.
Key Responsibilities:
Design, develop, and maintain large scale Spark applications using Scala and PySpark
Build and operate streaming heavy data pipelines using Kafka and Spark Structured Streaming
Implement stateful streaming patterns including windowing, watermarking, late data handling, and checkpointing
Develop robust event replay and reprocessing workflows using Kafka offsets and partitions
Build ingestion and routing flows using Apache NiFi, including Kafka based ingestion patterns
Implement end to end ETL/ELT pipelines with strong emphasis on low latency, fault tolerance, and scalability
Optimize Spark jobs through partitioning strategies, memory tuning, shuffle optimization, and efficient data formats
Integrate Spark workloads with distributed object storage systems such as Apache Ozone and Ceph
Ensure data quality, consistency, and auditability through validation, reconciliation, and metadata capture
Collaborate with platform, infrastructure, and operations teams on production readiness and capacity planning
Support production systems, including monitoring, incident analysis, and root cause resolution
Contribute to reusable frameworks, coding standards, and engineering best practices
Participate in architecture reviews, code reviews, and technical documentation
Required Qualifications:
Bachelor’s degree in computer science, Engineering, or equivalent practical experience
Strong hands on experience with Apache Spark in production environments
Advanced proficiency in Scala and PySpark
Solid understanding of distributed systems and data processing at scale
Strong experience with Kafka based streaming architectures
Hands on experience with Spark Structured Streaming
Experience building batch and real time pipelines
Hands on experience with Apache NiFi for data ingestion and flow management
Strong SQL skills and experience working with structured and semi structured data
Experience working with object storage or distributed storage platforms
Proficiency with Linux, shell scripting, and Git based version control
Preferred Qualifications
Experience with Apache Ozone and/or Ceph as storage backends for analytics workloads
Experience implementing exactly once / at least once streaming semantics
Strong background in Spark performance tuning (CPU, memory, I/O, shuffle)
Experience supporting mission critical production systems with strict SLAs
Familiarity with CI/CD pipelines and automated testing for data applications
Experience designing observability for streaming systems (lag, throughput, backpressure)
Technical Skills
Languages: Scala, Python (PySpark), SQL
Big Data: Apache Spark (Core, SQL, Structured Streaming)
OS & Tooling: Linux, Git, CI/CD, monitoring and logging tools"
EXPERIENCE:
7-10 years
Company Benefits & Culture
• Opportunity to work with a dynamic team in a fast-paced environment
• Exposure to cutting-edge technologies and methodologies
• Supportive and collaborative work culture
Appreciate your quick response and please feel free to reach me out for any query you may have.
Thanks
Numbers & Facts
Location
Ofallon, MO
Salary
$40–$45 Per Hour
Skills
Apacheunmatched
Apache Kafkaunmatched
Apache Sparkunmatched
Best Practicesunmatched
Big Dataunmatched
CPU (Central Processing Unit)unmatched
Capacity Managementunmatched
Code Reviewsunmatched
Coding Standardsunmatched
Computer Scienceunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Data Formatsunmatched
Data Managementunmatched
Data Processingunmatched
Data Qualityunmatched
Database Extract Transform and Load (ETL)unmatched