Carry out data profiling and understand schema, data interrelationships, and data flows using SparkSQL, HiveQL, Jupyter Document test plans, writing test case automation and working closely with other teams (engineering, project management, etc.), bug reporting and isolation This position demands a self-motivated individual with strong technical and communication skills who can contribute in a team environment. Hive, HDFS, Azkaban, SparkSQL, HiveQL, CQL) 5+ yrs experience with near real-time (NRT) and Batch data pipelines Experience black box testing 5+ yrs experience Client-Server products Knowledge in Data Quality, Data Profiling and Data Integration tools.