Responsibilities
Develop and maintain distributed data pipelines for ETL and feature engineering
Build frameworks and tools (mostly python) to process data and generate modeling data more efficiently
Develop and translate map-reduce and machine learning algorithms
Identify and validate trends and provide summary statistics for massive data sets of both structured and unstructured data
Experienced with productionizing machine learning algorithms is a big plus
Experience on data management platform (DMP) is a big plus
Qualifications
3+ years of experience, software engineering, and scripting languages such as Python
2+ years of experience with hadoop technologies
1+ year of experience processing and transforming data (ETL)
Exposure to Nosql technologies
Experience with data modeling/scripting language such as Matlab, R is a plus
| Location | San Mateo, California |