Job Title: Lead Data Engineer Data & AI, Supply Chain
Location: San Francisco, CA 94105 (Onsite)
Duration: 6 Months Contract
About the Role
Client is seeking an experienced Lead Data Engineer to join the Supply Chain Data & AI Organization. This role will support the design, development, and delivery of enterprise data products and analytics solutions across the Sourcing, Transportation, and Warehouse Management (WMS) domains.
The ideal candidate is a hands-on technical leader with deep expertise in building modern cloud-native data platforms on Google Cloud Platform (GCP). You will collaborate with Product Managers, Solution Architects, Data Architects, Business SMEs, and engineering teams to develop scalable, high-quality data solutions that enable advanced analytics, and AI-driven decision making.
Key Responsibilities
" Design, develop, and implement scalable data pipelines and data products on Google Cloud Platform (GCP).
" Build and optimize enterprise data solutions using Dataproc, BigQuery, SQL, and dbt.
" Design robust and scalable data models that support analytical and operational reporting requirements.
" Develop efficient ETL/ELT pipelines to ingest, transform, and publish data from multiple enterprise systems.
" Collaborate with Product Managers, Business Analysts, Enterprise Solution Architects, Data Architects, and business stakeholders to translate business requirements into scalable technical solutions.
" Lead technical design discussions and perform code reviews to ensure engineering quality and adherence to standards.
" Optimize data processing performance, reliability, scalability, and cost across cloud-based data platforms.
" Implement monitoring, testing, and operational best practices to support production workloads.
" Contribute to reusable frameworks, engineering standards, and documentation that improve team productivity and solution consistency.
" Support production issue resolution and continuous improvement initiatives.
" Work effectively within Agile delivery teams and participate in sprint planning, estimation, and backlog refinement.
" Mentor team members
Required Technical Skills
" 8 years of experience in Data Engineering with demonstrated technical leadership on enterprise data projects.
" Strong hands-on experience with Google Cloud Platform (GCP).
" Expert-level proficiency in:
o Dataproc
o BigQuery
o SQL
o dbt (Data Build Tool)
" Strong understanding of modern ETL/ELT architecture and large-scale data processing.
" Strong knowledge of data modeling techniques, including dimensional modeling, normalized data models, and analytical data warehouse design.
" Experience building scalable and maintainable cloud-native data pipelines.
" Experience with Git, CI/CD pipelines, and engineering best practices.
" Strong analytical, troubleshooting, and problem-solving skills.
" Excellent verbal and written communication skills with the ability to collaborate effectively across cross-functional teams.
Preferred Technical Skills
" Experience with Apache Airflow for workflow orchestration.
" Experience integrating enterprise data platforms with Apache Kafka or other streaming technologies.
" Working knowledge of PySpark for distributed data processing.
" Proficiency in Python for data engineering, automation, and utility development.
" Familiarity with data quality, metadata management, and data governance best practices.
Domain Experience (Highly Desirable)
Candidates with experience in one or more of the following areas will be strongly preferred:
" Retail industry (Apparel)
" Supply Chain data platforms
" Transportation and Logistics
" Warehouse Management Systems (WMS)
" Distribution Center operations
Desired Attributes
" Self-driven and able to work independently in a fast-paced environment.
" Strong ownership mindset with a focus on delivering high-quality solutions.
" Ability to balance technical excellence with business priorities.
" Effective collaborator who can work seamlessly with business partners, architects, product managers, and engineering teams.
" Passion for building scalable, reliable, and reusable data solutions that enable analytics and AI capabilities across the Supply Chain organization.