AWS Cloud Data Lake Lead

Diverse Lynx, LLC

Cambridge, VT

JOB DETAILS
SKILLS
AWS Lambda, Adoption, Algorithms, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Apache Spark, Application Programming Interface (API), Artificial Intelligence (AI), Best Practices, Big Data, Career Counseling, Cisco Unity, Cloud Computing, Code Reviews, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Data Analysis, Data Lake, Data Management, Data Processing, Data Science, Distributed Computing, Docker, Electronic Medical Records, Git, Leadership, Machine Learning, Mentoring, Model Validation, Operational Strategy, Performance Analysis, Performance Testing, Predictive Modeling, Production Systems, Project Planning, Python Programming/Scripting Language, SQL (Structured Query Language), Standards Development, Structured Data, Team Lead/Manager, Technical Delivery, Training Data Sets, Unstructured Data
LOCATION
Cambridge, VT
POSTED
6 days ago

Location: Cambridge, VT

Duration: 6 months

Role Descriptions:

AWS Cloud Data Lake (Expert) Deep hands-on experience with AWS services SageMaker, S3, EMR, Glue, Lambda, Redshift, Athena, Step Functions, Lake Formation, and IAMsecurity best practices. Proven experience designing and managing AWS Data Lake architectures. Databricks (Expert) Proficiency in Databricks for large-scale data engineering, ML model development, MLflow for experiment tracking, and Unity Catalog for governance. Big Data Processing (Expert) Strong experience with Apache Spark (PySparkScala), distributed computing, and processing large-scale structuredunstructured datasets. Machine Learning Solid expertise in ML algorithms, feature engineering, model evaluation, and frameworks such as scikit-learn, XGBoost, TensorFlow, or PyTorch. MLOps Proven experience with end-to-end ML lifecycle model deployment, monitoring, retraining, CICD pipelines, containerization (Docker), and orchestration (KubernetesECS). Python Expert-level proficiency for end-to-end data science and ML workflows. SQL Advanced SQL for data analysis and pipeline development. Code Management Proficiency with Git, branching strategies, and code review practices.

Key responsibilities

Leadership Strategy Lead and manage a team of 6 data scientists, ML engineers, and analytics professionals across onshoreoffshore locations, providing technical mentorship and career guidance. Define and drive the data science operations strategy, roadmap, and best practices aligned with business objectives. Partner with senior business stakeholders, product owners, and cross-functional teams to identify high-impact AIML opportunities and translate them into actionable project plans. Establish and govern standards for model development, deployment, monitoring, and responsible AI adoption across the organization.Hands-On Technical Delivery Architect and oversee scalable ML pipelines for data ingestion, feature engineering, model training, validation, and inference on AWS cloud and Databricks. Design and implement AWS Data Lake architectures and big data processing solutions for structured and unstructured data at petabyte scale using Spark, Databricks, and AWS-native services (S3, Lake Formation, EMR, Glue, SageMaker, Redshift, Athena). Lead the deployment of production ML systems including real-time inference APIs, batch prediction pipelines, and model-as-a-service architectures. Drive MLOps maturity CICD for ML, automated model retraining, drift detection, AB testing, and performance monitoring.

+Additional GRS Parameter

Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.

About the Company

D

Diverse Lynx, LLC