Lead enterprise-scale data platform modernization initiatives, driving migration from AWS EMR, Apache NiFi, and legacy ETL frameworks to the Databricks Lakehouse Platform.
Architect and implement scalable Lakehouse solutions using Databricks, Delta Lake, Unity Catalog, and Databricks Workflows.
Design and govern end-to-end data pipelines for batch, streaming, CDC, and real-time data integration workloads.
Provide architecture leadership for large-scale AWS-based data ecosystems leveraging S3, IAM, Redshift, Glue Catalog, Airflow, and Databricks.
Develop enterprise data architecture standards covering data modelling, metadata management, lineage, governance, security, and compliance.
Drive adoption of Databricks best practices including Delta Live Tables (DLT), Auto Loader, Unity Catalog, Serverless Compute, Lakehouse Federation, and advanced optimization techniques.
Lead replatforming and migration programmes involving PySpark applications, Airflow DAGs, EMR workloads, Redshift integrations, and NiFi pipelines.
Partner closely with business stakeholders, enterprise architects, product owners, and data science teams to define technology roadmaps and target architectures.
Establish governance frameworks using Unity Catalog, data quality controls, observability, monitoring, and operational excellence practices
Provide technical leadership, mentoring, architecture reviews, design governance, and solution sign-offs across multiple delivery teams.
Lead architecture workshops, executive presentations, solution assessments, and technology evaluations.
Mandatory Technical Skills
Databricks Lakehouse Platform
Delta Lake, Delta Live Tables (DLT)
Unity Catalog
Databricks Workflows
Auto Loader
PySpark, Spark SQL, Python
AWS (S3, IAM, Glue, Redshift, EMR, Lambda)
Apache Airflow
CDC & Data Migration Frameworks
Data Governance & Security
CI/CD, GitHub, Jenkins
JD:
Lead enterprise-scale Databricks Lakehouse Architecture design and implementation on AWS.
Drive large-scale data platform modernisation and cloud transformation initiatives.
Architect scalable Medallion Architecture (Bronze, Silver, Gold) data platforms.
Lead migration of legacy EMR, NiFi, Redshift, and ETL workloads to Databricks.
Design high-performance batch and real-time data processing solutions.
Build robust ingestion frameworks using Auto Loader, Delta Lake, and Structured Streaming.
Define enterprise data governance and security standards using Unity Catalog.
Architect metadata-driven and reusable PySpark-based ETL/ELT frameworks.
Establish best practices for performance tuning, scalability, and cost optimisation.
Design and implement data quality, lineage, and observability frameworks.
Drive adoption of CI/CD, DevOps, Infrastructure as Code, and automation practices.
Collaborate with business, analytics, and engineering teams to define target-state architectures.
Conduct architecture reviews and provide technical leadership across multiple projects.
Mentor architects and senior engineers on Databricks and AWS best practices.