Senior Platform Engineer
We are seeking a Senior Platform Engineer to build, administer, automate, secure, and operate enterprise Databricks and cloud data-platform environments. The role is responsible for Databricks workspace administration, Unity Catalog governance, identity and access management, compute and cluster policies, infrastructure automation, CI/CD enablement, monitoring, reliability, cost management, and production support.
The ideal candidate combines hands-on Databricks administration with strong AWS or Azure platform engineering, Terraform, Python, Kubernetes, CI/CD, security, networking, observability, and Site Reliability Engineering practices.
Key Responsibilities
Databricks Administration and Platform Engineering
- Administer Databricks accounts, workspaces, metastores, catalogs, schemas, external locations, storage credentials, connections, shares, recipients, and platform configurations.
- Provision and manage development, test, staging, and production workspaces using standardized, repeatable patterns.
- Configure workspace settings, repositories, jobs, notebooks, SQL warehouses, instance pools, job clusters, all-purpose compute, and serverless capabilities.
- Define and enforce cluster policies, approved runtime versions, libraries, init scripts, autoscaling, tagging, and compute-usage standards.
- Manage Databricks Runtime and platform upgrades, compatibility testing, release planning, maintenance windows, and rollback procedures.
- Support Databricks Workflows, Delta Live Tables or Lakeflow Declarative Pipelines, Databricks SQL, MLflow, model registry, feature engineering, and structured-streaming services.
- Troubleshoot workspace, permissions, connectivity, compute, storage, job execution, library, runtime, and performance issues.
- Maintain administration standards, platform documentation, knowledge articles, support procedures, and operational runbooks.
Unity Catalog, Identity, Security, and Governance
- Design and manage Unity Catalog metastores, catalogs, schemas, managed and external tables, volumes, external locations, storage credentials, and grants.
- Implement user and group provisioning through single sign-on, SCIM, identity-provider integration, and enterprise directory services.
- Administer account-level and workspace-level users, groups, service principals, permissions, entitlements, and access-control models.
- Implement role-based and attribute-based access controls, least-privilege permissions, separation of duties, and privileged-access procedures.
- Configure secure access patterns for secrets, tokens, credentials, service principals, private endpoints, storage, and external systems.
- Enable audit logging, lineage, system tables, tagging, data classification, row-level security, column masking, and compliance reporting.
- Partner with security and governance teams on encrypti on, key management, network controls, data loss prevention, retention, auditability, and regulatory requirements.
- Review access, monitor privileged activities, remediate policy violations, and support internal and external audits.
Cloud Infrastructure and Networking
- Build and operate Databricks on AWS or Azure, including secure integration with cloud storage, identity, networking, encryption, and monitoring services.
- On AWS, work with S3, IAM, KMS, VPC, PrivateLink, security groups, Route 53, CloudWatch, Secrets Manager, and related services.
- On Azure, work with ADLS Gen2, Microsoft Entra ID, managed identities, Key Vault, virtual networks, private endpoints, network security groups, Azure Monitor, and related services.
- Configure control-plane and data-plane connectivity, private networking, DNS, routing, firewall, proxy, and egress controls.
- Integrate Databricks with cloud data lakes, APIs, databases, message platforms, and enterprise applications.
- Support Kubernetes, Docker, EKS or AKS, and container-based platform services where required.
- Contribute to capacity planning, disaster recovery, high availability, backup, restoration, and business-continuity exercises.
Infrastructure as Code and Automation
- Build and maintain reusable Terraform modules using cloud and Databricks providers.
- Automate workspace, network, Unity Catalog, storage, identity, compute-policy, cluster, job, permission, and monitoring configurations.
- Manage Terraform state, workspaces, variables, modules, versioning, policy checks, drift detection, and controlled promotion across environments.
- Use Python, shell scripting, Databricks CLI, REST APIs, and SDKs to automate administrative and operational tasks.
- Implement self-service workspace, catalog, schema, access, and compute vending with appropriate approval and governance controls.
- Maintain configuration standards and reduce manual administration through repeatable automation.
DevOps, CI/CD, and Release Engineering
- Design and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or Harness.
- Automate deployment of notebooks, jobs, workflows, libraries, policies, infrastructure, and platform configuration.
- Support Databricks Asset Bundles, Git integration, artifact management, environment promotion, testing, approvals, and rollback.
- Integrate security scanning, policy validation, infrastructure testing, and release evidence into delivery pipelines.
- Enable engineering teams through templates, reusable pipelines, documentation, and self-service platform capabilities.
- Partner with application, data-engineering, and DevOps teams to ensure deploy ment standards are consistent and supportable.
Reliability, Monitoring, Operations, and FinOps
- Establish monitoring, alerting, dashboards, logs, metrics, traces, and health checks for Databricks and connected cloud services.
- Use CloudWatch, Azure Monitor, Datadog, Splunk, New Relic, or similar platforms to monitor availability, compute utilization, failures, security events, and cost.
- Define platform service-level indicators, service-level objectives, operational metrics, and error budgets.
- Lead incident response, problem management, root-cause analysis, corrective actions, and post-incident reviews.
- Manage vulnerability remediation, runtime patching, dependency updates, security exceptions, and platform lifecycle activities.
- Optimize cluster sizing, autoscaling, pools, SQL warehouses, job concurrency, serverless usage, storage, and workload scheduling.
- Implement budget controls, chargeback or showback tagging, utilization reporting, anomaly detection, and cost-optimization recommendations.
- Participate in operational support rotations and maintain escalation paths with Databricks and cloud providers.
- Coordinate platform upgrades, disaster-recovery tests, security reviews, and production-readiness assessments.
Collaboration and Technical Leadership
- Partner with architecture, data engineering, security, cloud, network, governance, FinOps, and service-management teams.
- Advise engineering teams on Databricks platform standards, secure patterns, deployment models, performance, and cost.
- Conduct technical reviews and ensure solutions meet enterprise architecture and operational-support requirements.
- Mentor platform engineers and administrators and lead knowledge-transfer sessions.
- Communicate platform health, risks, dependencies, incidents, and improvement roadmaps to technical and business stakeholders.
- Drive continuous improvement in automation, reliability, security, developer experience, and operational efficiency.
Required Qualifications
- Typically 7-10 years of cloud, DevOps, Site Reliability Engineering, infrastructure, or platform-engineering experience.
- At least 3 years of hands-on Databricks platform administration in an enterprise environment.
- Strong experience administering Databricks workspaces, Unity Catalog, compute, cluster policies, jobs, SQL warehouses, permissions, and service principals.
- Strong experience with AWS or Azure infrastructure, identity, storage, networking, encryption, monitoring, and private connectivity.
- Strong proficiency with Terraform and infrastructure-as-code practices.
- Experience with Python, shell scrip