Senior Platform Engineer

Tata Consultancy Services Ltd
  • Irvine, CA
    3 days ago

    Job Description

    Senior Platform Engineer

    We are seeking a Senior Platform Engineer to build, administer, automate, secure, and operate enterprise Databricks and cloud data-platform environments. The role is responsible for Databricks workspace administration, Unity Catalog governance, identity and access management, compute and cluster policies, infrastructure automation, CI/CD enablement, monitoring, reliability, cost management, and production support.

    The ideal candidate combines hands-on Databricks administration with strong AWS or Azure platform engineering, Terraform, Python, Kubernetes, CI/CD, security, networking, observability, and Site Reliability Engineering practices.

    Key Responsibilities

    Databricks Administration and Platform Engineering

    • Administer Databricks accounts, workspaces, metastores, catalogs, schemas, external locations, storage credentials, connections, shares, recipients, and platform configurations.
    • Provision and manage development, test, staging, and production workspaces using standardized, repeatable patterns.
    • Configure workspace settings, repositories, jobs, notebooks, SQL warehouses, instance pools, job clusters, all-purpose compute, and serverless capabilities.
    • Define and enforce cluster policies, approved runtime versions, libraries, init scripts, autoscaling, tagging, and compute-usage standards.
    • Manage Databricks Runtime and platform upgrades, compatibility testing, release planning, maintenance windows, and rollback procedures.
    • Support Databricks Workflows, Delta Live Tables or Lakeflow Declarative Pipelines, Databricks SQL, MLflow, model registry, feature engineering, and structured-streaming services.
    • Troubleshoot workspace, permissions, connectivity, compute, storage, job execution, library, runtime, and performance issues.
    • Maintain administration standards, platform documentation, knowledge articles, support procedures, and operational runbooks.

    Unity Catalog, Identity, Security, and Governance

    • Design and manage Unity Catalog metastores, catalogs, schemas, managed and external tables, volumes, external locations, storage credentials, and grants.
    • Implement user and group provisioning through single sign-on, SCIM, identity-provider integration, and enterprise directory services.
    • Administer account-level and workspace-level users, groups, service principals, permissions, entitlements, and access-control models.
    • Implement role-based and attribute-based access controls, least-privilege permissions, separation of duties, and privileged-access procedures.
    • Configure secure access patterns for secrets, tokens, credentials, service principals, private endpoints, storage, and external systems.
    • Enable audit logging, lineage, system tables, tagging, data classification, row-level security, column masking, and compliance reporting.
    • Partner with security and governance teams on encrypti on, key management, network controls, data loss prevention, retention, auditability, and regulatory requirements.
    • Review access, monitor privileged activities, remediate policy violations, and support internal and external audits.

    Cloud Infrastructure and Networking

    • Build and operate Databricks on AWS or Azure, including secure integration with cloud storage, identity, networking, encryption, and monitoring services.
    • On AWS, work with S3, IAM, KMS, VPC, PrivateLink, security groups, Route 53, CloudWatch, Secrets Manager, and related services.
    • On Azure, work with ADLS Gen2, Microsoft Entra ID, managed identities, Key Vault, virtual networks, private endpoints, network security groups, Azure Monitor, and related services.
    • Configure control-plane and data-plane connectivity, private networking, DNS, routing, firewall, proxy, and egress controls.
    • Integrate Databricks with cloud data lakes, APIs, databases, message platforms, and enterprise applications.
    • Support Kubernetes, Docker, EKS or AKS, and container-based platform services where required.
    • Contribute to capacity planning, disaster recovery, high availability, backup, restoration, and business-continuity exercises.

    Infrastructure as Code and Automation

    • Build and maintain reusable Terraform modules using cloud and Databricks providers.
    • Automate workspace, network, Unity Catalog, storage, identity, compute-policy, cluster, job, permission, and monitoring configurations.
    • Manage Terraform state, workspaces, variables, modules, versioning, policy checks, drift detection, and controlled promotion across environments.
    • Use Python, shell scripting, Databricks CLI, REST APIs, and SDKs to automate administrative and operational tasks.
    • Implement self-service workspace, catalog, schema, access, and compute vending with appropriate approval and governance controls.
    • Maintain configuration standards and reduce manual administration through repeatable automation.

    DevOps, CI/CD, and Release Engineering

    • Design and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or Harness.
    • Automate deployment of notebooks, jobs, workflows, libraries, policies, infrastructure, and platform configuration.
    • Support Databricks Asset Bundles, Git integration, artifact management, environment promotion, testing, approvals, and rollback.
    • Integrate security scanning, policy validation, infrastructure testing, and release evidence into delivery pipelines.
    • Enable engineering teams through templates, reusable pipelines, documentation, and self-service platform capabilities.
    • Partner with application, data-engineering, and DevOps teams to ensure deploy ment standards are consistent and supportable.

    Reliability, Monitoring, Operations, and FinOps

    • Establish monitoring, alerting, dashboards, logs, metrics, traces, and health checks for Databricks and connected cloud services.
    • Use CloudWatch, Azure Monitor, Datadog, Splunk, New Relic, or similar platforms to monitor availability, compute utilization, failures, security events, and cost.
    • Define platform service-level indicators, service-level objectives, operational metrics, and error budgets.
    • Lead incident response, problem management, root-cause analysis, corrective actions, and post-incident reviews.
    • Manage vulnerability remediation, runtime patching, dependency updates, security exceptions, and platform lifecycle activities.
    • Optimize cluster sizing, autoscaling, pools, SQL warehouses, job concurrency, serverless usage, storage, and workload scheduling.
    • Implement budget controls, chargeback or showback tagging, utilization reporting, anomaly detection, and cost-optimization recommendations.
    • Participate in operational support rotations and maintain escalation paths with Databricks and cloud providers.
    • Coordinate platform upgrades, disaster-recovery tests, security reviews, and production-readiness assessments.

    Collaboration and Technical Leadership

    • Partner with architecture, data engineering, security, cloud, network, governance, FinOps, and service-management teams.
    • Advise engineering teams on Databricks platform standards, secure patterns, deployment models, performance, and cost.
    • Conduct technical reviews and ensure solutions meet enterprise architecture and operational-support requirements.
    • Mentor platform engineers and administrators and lead knowledge-transfer sessions.
    • Communicate platform health, risks, dependencies, incidents, and improvement roadmaps to technical and business stakeholders.
    • Drive continuous improvement in automation, reliability, security, developer experience, and operational efficiency.

    Required Qualifications

    • Typically 7-10 years of cloud, DevOps, Site Reliability Engineering, infrastructure, or platform-engineering experience.
    • At least 3 years of hands-on Databricks platform administration in an enterprise environment.
    • Strong experience administering Databricks workspaces, Unity Catalog, compute, cluster policies, jobs, SQL warehouses, permissions, and service principals.
    • Strong experience with AWS or Azure infrastructure, identity, storage, networking, encryption, monitoring, and private connectivity.
    • Strong proficiency with Terraform and infrastructure-as-code practices.
    • Experience with Python, shell scrip

    Numbers & Facts

    LocationIrvine, CA

    Skills

    • Access Controlunmatched
    • Administrative Skillsunmatched
    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Application Programming Interface (API)unmatched
    • Automationunmatched
    • Autoscalingunmatched
    • Budgetingunmatched
    • Capacity Managementunmatched
    • Chargebacksunmatched
    • Cisco Unityunmatched
    • Cloud Computingunmatched
    • Cloud Storageunmatched
    • Computer Securityunmatched
    • Concurrencyunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Corrective Actionunmatched
    • Cost Controlunmatched
    • Cryptographyunmatched
    • DNS (Domain Name System)unmatched
    • Data Recoveryunmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Dockerunmatched
    • Documentationunmatched
    • Endpoint Securityunmatched
    • Engineeringunmatched
    • Enterprise Architectureunmatched
    • Enterprise Directory Servicesunmatched
    • External Auditunmatched
    • Firewallsunmatched
    • Gitunmatched
    • GitHubunmatched
    • High Availabilityunmatched
    • Identity Data Managementunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Information/Data Security (InfoSec)unmatched
    • Internal Auditunmatched
    • Jenkinsunmatched
    • Knowledge Transferunmatched
    • Loss Preventionunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Microsoft Product Familyunmatched
    • Microsoft Windows Azureunmatched
    • Network Connectivityunmatched
    • Network Routingunmatched
    • Network Securityunmatched
    • Operational Supportunmatched
    • Operations Processesunmatched
    • Performance Modelingunmatched
    • Process Improvementunmatched
    • Production Supportunmatched
    • Python Programming/Scripting Languageunmatched
    • REST (Representational State Transfer)unmatched
    • Regulatory Requirementsunmatched
    • Release Management/Engineeringunmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • SQL (Structured Query Language)unmatched
    • Scripting (Scripting Languages)unmatched
    • Security Attacksunmatched
    • Security Infrastructureunmatched
    • Security Policyunmatched
    • Single Sign-On (SSO)unmatched
    • Software Engineeringunmatched
    • Software Patchesunmatched
    • Splunkunmatched
    • Systems Administration/Managementunmatched
    • Technical Leadershipunmatched
    • Test Plan/Scheduleunmatched
    • Theater Productionunmatched
    • Training Data Setsunmatched
    • Unix Shell Programmingunmatched
    • VPN (Virtual Private Network)unmatched
    • Validation Testingunmatched
    • Warehousingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder