Eligability: This opportunity is open to candidates engaged on a C2C
We're looking for an experienced Databricks Platform Engineer to own, administer and optimise an enterprise Databricks platform hosted on AWS.
This is a hands-on platform engineering role focused on infrastructure, security, governance, automation and operational excellence rather than data engineering or analytics development.
You'll work closely with Cloud, DevOps, Security, Data Engineering and Analytics teams to ensure the Databricks platform is secure, scalable, reliable and cost-effective.
The successful candidate will have strong practical experience managing Databricks in enterprise production environments and be comfortable working across both platform administration and platform architecture.
What You'll Be Doing
Databricks Platform Administration
- Administer and support Databricks workspaces across Development, Test, UAT and Production.
- Design and maintain workspace standards, environment isolation and platform configurations.
- Manage Databricks Runtime (DBR) upgrades, version planning, testing and rollout.
- Maintain platform availability, reliability and performance.
- Support multiple workspaces and enterprise Databricks environments.
Compute & Cluster Management
- Configure and manage cluster policies, instance pools and job clusters.
- Implement compute governance and enforce organisational standards.
- Optimise cluster sizing, autoscaling, startup times and resource utilisation.
- Identify and prevent oversized or non-compliant compute configurations.
- Support performance and cost optimisation across the platform.
Unity Catalog, Governance & Security
- Administer Unity Catalog, including catalogs, schemas, tables and permissions.
- Manage Metastore architecture and enterprise data access.
- Implement RBAC and least-privilege access controls.
- Manage service principals, Personal Access Tokens (PATs) and Secret Scopes.
- Support enterprise authentication, authorisation and SSO.
- Maintain audit logging and support security and compliance requirements.
- Understand modern Databricks governance and data-sharing capabilities.
Data & Pipeline Platform Capabilities
- Support both batch and streaming workloads across the Databricks platform.
- Understand streaming architecture and implementation, including Watermarking and Change Data Feed (CDF).
- Support and administer Declarative Pipelines / DLT / Lakeflow.
- Understand operational considerations around pipeline deployment, monitoring and troubleshooting.
- Support Delta Sharing and external data-sharing mechanisms, including associated governance requirements.
- Work with teams using Lakehouse architecture and Delta Lake.
AWS & Cloud Infrastructure
- Integrate Databricks with AWS IAM and enterprise identity platforms.
- Work across AWS services including S3, VPC, EC2, KMS, CloudWatch and Security Groups.
- Support secure networking and connectivity across Databricks and AWS.
- Apply cloud security and infrastructure best practices.
- Experience with PrivateLink and VPC endpoints is advantageous.
DevOps, Terraform & Automation
- Build and maintain Databricks infrastructure using Terraform / Infrastructure as Code.
- Develop and maintain reusable Terraform configurations/modules where appropriate.
- Support Git integration through Databricks Repos.
- Implement and maintain CI/CD pipelines for notebooks, workflows and infrastructure.
- Work with DevOps tooling such as GitHub Actions, Azure DevOps or Jenkins.
- Automate platform provisioning, configuration and deployment wherever possible.
Monitoring, Operations & Troubleshooting
- Monitor workspace health, jobs, clusters and overall platform performance.
- Troubleshoot platform and production issues and perform root cause analysis.
- Analyse logs and operational data to identify platform issues.
- Support production releases, platform maintenance and incident management.
- Develop operational dashboards, alerts and monitoring processes.
- Drive continuous improvement across platform reliability and performance.
Cost Optimisation
- Monitor Databricks usage and AWS cloud spend.
- Optimise cluster sizing, autoscaling and compute utilisation.
- Optimise job scheduling and platform resource consumption.
- Identify opportunities to reduce Databricks and AWS costs.
- Apply FinOps principles to enterprise cloud environments.
What We're Looking For
- 5+ years' overall IT experience, with 3+ years of hands-on Databricks platform administration in production AWS environments.
- Strong hands-on experience administering and managing enterprise Databricks platforms.
- Experience managing multiple Databricks workspaces and production environments.
- Strong experience with Unity Catalog, Metastore architecture, RBAC and data governance.
- Hands-on experience with cluster management, cluster policies, instance pools and compute governance.
- Strong AWS experience, particularly IAM, S3, VPC, EC2, KMS and CloudWatch.
- Strong understanding of Databricks security, identity and access management.
- Hands-on experience with Terraform and Infrastructure as Code.
- Experience with Git, CI/CD and DevOps practices.
- Experience with platform monitoring, troubleshooting, incident management and production support.
- Strong understanding of platform performance and cloud cost optimisation.
- Ability to explain real-world Databricks implementation, architecture and troubleshooting scenarios, rather than purely theoretical knowledge.
- Comfortable working closely with Cloud, DevOps, Security and Data Engineering teams.
Preferred / Additional Experience
- DLT / Lakeflow Declarative Pipelines.
- Strong batch and streaming architecture experience.
- Watermarking and Change Data Feed (CDF).
- Delta Sharing and external data-sharing mechanisms.
- Hive Metastore and Unity Catalog migration experience.
- Lakehouse architecture and Delta Lake.
- AWS Organizations or Control Tower.
- PrivateLink and VPC Endpoints.
- Datadog, Splunk or other enterprise monitoring platforms.
- Apache Spark internals.
- Experience supporting AI/ML workloads on Databricks.
- FinOps / cloud cost optimisation experience.
- Databricks Certified Data Engineer Professional or Platform Administrator.
- AWS Solutions Architect or SysOps Administrator certification.
- Terraform Associate certification.