Databricks Platform Engineer
- Posted 8 hours ago
- Be among the first 10 applicants
Job Description
We are seeking a hands-on Databricks Platform Engineer to provide offshore operational support, platform administration, automation, monitoring, and continuous improvement for an enterprise Databricks Lakehouse platform on AWS.
The engineer will work closely with the onshore Senior Data Platform Engineer while independently managing assigned platform engineering and operational responsibilities. The role focuses on platform reliability, automation, incident management, security, governance, cost optimization, and supporting Data Engineering and Data Science teams.
Key Responsibilities
Databricks Administration | Databricks Production Support | AWS | Unity Catalog | Terraform | AWS IAM | VPC/Networking | S3 | KMS | PrivateLink | CloudWatch | Databricks REST API/CLI | Cost Optimization | Python/Shell | Monitoring & Incident Management | SQL/ETL
Skills: databricks admin,aws,terraform,production support
The engineer will work closely with the onshore Senior Data Platform Engineer while independently managing assigned platform engineering and operational responsibilities. The role focuses on platform reliability, automation, incident management, security, governance, cost optimization, and supporting Data Engineering and Data Science teams.
Key Responsibilities
- Provide day-to-day Databricks platform administration and production operational support.
- Monitor platform cost, usage, capacity, performance, and resource utilization.
- Identify and recommend opportunities for Databricks and AWS cost optimization.
- Implement approved platform architecture, engineering standards, security controls, and operational processes.
- Monitor platform health and respond to incidents, service requests, and operational alerts.
- Troubleshoot and restore platform services across Databricks, AWS, networking, security, compute, and workloads.
- Develop and maintain automation using Terraform, Databricks REST APIs, CLI, Python/Shell scripting, and other appropriate tools.
- Support platform deployments, configuration changes, environment promotion, and release activities.
- Administer Databricks workspaces, users, groups, permissions, compute, clusters, policies, and platform configurations.
- Manage Unity Catalog, including catalogs, schemas, external locations, storage credentials, grants, and access governance.
- Configure and manage compute policies and job policies.
- Support AWS-integrated Databricks services and troubleshoot connectivity and security issues.
- Assist Data Engineering and Data Science teams with platform-related issues, pipeline failures, compute problems, access issues, and workload troubleshooting.
- Perform initial investigation of complex platform issues and escalate to the appropriate engineering teams when required.
- Participate in proof-of-concept (POC) initiatives and evaluate emerging Databricks capabilities.
- Create and maintain runbooks, operational procedures, troubleshooting guides, and platform documentation.
- Provide effective shift handoffs and offshore/onshore operational coverage.
- Continuously identify opportunities to improve platform reliability, automation, operational efficiency, and support processes.
- 3+ years of hands-on experience administering and supporting Databricks in a production environment.
- Strong experience supporting enterprise-scale Databricks environments on AWS.
- AWS certification and Databricks certification are required.
- Strong hands-on knowledge of:
- Databricks Account and Workspace Administration
- Unity Catalog and access governance
- Compute and Job Policies
- Databricks cost and usage optimization
- Databricks security and platform administration
- Databricks REST APIs and CLI
- Terraform-based infrastructure automation
- Strong AWS knowledge across:
- AWS IAM
- VPC, networking and security
- KMS
- DynamoDB
- MSK
- S3
- AWS PrivateLink
- CloudWatch
- Experience with Python/Shell scripting and automation.
- Good understanding of data pipelines, SQL, ETL processes, and production workload troubleshooting.
- Experience with monitoring, incident management, root-cause investigation, and service restoration.
- Strong troubleshooting skills across platform, compute, networking, security, access, and workload layers.
- Good understanding of enterprise cloud security, governance, and operational best practices.
- Strong communication skills with the ability to work effectively with onshore engineering, Data Engineering, Data Science, and support teams.
- Experience with Databricks on AWS at enterprise scale.
- Experience with Databricks Jobs, Workflows, Delta Lake, SQL Warehouses, and serverless capabilities.
- Experience implementing Infrastructure as Code (IaC) using Terraform.
- Experience with CI/CD and DevOps tools such as Git, GitHub/GitLab, Azure DevOps, or Jenkins.
- Experience with Databricks cluster troubleshooting and Spark workload performance.
- Knowledge of AWS PrivateLink and secure Databricks networking architectures.
- Experience with platform observability and monitoring solutions.
- Experience participating in Databricks platform upgrades, migrations, and environment provisioning.
- Exposure to emerging Databricks features and platform POCs.
Databricks Administration | Databricks Production Support | AWS | Unity Catalog | Terraform | AWS IAM | VPC/Networking | S3 | KMS | PrivateLink | CloudWatch | Databricks REST API/CLI | Cost Optimization | Python/Shell | Monitoring & Incident Management | SQL/ETL
Skills: databricks admin,aws,terraform,production support
More Info
Key Skills
Monitoring Incident Management
Databricks REST API CLI
Databricks Administration
Databricks Production Support
Unity Catalog
SQL ETL
Python Shell
PrivateLink
VPC Networking




