The Role
Architect and operate large-scale distributed data platforms using Databricks and Spark.
What You Will Do
- Architect Databricks deployments
- Optimize Spark and Delta Lake workloads
- Implement Infrastructure as Code
- Lead incident response
- Drive performance standards
What You Will Have
- 6+ years of data or cloud engineering experience
- 2+ years operating Databricks in production
- 2+ years working with Spark
- 2+ years implementing Infrastructure as Code
- Production ownership experience
- Responsible for the end-to-end administration, support, and maintenance of enterprise data platforms, including Databricks and Cloudera (CDP/CDH).
- Manages Databricks workspaces, compute, Unity Catalog, security, storage integrations, and CI/CD, while supporting Cloudera services such as HDFS, YARN, Hive, Impala, Spark, Oozie, Ranger, Atlas, Kerberos, and TLS.
- Performs platform monitoring, incident troubleshooting, Spark job analysis, log investigation, disk space management, Linux-based operational support (services, processes, file systems, permissions, and logs), performance tuning, upgrades, automation (Ansible/chef/Terraform), capacity planning, and ensures high availability, security, governance, and operational excellence across AWS, Azure, GCP and on-perm Linux environments.
- On-Call Support
What Will Set You Apart
- Certifications: Databricks, Cloudera, AWS
- Streaming architecture experience
- AI/ML workload enablement
- FinOps cost optimization expertise