

Search by job, company or skills

This job is no longer accepting applications
Key Responsibilities
• Ensure platform uptime and application health as per SLOs/KPIs
• Monitor infrastructure and applications using ELK, Prometheus, Zabbix, etc.
• Debug and resolve complex production issues, performing root cause analysis
• Automate routine tasks and implement self-healing systems
• Design and maintain dashboards, alerts, and operational playbooks
• Participate in incident management, problem resolution, and RCA documentation
• Own and update SOPs for repeatable processes
• Collaborate with L3 and Product teams for deeper issue resolution
• Support and guide L1 operations team
• Conduct periodic system maintenance and performance tuning
• Respond to user data requests and ensure timely resolution
• Address and mitigate security vulnerabilities and compliance issues Technical Skillset
• Hands-on with Spark, Hive, Cloudera Hadoop, Kafka, Ranger
• Strong Linux fundamentals and scripting (Python, Shell)
• Experience with Apache NiFi, Airflow, Yarn, and Zookeeper
• Proficient in monitoring and observability tools: ELK Stack, Prometheus, Loki
• Working knowledge of Kubernetes, Docker, Jenkins CI/CD pipelines
• Strong SQL skills (Oracle/Exadata preferred)
• Familiarity with DataHub, DataMesh, and security best practices is a plus
• Strong problem-solving and debugging mindset
• Ability to work under pressure in a fast-paced environment.
• Excellent communication and collaboration skills.
• Ownership, customer orientation, and a bias for actionKey Responsibilities
• Ensure platform uptime and application health as per SLOs/KPIs
• Monitor infrastructure and applications using ELK, Prometheus, Zabbix, etc.
• Debug and resolve complex production issues, performing root cause analysis
• Automate routine tasks and implement self-healing systems
• Design and maintain dashboards, alerts, and operational playbooks
• Participate in incident management, problem resolution, and RCA documentation
• Own and update SOPs for repeatable processes
• Collaborate with L3 and Product teams for deeper issue resolution
• Support and guide L1 operations team
• Conduct periodic system maintenance and performance tuning
• Respond to user data requests and ensure timely resolution
• Address and mitigate security vulnerabilities and compliance issues Technical Skillset
• Hands-on with Spark, Hive, Cloudera Hadoop, Kafka, Ranger
• Strong Linux fundamentals and scripting (Python, Shell)
• Experience with Apache NiFi, Airflow, Yarn, and Zookeeper
• Proficient in monitoring and observability tools: ELK Stack, Prometheus, Loki
• Working knowledge of Kubernetes, Docker, Jenkins CI/CD pipelines
• Strong SQL skills (Oracle/Exadata preferred)
Job ID: 114642855
Skills:
Load Balancing, Prometheus, OpenStack, Dns, Grafana, Devops, Jenkins, Gcp, Linux, Terraform, Ansible, Azure, Kubernetes, Python, AWS, Go, Sumo Logic, Elastic Stack, AI-assisted development tools, Site Reliability Engineering, Infrastructure Engineering, Networking fundamentals, Platform Engineering
Skills:
Bash Shell Scripting, Gitlab, Databricks, Python, AWS, monitoring and observability tools, Infrastructure-as-Code, CI CD
Skills:
Databricks, Gitlab, AWS, Bash Shell Scripting, Python, monitoring and observability tools, CI CD, Infrastructure-as-Code
Skills:
Powershell, Azure Databricks, Azure Data Factory, Terraform, Azure Data Lake, Microsoft Azure, Python, Azure DevOps, Azure SQL Data Warehouse, Virtual Network, Azure SQL Database, Azure Service Bus, Azure DNS, Blob Storage, Azure App Service, Azure Virtual Machines
Skills:
AWS Glue, Sql, Jfrog Artifactory, Databricks, Apache Airflow, Azure Data Factory, Lambda, AWS, Python, Bash, Azure, Gcp, Docker, Terraform, Jenkins, GitHub Actions, Step Functions