
Search by job, company or skills
Description
What You'll Do –
Customer Support and Cloud Operations
• Monitor, support, and troubleshoot cloud-based services across Microsoft Azure and related platforms.
• Support incident response, service restoration, root cause analysis, and follow-up activities.
• Use dashboards, logs, alerts, and runbooks to identify issues, validate service health, and support customer-facing operations.
• Maintain accurate handover notes, operational documentation, and support records.
Customer Enablement and Delivery Support
• Partner with Cloud and Solution Architects to support effective customer onboarding and delivery execution.
• Assist customers by troubleshooting platform usage issues, onboarding queries, and initial implementation or integration steps.
• Serve as a technical liaison between customers and internal teams to ensure issues are clearly understood and resolved efficiently.
• Contribute to improved customer experience, customer satisfaction, and delivery confidence.
SLA Management and Support Operations
• Ensure support requests are managed in accordance with defined SLAs, severity levels, and operational processes.
• Provide customers with timely updates, clear communication, and effective resolution for critical incidents.
• Follow SOPs for ticket prioritization, incident response, escalation workflows, and customer communication.
• Maintain operational discipline to ensure support workflows remain reliable, predictable, and scalable.
Technical Troubleshooting and Escalation
• Support Kubernetes-based cloud services through basic health checks, deployment validation, and first-level troubleshooting.
• Troubleshoot common issues related to access, configuration, connectivity, workload failures, resource utilization, logs, and monitoring alerts.
• Hands on experience on Grafana, GitOps, no-SQL DB's, CI/CD tools, service mesh concepts and hybrid infrastructure.
• Escalate defects, feature gaps, or complex platform issues with clear context, reproduction steps, logs, and diagnostics.
Knowledge Management and Continuous Improvement
• Develop and maintain troubleshooting guides, support playbooks, root cause summaries, and customer-facing explanations.
• Identify recurring issue patterns and recommend improvements to reduce repeat incidents.
• Collaborate with Support, Engineering, Product, and Delivery/Customer Success teams to strengthen product quality and internal troubleshooting capability.
Shift-Based Support Coverage
• Provide cloud support coverage as part of a planned shift roster.
• Work rotational shifts in alignment with Customer Support coverage requirements and business needs.
• Provide weekend coverage when assigned, ensuring timely handovers and continuity of service.
What we're looking for:
• Approximately 3+ years of experience in cloud support, infrastructure operations, platform support, or a related technical support role.
• Working knowledge of Microsoft Azure and foundational exposure to Kubernetes, containers, or cloud-native services.
• Sound understanding of incident management, monitoring and operational support processes.
• Hands on experience on full-stack observability platforms like Grafana (including but not limited to querying logs, metrics, traces, setting up dashboards & alerts).
• Strong ability to analyze, maintain, and develop scripts utilizing PowerShell, Python, Bash, or equivalent tools to streamline operations.
• Strong troubleshooting, communication, ownership, and collaboration skills. Willingness to work rotational shifts and provide weekend support when required.
Success Measures
• Critical and complex support issues are resolved promptly, clearly, and efficiently.
• Customers experience minimal disruption, timely communication, and high service reliability.
• Customer satisfaction and response times show continuous improvement.
• Support knowledge, troubleshooting guides, and playbooks are structured for effective reuse.
• Recurring issues are identified, escalated, and reduced over time.
Requirements
Desirable Experience
• Exposure to AI, ML, data-intensive workloads, or GPU-enabled environments.
• Familiarity with Infrastructure as Code concepts (including Terraform, Bicep, ARM templates, or Pulumi).
• Experience supporting enterprise, industrial, regulated, or mission-critical environments.
• CKA certification will be an added advantage.
Job ID: 151737017
Skills:
Unix, Load Balancing, Apis, Routing, Http Protocol, Dns, Linux, Firewalls, Shell scripting, Python, Kubernetes, Container networking, Software-defined networking, Network namespaces, Packet captures, Network specific kernel parameters
Skills:
Splunk, Data Structure, Ml, Sql, Pl Sql, Ms Excel, Java, Python, Azure, Appdynamics, Mulesoft, Ai
Skills:
Vmware Vsphere, DHCP, Ticketing Tools, AWS, Dns, Azure, Group Policies, Itil, PowerShell, patching tools, Active Directory