Search by job, company or skills

Service Engineer 2

Service Engineer 2

Providence India
Fresher
Not Disclosed
  • Posted a month ago
  • Be among the first 10 applicants

Job Description

2P Incident Commander - Job Description

Organization Background:

  • Service Engineering teamsdeliverreliable, secure, and scalable technology services across enterprise platforms, supportingavailability,automation,operational excellence,andcontinuous improvement across cloud, infrastructure,platform, monitoring, and service management ecosystems.
  • The team is also driving the adoption of AI, Copilot, Automation, and Monitoring capabilities to improve operational efficiency, trend analysis, and incident response processes.

What will you be responsible for

  • Primary Responsibility: Serve as a first-line Service Engineering professional responsible for monitoring alerts, triaging incidents, coordinating incident response as an Incident Commander whenrequired, and resolving L1 operational tasks across application and infrastructure environments.

    • Perform real-time alert monitoring, initial triage, classification, prioritization, and routing acrossapplication, infrastructure, cloud, platform, monitoring, and service management environments.

    • Act as the first point of response for incidents byvalidatingimpact, gathering key details, engaging resolver teams, and ensuringtimelycommunication and escalation.

    • Serve as Incident Commander for defined L1/L2 operational incidents by coordinating bridges, tracking actions, driving updates, and ensuring incident process adherence.

    • Resolve L1 operational tasks and service requests across application and infrastructure areas, including standard checks, restarts, access validations, health checks, basic configuration updates, and runbook-driven remediation.

    • Use monitoring dashboards, logs, alerts, SOPs, knowledge articles, and runbooks to diagnose recurring issues and restore services within agreed operational targets.

    • Document incident timelines, actions taken, resolution steps, handoffs, and improvement opportunities in the service management system.

    • Identifyrepetitive alerts, noisy events, process gaps, and automation opportunities to improve operational efficiency, service reliability, and incident response quality.

    • Lead End to End Major incident management process effectively.

    • Drive structured troubleshooting, escalation, and decision-making during high-pressure situations.

    • Clear stakeholder & leadership communication.

    • Accountable for the overall quality of the process and in compliance with the procedures, data models, policies, and technologies associated with the process.

      Key Performance Indicators:

      • Alert triage quality: Accurately classify, prioritize, and route alerts with minimal rework, ensuring alerts are actioned within defined operational response targets.

      • Incident response timeliness: Acknowledge incidents promptly,initiaterequired coordination, and support service restoration within agreed SLA or operational targets.

      • L1 resolution effectiveness: Resolve eligible L1 application and infrastructure tasks through SOPs, runbooks, and standard remediation steps while reducing unnecessary escalations.

      • Incident command discipline:Maintainclear ownership during assigned incidents by coordinating bridge calls, tracking actions, publishing updates, and ensuringtimelyescalation to resolver teams.

      • Documentation accuracy:Maintaincomplete andaccurateincident notes, timelines, troubleshooting steps, handoff details, and resolution summaries in the service management system.

      • Escalation and communication quality: Providetimely, concise, and audience-appropriate updates to stakeholders, resolver teams, and shift handoff teams.

      • Operational hygiene: Keep monitoring queues, incident backlogs, pending tasks, and shift handoffs current, reducing missed alerts, stale tickets, and delayed follow-ups.

      • Continuous improvement contribution:Identifyrecurring alerts, noisy events, process gaps, and automation opportunities that improve service reliability and operational efficiency

      What would your day look like

      • Start the shift by reviewing monitoring queues, open alerts, incident dashboards, service health signals, and pending L1 tasks.

      • Triage alerts byvalidatingsymptoms, checking logs or dashboards, correlating events, confirming business impact, anddeterminingnextactions.

      • Coordinate incident response by opening or joining bridge calls, engagingapplicationand infrastructure teams, tracking owners, and publishingtimelyupdates.

      • Execute runbook-based remediation for L1 application and infrastructure issues, including standard service checks, restart procedures, access checks, backup or job validations, and basic platform health verification.

      • Update incidents, tasks, handoff notes, dashboards, and knowledge articles withaccuratetroubleshooting details and resolution outcomes.

      • Collaborate with application, database, network, compute, cloud, security, and service desk teams to restore service quickly and reduce recurring operational issues

      • Identify critical impacting issues, Drive bridges effectively, pull SMEs and drive end to end incidents, document steps for troubleshooting, and send timely communications.
      • Coordinating with Service owners on repetitive issues and driving the root cause by following the 5Y method.
      • Works in conjunction with Continual Service Improvement (CSI)

Who are we looking for

  • 2P-level operations professional with foundational experience in IT operations, service desk, application support, infrastructure support, NOC, command center, monitoring, or incident management.

    • 4+ years of experience in monitoring, Incident management and all modules under ITIL
    • Good understanding of alert triage, incident lifecycle, escalation management, service restoration, ticket documentation, and operational handoffs.Basic working knowledge of applications, operating systems, servers, databases, networks, cloud platforms,monitoringtools, logs, and IT service management processes.
    • Ability to follow SOPs and runbooks to resolve L1 tasks, perform standard validations, communicate status clearly, and escalate when deeper technical intervention isrequired.
    • Exposure to tools such as ServiceNow, Azure Monitor, Grafana, Splunk, AppDynamics, Dynatrace, SolarWinds, Teams, Outlook, PowerShell, or similar monitoring and collaboration platformsispreferred.
    • Strong ownership, calmness under pressure, communicationdiscipline, attention to detail, collaboration mindset, and willingness to work in a shift-based operational model.
    • Exposure or knowledge on AI tools, Copilot, Automation platforms, or operational AI capabilities is preferred.
    • Experience working in NOC/Monitoring environments will be an added advantage.
    • Strong communication skills with excellent interpersonal skills both in written and verbal correspondence.
    • Flexible to work in shifts, weekends, and holidays.
    • A bachelor's degree in computer science or information science or related field.

Providencene
Providence's vision to create Health for a Better World aids us to provide a fair and equitable workplace for all in our employment, whether temporary, part-time or full time, and to promote individuality and diversity of thought and background, and acknowledge its role in the organization's success. This makes us committed towards equal employment opportunities, regardless of race, religion or belief, color, ancestry, disability, marital status, gender, sexual orientation, age, nationality, ethnic origin, pregnancy, or related needs, mental or sensory disability, HIV Status, or any other category protected by applicable law. In furtherance to our mission in building a more inclusive and equitable environment, we shall, from time to time, undertake programs to assist, uplift and empower underrepresented groups including but not limited to Women, PWD (Persons with Disabilities), LGTBQ+ (Lesbian, Gay, Transgender, Bisexual or Queer), Veterans and others. We strive to address all forms of discrimination or harassment and provide a safe and confidential process to report any misconduct.
Contact ouralso, read our.
Providence

More Info

Key Skills

Copilot Automation

AI tools

Microsoft Azure IaaS

About Company

Providence, one of the US's largest not-for-profit healthcare systems, is committed to high quality, compassionate healthcare for all. Driven by the belief that health is a human right and the vision, ‘Health for a better world', Providence and its 121,000 caregivers strive to provide everyone access to affordable quality care and services.

Similar Jobs

India
Skills:
Python, Data Factory, PowerShell, Power Bi, Data Warehouse, Azure DevOps, Lakehouse, Azure Cloud Infrastructure, Microsoft Fabric, Azure Automation, Microsoft Entra ID, Log Analytics, Infrastructure as Code, Azure Policy, rbac, Key Vault
Coimbatore, India
Skills:
Power Bi, Incident Management, Excel, Servicenow, Itil Processes, Change Management, Major Incident Management, Problem Management, Service dashboards, SLA KPI governance, Monitoring, Operational reporting