NOC Engineer (Alert & Incident Management)
Athenahealth Technology Private Limited- Posted 9 hours ago
- Be among the first 10 applicants
Job Description
Join us as we work to create a thriving ecosystem that delivers accessible, high-quality, and sustainable healthcare for all.
Position Summary: We are looking for an NOC Engineer to join our Alert & Incident Management team, part of Cloud Infrastructure Engineering (Network Operations Centre). You will promote a teaching and learning culture within the team and serve as a NOC liaison to all internal stakeholders. As a NOC liaison, you will identify and execute opportunities to adopt new processes and improve operational practices within the team.
This is NOT a directly client-facing role. However, every action you take will be in the interest of delivering an amazing user experience for athenahealth clients and internal users alike.
About the Team: NOC (Alert & Incident Management) is a Global Operations team responsible for 24/7/365 support of the core athenaNet application (Collector & Clinicals), all supporting Platform and Product technology services, the entire supporting infrastructure across public and private cloud for core athenaNet, athenahealth Enterprise applications, and more services as the environment continues to grow.
The team is focused on availability, reliability, stability, scalability, customer experience, and solving large-scale distributed computing problems, while maintaining and improving the performance of the applications and infrastructure it supports.
The team ideally spends approximately 50% of its time reacting to issues and 50% of its time making the overall system better through continuous improvement, automation, process enhancement, and operational excellence.
The team operates within an agile environment and focuses on maintaining and evangelizing Infrastructure-as-Code, while continuously improving availability, scalability, reliability, and customer experience.
Essential Job Responsibilities
- Alert Management & Resolution - 40%: Monitor, identify, triage, investigate, and resolve alerts and operational issues in a timely manner.
- Perform incident management activities to restore services and minimize impact to applications, infrastructure, and users.
- Support 24/7/365 operational requirements for core athenaNet, supporting Platform and Product technology services, infrastructure, and Enterprise applications.
- Apply strong system administration and Linux administration skills to troubleshoot and resolve production issues.
- Work with database technologies such as Oracle, MySQL, PostgreSQL, and similar technologies when troubleshooting operational issues.
- Use Fault Management and Monitoring tools to identify, analyze, and respond to alerts and system events.
- Support applications and infrastructure in a production environment, maintaining availability, reliability, stability, and scalability.
- Incident Management - 30%: Coordinate incident response, troubleshooting, escalation, communication, and resolution across relevant teams.
- Work effectively with cross-functional groups and teams to achieve common operational and business goals.
- Serve as a NOC liaison to internal stakeholders, ensuring effective communication and coordination during operational events.
- Deliver clear communications to both technical and non-technical stakeholders.
- Continuous Improvement - 20%: Identify opportunities to improve processes, operational practices, system reliability, and team effectiveness.
- Promote a teaching and learning culture within the team.
- Identify and execute opportunities to adopt new processes and improve existing operational workflows.
- Contribute to Infrastructure-as-Code practices and the continuous improvement of the supported environment.
- Create and maintain technical documentation and Standard Operating Procedures (SOPs).
- Turnover, Team Meetings, Vendor Management & Ticket Work - 10%: Complete operational handovers, participate in team meetings, manage vendor-related activities, and perform ticket-related work.
- Follow established operational best practices and contribute to maintaining consistent and reliable operational processes.
Additional Job Responsibilities
- Participate in standard weekday shifts and rotational weekend shifts as required.
- Provide operational support outside of the regular schedule when business or operational needs arise.
- Maintain effective shift turnover and handoff to ensure continuity of operations.
- Work with internal teams and vendors to coordinate issue resolution and operational activities.
- Support the adoption of new tools, processes, and technologies that improve operational efficiency.
- Contribute to knowledge sharing and continuous learning within the NOC team.
- Assist with maintaining accurate tickets, operational records, documentation, and incident information.
- Support initiatives focused on availability, reliability, stability, scalability, and customer experience.
- Help improve the overall system by identifying recurring issues and opportunities for operational improvement.
Expected Education & Experience
- 3-7+ years of professional experience in an IT/Technology environment.
- Bachelor's or Master's Degree in a relevant field, such as B.E., B.Tech., or B.C.A.
- Strong System Administration and Linux Administration skills.
- Previous experience working with database technologies, including Oracle, MySQL, PostgreSQL, or similar databases.
- Proven operational background with knowledge of operational best practices.
- Experience working with Fault Management/Monitoring tools, such as: Assure1, Netcool, Zabbix, SolarWinds, Nagios / Icinga, Kibana, Grafana, Splunk, xMatters, WorldPing, AppDynamics, ThousandEyes, Prometheus, Datadog
- Working knowledge of Confluence, Jira, or similar ticketing systems.
- Experience working with cross-functional groups and teams to achieve common goals.
- Experience supporting a web application in a production environment.
- Experience creating technical documentation and SOPs.
- Experience delivering communications to both technical and non-technical stakeholders.
- Ability to work in an agile environment and support continuous improvement initiatives.
- Willingness and ability to work standard weekday shifts and rotational weekend shifts.
- Schedule flexibility: Every attempt will be made to maintain a consistent schedule for all team members however, business or operational needs may require you to work outside of your regular schedule.
-
More Info
Key Skills
ThousandEyes
xMatters
Fault Management Monitoring tools
Assure1
WorldPing
