Search by job, company or skills

Site Reliability Engineer

Early Applicant
  • Posted 16 hours ago
  • Be among the first 10 applicants

Job Description

JD for reliability Engineer ( NewRelic SME ) –

Location- Bangalore, Hyderabad, Pune, Noida, Chennai

  1. Monitoring and Automation: Proactively monitor software systems to prevent incidents and automate routine tasks.
  2. Effective Monitoring: Build monitoring systems that alert based on symptoms rather than outages.
  3. Application Performance Monitoring (APM): Implement and utilize APM tools such as New Relic or Dynatrace to monitor application performance, identify bottlenecks, and optimize resource usage.
  4. Log Analysis with Splunk: Analyse logs using Splunk to troubleshoot issues, detect anomalies, and improve system reliability.
  5. Dashboards Preparation: Create informative dashboards to visualize system health, performance, and key metrics.
  6. Alerts Setup: Configure alerts based on thresholds and anomalies to promptly address issues.
  7. Reports Scheduling: Set up regular reports to provide insights into system performance and reliability.
  8. Reliability Metrics: Establish and track reliability metrics (e.g., SLOs, SLIs, error budgets) to measure system performance.
  9. Observability Skills: Proficiency in observability practices, including distributed tracing, logging, and metrics collection.
  10. Collaboration: Partner with development, support teams to improve services through rigorous testing and release procedures.
  11. Capacity Planning: Participate in system design consulting and capacity planning.
  12. Debugging and Incident Response: Understand debugging information, handle incidents, and roll back faulty software pushes.
  13. Automation:Handson experience in automation using Python, Bash
  14. AIOps:Familiar with GenAI, Agentic AI frameworks. Understanding of leveraging agents in AIOps implementation across monitoring, observability, ITSM, Change/Incident Management, PagerDuty
  15. Mentoring L1/L2 Support Teams: Provide guidance, mentorship to establish best practices on monitoring and observability.
  16. Infrastructure Management: Run and manage infrastructure using tools like Chef, Ansible, Terraform, GitLab CI/CD, and Kubernetes.
  17. Documentation: Document processes and procedures to avoid redundancy.
  18. Enthusiastic Attitude: Approach challenges with enthusiasm and a proactive mindset.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152939787

Similar Jobs

Hyderabad, India

Skills:

ElkPrometheusBashGrafanaDatadogJenkinsGcpTerraformDockerAnsibleAzureKubernetesPythonAWSGoGitLab CIGitHub ActionsOpenTelemetry

Hyderabad, India

Skills:

Performance TestingMicroservicesJenkinsDockerTerraformAutomation FrameworksHelmKubernetesAzure DevOpsobservability frameworksIaCCI CDGitHub Actionschaos engineering

Hyderabad, India

Skills:

GitPowerShellBashMicrosoft AzureLog AnalyticsAzure Monitor

Hyderabad, India

Skills:

GithubPrometheusBashAmazon CloudWatchGrafanaLinux AdministrationTerraformIamKubernetesPythonAWSAmazon EKSSecrets managementOpenSearchArgo CDOpenTelemetry

Hyderabad, India

Skills:

GithubElkPowerShellGrafanaDatadogJenkinsCloudwatchTerraformDockerAnsibleKubernetesPythonSite24x7AWS cloud services

Beware of Scammers

We don’t charge money for job offers