Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

PubMatic
3-6 Years
Not Disclosed
Quick Apply
  • Posted a month ago
  • Over 200 applicants have applied

Job Description

As an SRE Engineer, you will be responsible for ensuring the seamless operation and optimal performance of large-scale distributed software applications. Your role revolves around maintaining a robust and high-performing environment, contributing to the reliability of our services, and innovating solutions to guarantee 24/7 availability. You will leverage your technical expertise to maintain a seamless experience for our users while upholding the highest standards of operational excellence.

Responsibilities:

Monitoring and Alerting:

  • Review existing monitoring tools and set up new systems to track system performance and key metrics.

Incident Management:

  • Monitor alerts and logs to promptly identify incidents or anomalies.
  • Prioritize incidents based on severity and potential impact on stability and reliability.
  • Engage in incident resolution, applying necessary fixes and mitigations to restore normal operations.

On-Call Responsibilities:

  • Organize on-call schedules to ensure 24/7 coverage for incident response.
  • Respond to alerts, troubleshoot issues, and coordinate with NOC and Engineering teams for incident resolution.
  • Conduct post-incident reviews to identify root causes, learn from incidents, and implement preventive measures.

Automation and Tooling:

  • Review and build new automation scripts and tools to streamline tasks, enhance efficiency, and reduce manual errors.
  • Regularly update and maintain monitoring, deployment, and incident management tools.

Performance Optimization:

  • Analyze application performance using profiling and monitoring tools to identify bottlenecks and areas for improvement.
  • Work on optimizations, infrastructure upgrades, and architectural improvements to enhance system performance and efficiency.

Capacity Planning and Scaling:

  • Monitor resource utilization and trends to predict capacity needs and plan for scaling.
  • Scale resources (servers and databases) based on usage patterns and anticipated growth.
  • Automate the sizing process to ensure efficiency.

Disaster Recovery and Redundancy:

  • Develop and maintain disaster recovery plans to ensure business continuity.
  • Implement redundancy and failover strategies to minimize downtime and maintain service availability during failures.

Knowledge Sharing and Documentation:

  • Create and maintain comprehensive documentation for configurations, procedures, incidents, and best practices.
  • Foster a culture of knowledge sharing within the team through regular sessions and training programs.

Feedback Loop and Continuous Improvement:

  • Collect feedback from incidents, post-mortems, and NOC/Dev team interactions to identify areas for improvement.
  • Continuously iterate on processes, tools, and systems based on feedback to drive continuous improvement.

Collaboration and Communication:

  • Collaborate closely with Engineering and DC/NOC teams to align goals and priorities.
  • Ensure open communication within the team and with stakeholders, providing regular updates on incidents, progress, and initiatives.

Requirements:

  • Bachelor's degree in Computer Science or related disciplines.
  • 3+ years of experience in software application/product support.
  • Proficiency in programming using Go, Shell, or Python scripting languages.
  • Experience in technical engineering (preferred).
  • A proactive approach to identifying problems, performance bottlenecks, and areas for improvement.
  • Strong knowledge of Networking, Database (MySQL), Linux System concepts, and experience in debugging and analyzing core dumps.
  • Hands-on experience with monitoring and observability tools like Grafana, Nagios, Influx, ELK, etc.
  • Familiarity with orchestration tools like Docker, Grafana, and incident management systems like Zenduty.
  • Excellent communication and collaboration skills with the ability to work effectively across teams.
  • Self-motivated with a positive mindset to examine and solve incidents.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company