Search by job, company or skills

Principal Site Reliability Engineer

  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

Oracle Cloud Infrastructure (OCI) is looking for a Principal Site Reliability Engineer (IC4) to help build, operate, and evolve highly available, scalable, and resilient cloud services.

As a Principal SRE, you will take technical ownership of reliability across complex, distributed systems operating at cloud scale. You will work closely with software engineering, architecture, security, and operations teams to influence service design, improve availability and performance, automate operational work, and ensure our services meet their reliability objectives.

This role goes beyond operating production systems. You will identify systemic reliability risks, influence architecture and engineering decisions, lead complex incident investigations, build automation, improve observability, and drive long-term engineering improvements.

You will also serve as a technical leader within the team, mentoring engineers and helping establish strong SRE practices across services.

Key Responsibilities

Reliability Architecture & Capacity Engineering

  • Design and influence architectures for highly available, resilient, scalable, and operationally efficient cloud services.
  • Partner with software development teams during design and implementation to ensure reliability, scalability, observability, security, and operability are built into services from the beginning.
  • Identify architectural and operational risks across multiple services and drive engineering improvements to address them.
  • Define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), monitoring strategies, and reliability standards.
  • Forecast infrastructure and service capacity based on workload growth, utilization trends, architecture changes, and customer demand.
  • Identify capacity risks and bottlenecks before they impact customers and drive appropriate mitigation plans.
  • Lead technical prototypes and evaluations for new infrastructure, reliability patterns, and operational technologies.

Production Engineering & Service Lifecycle

  • Own and continuously improve the operational health of production services.
  • Analyze service telemetry, operational data, and reliability trends to identify systemic issues and improvement opportunities.
  • Drive improvements across availability, latency, performance, scalability, security, recoverability, and operational efficiency.
  • Establish mechanisms to detect reliability degradation before it becomes customer impacting.
  • Lead complex service lifecycle activities including upgrades, migrations, security updates, disaster recovery, capacity expansion, and decommissioning.
  • Identify recurring operational issues and convert them into engineering problems with sustainable solutions.

Automation & Toil Reduction

  • Identify high-impact opportunities to eliminate repetitive operational work through software engineering and automation.
  • Design and build scalable automation, tooling, and frameworks for deployment, monitoring, diagnostics, mitigation, remediation, and service lifecycle management.
  • Develop automated mechanisms for detecting and recovering from common failure scenarios.
  • Establish engineering standards for operational tooling, ensuring automation is reliable, testable, maintainable, observable, and safe.
  • Measure operational toil and drive initiatives that improve engineering efficiency and reduce manual intervention.
  • Review and improve automation developed by other engineers.

Observability & Performance Engineering

  • Define and improve observability strategies across services using metrics, logs, traces, dashboards, and alerting.
  • Develop meaningful service health indicators that accurately reflect customer experience.
  • Analyze production workloads to identify performance bottlenecks, resource inefficiencies, scaling limitations, and reliability risks.
  • Drive improvements to monitoring and alerting to improve signal quality and reduce operational noise.
  • Use production data and reliability trends to influence architecture, capacity planning, and engineering priorities.

Technical Leadership & Engineering Excellence

  • Provide technical leadership for reliability initiatives spanning multiple services or engineering teams.
  • Influence architecture and design decisions by identifying reliability, scalability, operational, and failure-mode considerations.
  • Lead technical discussions and design reviews for complex infrastructure and reliability challenges.
  • Establish and promote engineering best practices for operating large-scale distributed systems.
  • Mentor SREs and software engineers in troubleshooting, incident management, automation, observability, and reliability engineering.
  • Review designs, operational readiness, automation, and implementation approaches and provide actionable technical feedback.
  • Raise the overall technical and operational maturity of the team.

Cross-Team Collaboration & Communication

  • Partner with software engineering, architecture, security, networking, infrastructure, and operations teams to solve complex reliability problems.
  • Clearly communicate service health, operational risks, capacity constraints, incident impact, and reliability priorities to technical and non-technical stakeholders.
  • Anticipate the operational impact of infrastructure, architecture, feature, and tooling changes across multiple services.
  • Drive alignment across teams when reliability improvements require changes across organizational boundaries.
  • Provide clear technical recommendations supported by production data and engineering analysis.

Continuous Improvement & Innovation

  • Evaluate emerging technologies, engineering approaches, and SRE practices that can improve reliability, scalability, security, or operational efficiency.
  • Identify systemic weaknesses in existing operational processes and drive improvements.
  • Use operational data, incident trends, and engineering metrics to prioritize reliability investments.
  • Contribute reusable tools, patterns, standards, and best practices that benefit teams beyond your immediate area.
  • Stay current with developments in cloud infrastructure, distributed systems, observability, automation, and Site Reliability Engineering.

What We're Looking For

  • Strong experience operating and troubleshooting large-scale production systems.
  • Strong understanding of distributed systems, high availability, scalability, fault tolerance, and reliability engineering principles.
  • Experience designing or operating cloud infrastructure and services at scale.
  • Strong understanding of Linux systems, networking, storage, compute, and cloud infrastructure concepts.
  • Experience with monitoring, observability, metrics, logging, tracing, and alerting systems.
  • Experience defining or working with SLIs, SLOs, availability targets, and service health metrics.
  • Strong troubleshooting and debugging skills across application and infrastructure layers.
  • Experience leading or significantly contributing to complex production incident response and root cause analysis.
  • Experience identifying and eliminating operational toil through engineering and automation.
  • Ability to influence technical decisions and drive engineering initiatives across teams.
  • Strong written and verbal communication skills.
  • Demonstrated ability to mentor engineers and raise engineering standards within a team.

Preferred Qualifications

  • Experience working with large-scale cloud platforms such as Oracle Cloud Infrastructure (OCI), AWS, Azure, or GCP.
  • Experience with Kubernetes, containers, infrastructure-as-code, and modern deployment technologies.
  • Experience building automation and internal platforms for large-scale infrastructure operations.
  • Experience with capacity planning, performance engineering, disaster recovery, and resilience testing.
  • Experience designing highly available distributed systems and understanding complex failure modes.
  • Experience improving operational readiness and reliability across multiple services.
  • Experience driving reliability initiatives that span multiple engineering teams.

Career Level

IC4 - Principal Site Reliability Engineer


About Oracle

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. With AI embedded across our products and services, we help customers turn that promise into a better future for all.

At Oracle Cloud Infrastructure, you will have the opportunity to work on cloud services operating at significant scale and solve challenging distributed systems and reliability problems that directly affect our customers.

Oracle is committed to creating an inclusive workplace where everyone has the opportunity to contribute, grow, and succeed.

Career Level - IC4

More Info

About Company

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Job ID: 152286175

Similar Jobs

Bengaluru, India

Skills:

NginxPrometheusKafkaGrafanaRedisElasticsearchRabbitmqDockerTerraformLinuxAnsiblePostgresHelmKubernetesEvent HubVulnerability management toolsSecurity best practices

Bengaluru, India

Skills:

PrometheusAmazon CloudWatchGrafanaJenkinsTerraformAnsibleShell scriptingMongoDBKubernetesPythonAWSAWS RDS PostgreSQLGitOpsAWS RDS MySQLAlertmanagerAmazon AuroraGitHub Actions

Beware of Scammers

We don’t charge money for job offers