Search by job, company or skills

Tech Lead

Fresher

This job is no longer accepting applications

Job Description

Site Reliability Engineer (SRE)

The Role

As a Site Reliability Engineer (SRE), you will be responsible for ensuring the reliability,

availability, performance, and operability of production systems across our AI platforms and

hosted applications.

You will apply software engineering principles to operational challenges, with a strong focus on

automation, observability, and proactive reliability management. Reliability is treated as a feature,

measured objectively, and improved continuously through engineering best practices rather than

reactive support.

This role works closely with product engineering and platform teams to enable teams to ship

changes safely, operate services at scale, and learn from incidents through a blameless culture.

The role includes participation in a global oncall model supporting production systems across

multiple time zones.

Key Responsibilities

Own and improve the reliability, availability, and performance of production services.

Operate and evolve production cloud infrastructure at scale (primarily GCP).

Participate in incident management, including detection, triage, mitigation, escalation,

and recovery.

Use and improve incident workflows and tooling (e.g. ServiceNow) to ensure clear

ownership and timely communication.

Design, implement, and operate observability solutions including monitoring, logging,

tracing, synthetics, and dashboards (e.g. Splunk Observability, OpenTelemetry).

Define, monitor, and report SLIs, SLOs, and error budgets for critical services.

Reduce operational toil through automation and engineeringled solutions.

Contribute to production readiness reviews, capacity planning, and deployment safety

mechanisms.

Perform rootcause analysis and lead or contribute to blameless postincident reviews.

Communicate complex technical issues clearly to engineers, stakeholders, and

leadership.

Requirements

Must haves

Strong experience with cloudnative concepts and technologies, preferably

on GCP.

Handson experience operating production cloud infrastructure at scale.

Proven experience with incident management, ideally using platforms such

as ServiceNow.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151463599

Similar Jobs

Chennai, India

Skills:

BitbucketSqlAgileSpring BootSplunkGithubPostgresJ2EESql TuningDatadogAWSOracleAzureDB2DockerGcpJenkinsRestful ApisClaudeCodeDevSecOps practicesMessaging queue or streaming toolsMicroservices Design PatternsOpenCode

Bengaluru, India

Skills:

LanO365DnsMac OsWindowsWanJIRADHCPNetworking ConceptsServicenowVpnMs OfficeWi-FiActive DirectoryITSM toolscommon enterprise applicationsRemedyAWS Connect

India

Skills:

TensorflowAzure MLPytorchFastAPIPythonLangChainMLflowPineconeAutoGenVector DatabasesGCP Vertex AIAWS SageMakerLLM Fine-tuningGenAI OrchestrationWeaviateLlamaIndex

Beware of Scammers

We don’t charge money for job offers