Search by job, company or skills

Site Reliability Engineer

8-10 Years
Early Applicant
  • Posted 3 days ago
  • Be among the first 10 applicants

Job Description

Experience: 8+ years

.NET Application Reliability

· Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.

· Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.

· Partner with application engineers to embed reliability into new feature design and deployment practices.

· Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.

AWS Compute & Infrastructure

· Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit.

· Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.

· Automate operational tasks, deployment pipelines, and disaster recovery procedures.

· Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work.

RDS SQL Server Operations

· Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.

· Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly.

· Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.

· Capacity plan and scale database infrastructure to support transaction volume growth.

Observability & Monitoring

· Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.

· Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.

· Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.

· Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.

Networking & Security

· Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs.

· Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM.

· Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning.

· Support compliance and security controls relevant to a PCI-regulated payments environment.

Incident Response & On-Call

· Participate in on-call rotation to respond to production incidents and drive swift resolution.

· Define and track error budgets; use them to balance velocity and reliability investment.

· Communicate status updates clearly during incidents and coordinate cross-functional response.

· Maintain and improve runbooks, escalation paths, and on-call health over time.

Cross-Functional Partnership

· Collaborate with platform engineering teams on architecture decisions and scalability requirements.

· Share observability and reliability best practices with application teams.

· Mentor engineers on SRE principles and operational excellence.

· 8+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role.

· C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.

· AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.

· RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.

· Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.

· AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.

· SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.

· Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.

· Excellent communication skills and a collaborative, blameless engineering mindset.

· Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152083901

Similar Jobs

Pune, India

Skills:

PrometheusGrafanaJenkinsTerraformLinuxShell scriptingHelmPythonKubernetesAWSLokiGoGitHub ActionsTempoOpenTelemetry

Pune, India

Skills:

GcpTerraformIamNetworkingKubernetesCloudSQLGKEinfrastructure automationSpanner

Pune, India

Skills:

containerization KibanaPerforcePrometheusKafkaTableauGrafanaNosqlDockerInfrastructure ManagementSystem AdministrationMySQLAWSAutomationZabbixDevopsJenkinsGitGcpAnsibleElastic SearchPuppetAzureKubernetesVirtualizationChefFilebeatMonitoring

Pune, India

Skills:

JavaCloud FoundryMavenKafkaSpring BootRedisMicroservicesJenkinsGitJmeterGcpDockerBlazemeterDynatraceSplunkRestful ApisAzureKubernetesAWS

Pune, India

Skills:

ElkPrometheusGrafanaDatadogGcpDockerTerraformAnsibleAzureKubernetesPythonAWSGoGitLab CIGitHub ActionsArgoCD

Beware of Scammers

We don’t charge money for job offers