Search by job, company or skills

Site Reliability Engineer

10-12 Years
  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

At Bajaj Broking, we're building capital markets infrastructure for millions of traders. when market volatility spikes, our platform doesn't get to slow down, queue requests, or crash. zero downtime isn't a KPI here—it's the baseline requirement.

we don't view reliability as an operational patch. we treat it as an engineering problem.

we're looking for an SRE who thinks in distributed systems, automates relentlessly, and obsesses over microsecond latencies, system resilience, and high-availability architecture.

What you'll own

  • architect for high concurrency: design, scale, and maintain cloud-native systems capable of handling massive order volumes and real-time market data feeds during peak trading hours.
  • kill operational toil: if a manual task happens twice, write software to automate it. build self-healing infrastructure, deployment pipelines, and auto-scaling frameworks.
  • drive observability: move past basic uptime dashboards. design deep distributed tracing, custom telemetry, and predictive alerting to catch anomalies before users do.
  • own incident engineering: lead post-mortems with a strict blameless, root-cause culture. turn every outage or degradation into a permanent architectural fix.
  • partner with dev squads: work closely with backend teams to establish SLOs, SLAs, and error budgets—ensuring system velocity never compromises platform stability.

Signals we look for

  • systems fundamentals: 10+ years of experience engineering high-throughput, low-latency production environments (preferably in FinTech, high-frequency trading, or hyper-growth consumer platforms).
  • cloud & orchestration: deep expertise in AWS/GCP, Kubernetes, Docker, Terraform, and infrastructure-as-code paradigms.
  • code fluency: strong programming foundation in Go, Python, or Java. you build tools and services, not just write bash scripts.
  • telemetry mastery: hands-on expertise with Prometheus, Grafana, ELK stack, Jaeger, or OpenTelemetry at scale.
  • database resilience: strong grasp of database tuning, high-availability setups, and failover strategies for PostgreSQL, Redis, or Kafka.

EQUAL OPPORTUNITY EMPLOYER

We are an equal opportunity employer committed to fostering an inclusive, diverse, and equitable workplace. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability status, or any other characteristic protected by law.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 153515709

Similar Jobs

Pune, India

Skills:

ElkPrometheusGrafanaLinux AdministrationJenkinsGitCloudwatchBitbucketTerraformAWS CloudFormationKubernetesAWSCI CD

Pune, India

Skills:

.NETBackupNlbCDKVpcAuto ScalingLambdaEc2Route 53TlsPythonAWSCloudformationPowerShellBashSecurity GroupsCloudwatchECSIamMulti-AZ failoverPITRACMRDS SQL ServerFargateALBX-Ray

Pune, India

Skills:

JavaDatabasesCloudLinuxNetworkingBashKubernetesPythonGo

Pune, India

Skills:

containerization KibanaPerforcePrometheusKafkaTableauGrafanaNosqlDockerInfrastructure ManagementSystem AdministrationMySQLAWSAutomationZabbixDevopsJenkinsGitGcpAnsibleElastic SearchPuppetAzureKubernetesVirtualizationChefFilebeatMonitoring

Pune

Skills:

KubernetesCI/CD / Git / Helm tools/ JenkinsPrometheus / Grafana / Datadog/ELK- DatadogAPPD)SLA / SLO / MTTRPython / BashIncident ResponseAzure CloudSplunkRoot Cause AnalysisReliability EngineeringOpen TelemeteryGolden SignalsSRE Practicesargo cdOpen Search

Beware of Scammers

We don’t charge money for job offers