Search Jobs

Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

eton solutions lp
  • Posted 15 hours ago
  • Be among the first 10 applicants

Job Description

About the Company:

Eton Solutions is a hypergrowth fintech transforming the Family Office segment of the Wealth Management industry. Eton's AtlasFive® is a comprehensive enterprise management platform specifically designed to allow today's modern Family Office meet the unique and varied challenges of Ultra High Net Worth families.

For more details visit: https://eton-solutions.com/

Senior SRE / Observability Engineer

Role Summary

We are seeking a Senior Site Reliability Engineer (SRE) with deep expertise in observability, monitoring framework design, and Azure platform operations. This is a hands-on technical leadership role responsible for owning and driving the end-to-end monitoring framework for all Atlas5 production environments. The ideal candidate has a strong background in monitoring tool evaluation and implementation, alert engineering, incident response, and SRE best practices, and can operate independently to design, deploy, and continuously improve the observability stack across a complex, multi-layer cloud architecture.

Key Responsibilities

Monitoring Framework Ownership

  • Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers:
  • Layer 1: Infrastructure (VMs, compute, networking)
  • Layer 2: Data Tier (SQL DB, Cosmos DB, Elastic Pools, Microsoft Fabric)
  • Layer 3: Application (App Services, Function Apps, BRE)
  • Layer 4: Integration (EDA, Logic Apps, APIM)
  • Layer 5: File Exchange (FTP/SFTP, file drop, inbound feeds)
  • Define and configure alert thresholds, severity levels, and escalation paths for all layers
  • Ensure every alert triggers a Jira incident ticket and DevOps channel notification within defined SLA
  • Enforce mandatory RCA completion for all P1-Critical incidents before ticket closure

Tool Selection & Implementation

  • Evaluate, recommend, and implement the right monitoring toolset for each layer — Azure Monitor, Prometheus, Grafana, New Relic, SigNoz, Zabbix, Dynatrace, or equivalent
  • Build a cost-effective hybrid monitoring stack that balances coverage, operational overhead, and cost vs. native Azure Monitor
  • Deploy and configure monitoring agents across all Atlas5 production infrastructure
  • Integrate monitoring tools with Jira Service Management for automated incident creation
  • Build and maintain unified monitoring dashboards for real-time 5-layer status visibility

SRE Practice & Incident Management

  • Define and maintain SLOs, SLIs, and error budgets for all Atlas5 production services
  • Lead incident response for P1-Critical production incidents — own the investigation, RCA, and permanent fix process
  • Drive the monitoring gap assessment as a mandatory step in every incident RCA
  • Configure auto-escalation for unacknowledged Critical alerts to DevOps Lead and secondary on-call
  • Conduct regular alert quality reviews — target false positive rate below 5%

Azure Observability Engineering

  • Configure and manage Azure Monitor, Log Analytics workspaces, Application Insights, and Diagnostic Settings
  • Design and manage KQL queries for alerting, dashboards, and operational reporting
  • Optimise monitoring costs — migrate Log Search Alerts to Metric Alerts (30x cheaper) and Activity Log Alerts (free) where applicable
  • Configure Prometheus scraping for Azure-native services and manage the Prometheus + Grafana stack
  • Instrument applications with Open Telemetry for vendor-neutral, future-proof telemetry

Reliability & Performance Engineering

  • Proactively identify reliability risks — throttling, capacity saturation, latency spikes, and dependency failures
  • Partner with DevOps and application teams to instrument new services and define monitoring parameters
  • Work with development team to define and implement EDA logging and integration layer monitoring parameters
  • Contribute to deployment reliability — flag monitoring gaps in pre-deployment validation and change management

Required Experience & Skills

Core Requirements

  • 8 – 10 years of experience in SRE, observability engineering, or cloud operations roles
  • Proven experience designing and implementing monitoring frameworks from scratch in Azure environments
  • Strong Azure Monitor expertise — Log Analytics, Application Insights, Metric Alerts, Log Search Alerts, Diagnostic Settings
  • Hands-on experience with at least 2 of: Prometheus, Grafana, New Relic, Datadog, Dynatrace, Zabbix, SigNoz
  • Experience integrating monitoring with incident management tools (Jira, PagerDuty, or equivalent)

Technical Skills

  • Azure: Azure Monitor, Log Analytics (KQL), Application Insights, Azure Alerts, Diagnostic Settings, Service Health
  • Monitoring tools: Prometheus, Grafana, New Relic, Dynatrace, Zabbix, SigNoz — hands-on with at least 2
  • Instrumentation: Open Telemetry (OTel) — logs, metrics, traces; Application Insights SDK
  • Alerting: Alert rule design, severity taxonomy, escalation policies, on-call automation
  • Scripting: KQL (mandatory), PowerShell or Python for alert automation and tooling
  • Azure infrastructure: App Services, Function Apps, Cosmos DB, SQL DB, Microsoft Fabric, AKS — operational monitoring knowledge
  • Cost optimisation: experience reducing Azure Monitor costs through Metric Alert migration and Log tier management

Nice to Have

  • Experience in fintech or regulated financial services environments
  • Familiarity with EDA (Event-Driven Architecture) and integration layer monitoring (Service Bus, Logic Apps, APIM)
  • Experience with SigNoz or Open Telemetry-native observability stacks
  • Dynatrace certification or hands-on Dynatrace deployment experience

if interested, share your updated resume to [Confidential Information]

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

Azure Alerts

Open Telemetry

KQL

Diagnostic Settings

Log Analytics

SigNoz

Azure Monitor

Application Insights

About Company

Similar Jobs

8-10 yrs
Bengaluru, India
Skills:
HadoopPrometheusGrafanaDatadogApache AirflowCloudwatchTerraformLinuxSparkSplunkPythonKubernetesAWSAWS EMRFinOpsAmazon EKSOpenSearch
10-12 yrs
Bengaluru, India
Skills:
JavaUnixElkNodejsGrafanaAutosysShellCloudwatchLinuxRestful ServicesPythonLoggingAWSAlertingEFKPagerDutySLOsOracle FCMREMonitoring
7-9 yrs
Bengaluru, India
Skills:
GolangRustLinux System AdministrationDnsFirewallsroutingJenkinsDHCPIp TablesKubernetesPythonScriptingSLURMcore linux networkingCI CD toolsArgoCDlarge-scale cluster management
8-10 yrs
Bengaluru, India
Skills:
Change ManagementTerraformAnsibleIncident ManagementProblem ManagementPythonAzureAWSAlertingTroubleshootingObservabilityMonitoringSRE conceptsITIL principles
6-10 yrs
Bengaluru, India
Skills:
Service Mesh (IstioKong Mesh)Large Language Models (LLMs)BashFluxTerraformHelmKubernetesPythonGKEGitOpsAI-assisted workflowsGoCIS complianceSecurity vulnerability managementGitLab CIGitHub ActionsAPI Gateway KongCI/CDArgoCDAIOps