Staff Software Engineer (SRE) - Business Intelligence
Staff Software Engineer (SRE) - Business Intelligence
Intuit- Posted an hour ago
- Be among the first 10 applicants
Job Description
Overview
Business Intelligence today spans a portfolio of Tier 1 and Tier 2 assets whose reliability, resiliency, and cost posture are owned ad hoc across individual app teams. We're hiring a dedicated SRE to be the single accountable owner for the health, availability, and operational excellence of these assets end-to-end — closing the gap between best effort reliability and a proven, tested, continuously improving one. This role sits horizontally across all BI scrum teams, partners tightly with our DBA, and is expected to build the automation and self-serve tooling that reduces the manual toil currently driving our incidents.
Responsibilities
Reliability Ownership
Required
Business Intelligence today spans a portfolio of Tier 1 and Tier 2 assets whose reliability, resiliency, and cost posture are owned ad hoc across individual app teams. We're hiring a dedicated SRE to be the single accountable owner for the health, availability, and operational excellence of these assets end-to-end — closing the gap between best effort reliability and a proven, tested, continuously improving one. This role sits horizontally across all BI scrum teams, partners tightly with our DBA, and is expected to build the automation and self-serve tooling that reduces the manual toil currently driving our incidents.
Responsibilities
Reliability Ownership
- Serve as the single accountable owner for the health, availability, and completeness of all Tier 1/Tier 2 BI assets, replacing today's diffused, ad hoc ownership across app teams.
- Keep asset health and resiliency practices current as CI/CD evolves — progressive rollouts, canary deployments, and failure mode and effects analysis (FMEA).
- Partner closely with the DBA to ensure system health is covered end to end (app layer + data layer), not siloed.
- Design and run mock incidents / game days across BI assets to surface gaps proactively, before they become customer-facing.
- Own disaster recovery (DR) for every asset in the BI org: plan it, test it on a recurring cadence, and produce evidence it actually works (not theoretical DR).
- Build automation that removes manual, error-prone operational steps — the primary source of today's incidents and customer impact.
- Own cloud resiliency, performance engineering, and security posture as a standing discipline across BI, not a side-of-desk responsibility for app teams.
- Establish and enforce operational and cost-optimization patterns across all assets (e.g., Splunk log volume/retention hygiene, right-sizing compute).
- Build reusable, self-serve libraries and tooling so individual teams stop re-solving the same reliability/ops problems independently.
- Track Intuit-wide platform and tooling initiatives (e.g., IKS, AIR) and drive horizontal adoption across BI so each team isn't independently absorbing migration cost.
Required
- 5+ years in an SRE, DevOps, or infrastructure/platform engineering role, ideally supporting data/BI or analytics platforms.
- Demonstrated ownership of production reliability for multiple services/assets — on-call, incident response, postmortems, and follow-through on remediation.
- Hands-on experience designing and executing disaster recovery plans (not just documenting them) — including recovery testing and validation.
- Experience with progressive delivery practices: canary deployments, blue/green rollouts, feature-flag-gated releases.
- Proficiency building automation/tooling (scripting or full software engineering) to eliminate manual operational toil.
- Working knowledge of cloud infrastructure cost management — right-sizing, log/data volume hygiene, resource tagging and chargeback.
- Experience with observability stacks (e.g., Splunk or equivalent) for both incident diagnosis and cost/volume management.
- Strong cross-team collaboration skills — this role influences without direct authority across multiple scrum teams and a DBA function.
- Experience running game days / chaos engineering exercises.
- Familiarity with FMEA or similar systematic risk-assessment methodologies.
- Exposure to Intuit-specific platforms (IKS, AIR) or comparable enterprise Kubernetes/AI infra migrations.
- Security posture / cloud security engineering background.
- Experience building self-serve internal developer platforms or shared libraries consumed by multiple teams.
More Info
Job Type:
Industry:
Employment Type:
Key Skills
self-serve internal developer platforms
shared libraries
blue green rollouts
right-sizing
cloud security engineering
infrastructure platform engineering
progressive delivery practices
feature-flag-gated releases
resource tagging
log data volume hygiene
cloud infrastructure cost management
Air
game days
enterprise Kubernetes
observability stacks
IKS
automation tooling
canary deployments
chaos engineering
