Search by job, company or skills

Senior Manager of SRE

Senior Manager of SRE

JP Morgan Chase & Co.
Fresher
Not Disclosed
  • Posted 21 hours ago
  • Be among the first 10 applicants

Job Description

Guide and shape the future of technology at a globally recognized firm, driven by pride in ownership.

As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Infrastructure Platforms team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your team's strategic planning, driving continual improvement in customer experience, resiliency, security, scalability, monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.

Job responsibilities
  • Shows expertise in site reliability principles and demonstrates what it means to strike the balance between features, efficiency, and stability, negotiating with peers and executive partners to ensure optimal outcomes
  • Drives the adoption of site reliability practices throughout the organization and ensures teams demonstrate site reliability best practices empirically through stability and reliability metrics
  • Manage a small team of enthusiastic SREs with varying experiences.
  • Accountable for their Annual performance reviews, 1:1 and career development.
  • Drives a culture of continual improvement by encouraging real-time feedback to improve the customer's experience
  • Ensures your team collaborates with other teams within your specialization and avoids duplication of work where possible
  • Follows an objective, data-driven, post-mortem strategy by conducting regular team debriefs that enable team members to learn from successes and mistakes
  • Coaches and develops entry to mid-level team members through tailored feedback and growth plans
  • Ensures your team documents and shares their knowledge and innovations via internal forums, communities of practice, guilds, and conferences
  • Establishes team standards for AI-assisted reliability workflows across automation and delivery practices, ensuring traceability/auditability, resiliency, and security controls.
  • Drives reuse-first adoption of enterprise-authorized AI capabilities within the work environment to improve reliability operations and customer experience outcomes, with human-in-the-loop validation and appropriate handling of sensitive data.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience . I
  • Demonstrates advanced proficiency in site reliability culture and principles and can demonstrate how to implement site reliability across application and platform teams while avoiding common pitfalls
  • Experience leading teams in the safe use of enterprise-authorized AI capabilities within the work environment for reliability engineering workflows, including validation habits and awareness of data sensitivity.
  • Ability to set and reinforce organization-level practices for reviewing AI-assisted recommendations and escalating uncertain decisions while maintaining resiliency, security, and auditability outcomes.
  • Experience leading technologists to manage and solve complex technological issues on an organizational level
  • Influences the team's culture by championing innovation and change for success
  • Experience hiring, developing, and recognizing talent
  • Proficient in at least one programming language such as Python, Java etc., and Web frameworks like Django Flask etc.,
  • Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, etc.)
  • Proficient with CI/CD practices and related tooling, container/container orchestration, and troubleshooting common networking technologies and issues
Preferred qualifications, capabilities, and skills
  • Ability to initiate and implement ideas to solve business problems
  • Passion for learning new technologies and driving innovative solutions.

Key Skills

site reliability engineering concepts

AI-assisted reliability workflows

container orchestration

site reliability culture and principles

enterprise-authorized AI capabilities

CI/CD practices and related tooling