Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital & Platform Services team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them.
Job responsibilities
- Define the Production Management SRE vision, north-star outcomes, and multi‑year roadmap aligned to CIB and JPMorganChase Global Technology priorities, and establish the global SRE operating model (ways of working, intake, prioritization, and engagement with engineering/production support).
- Build and scale a high-impact SRE capability by developing a small core team and/or a global community of practice, and by partnering with business-aligned Production Support leads to embed reliability practices and “engineer out” operational load.
- Set and implement firmwide reliability standards and patterns across service cataloging, SLO/SLI and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, and capacity/performance/scalability engineering.
- Drive evidence-based service health and reliability governance through regular service reviews and reporting (availability, latency, incident trends, MTTR/MTTD, change failure rate, and customer impact), and continuously improve observability and alerting quality (logs/metrics/traces, golden signals, end-user journey monitoring, actionable routing, and reduced false positives).
- Champion enterprise-authorized AI adoption to reduce operational toil and improve incident response (triage, troubleshooting, post-incident analysis), and lead reuse-first AI-assisted reliability workflows across the SDLC/toolchain (CI/CD quality checks, automation, operational readiness) with traceability/auditability and required resiliency/security controls.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience ( NAMR/APAC – India/ LATAM/ Hong Kong)
Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
Strong engineering background: ability to design, build, and deliver automation and reliability solutions.
Fluency & expertise in Python
Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction.
Proficiency and experience with telemetry (logs/metrics/traces) collection using tools and standards such as Prometheus, OpenTelemetry, Datadog, Dynatrace, Splunk.
Experience delivering automation at scale (scripting, workflow automation, runbook automation, CI/CD-integrated guardrails).
Proven leadership skills: influencing without authority, coaching leaders, and building communities of practice.
Strong judgment around risk, security, and controls—especially when applying AI to production workflows.
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
Preferred qualifications, capabilities, and skills
- Proficient with container and container orchestration
- Experience with troubleshooting common networking technologies and issues
- Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, and actively self-educates to evaluate and recommend suitable new technologies
- Familiarity with Athena / prior experience in Athena