Site Reliability Engineering Lead
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the United States.
This leadership role is responsible for building reliable, scalable, and secure cloud platforms while developing a high-performing SRE team.
You will lead engineers, establish priorities, and drive initiatives that improve service reliability, resilience, automation, and operational efficiency.
The role combines people leadership with hands-on technical direction across modern cloud and infrastructure environments.
You will work closely with Development, Security, Product, and other engineering teams to strengthen platform performance and incident response.
A key focus will be reducing operational toil through automation, observability, infrastructure as code, and self-healing capabilities.
You will also guide post-incident reviews, root-cause analysis, and continuous improvement efforts across services and infrastructure.
This is an opportunity to shape SRE practices at scale while supporting engineers in their technical and professional growth.
-
Lead, mentor, and develop a small to medium-sized team of Site Reliability Engineers through regular 1:1s, performance reviews, career planning, and ongoing coaching.
-
Own hiring, onboarding, team capacity, resourcing, and workforce planning decisions to ensure the team can effectively support business and platform priorities.
-
Establish team objectives, prioritize the engineering backlog, coordinate planning, and ensure projects and operational tasks remain aligned with reliability goals.
-
Lead reliability initiatives across infrastructure and services, improving availability, scalability, resilience, security, and operational performance.
-
Drive incident response activities and facilitate blameless post-incident reviews, ensuring timely root-cause analyses and actionable follow-up.
-
Partner with Development, Security, Product, and other engineering teams to resolve cross-functional issues and strengthen collaboration.
-
Champion automation and operational excellence by reducing manual work, eliminating recurring toil, and introducing self-healing systems and infrastructure automation.
-
Support the design and evolution of scalable, secure, cloud-native environments and continuously identify opportunities to improve performance, reliability, and cost efficiency.
Requirements
~1 min read-
Demonstrated experience in SRE, DevOps, infrastructure engineering, or a related discipline, including experience leading engineering teams.
-
Expert-level knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and large-scale troubleshooting.
-
Advanced experience with Terraform, including modular infrastructure-as-code design, state management, multi-environment provisioning, and policy-as-code.
-
Deep knowledge of Azure Cloud services, including compute, networking, identity and access management, storage, and cost optimization.
-
Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and automated rollback approaches.
-
Strong knowledge of observability platforms such as Prometheus, Grafana, and OpenTelemetry, along with experience managing SLOs, SLAs, and error budgets.
-
Strong automation capabilities and advanced proficiency in Python, Bash, and/or PowerShell for infrastructure tooling and operational automation.
-
Deep understanding of networking fundamentals, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking.
-
Proven experience leading incident response, conducting root-cause analysis, and implementing measurable reliability improvements.
-
Strong people leadership, communication, prioritization, and problem-solving skills, with the ability to support engineers while coordinating effectively across technical teams.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- First seen
- October 2, 2026
- Last seen
- October 2, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- October 2, 2026
Signal breakdown
Similar Site Reliability Engineering jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.