~44m ago
New

Site Reliability Engineering Lead

United StatesUnited StatesRemoteFull-timelead
OtherSite Reliability Engineering
2 views0 saves0 applied

Quick Summary

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the United States.

Technical Tools
OtherSite Reliability Engineering

This leadership role is responsible for building reliable, scalable, and secure cloud platforms while developing a high-performing SRE team.
You will lead engineers, establish priorities, and drive initiatives that improve service reliability, resilience, automation, and operational efficiency.
The role combines people leadership with hands-on technical direction across modern cloud and infrastructure environments.
You will work closely with Development, Security, Product, and other engineering teams to strengthen platform performance and incident response.
A key focus will be reducing operational toil through automation, observability, infrastructure as code, and self-healing capabilities.
You will also guide post-incident reviews, root-cause analysis, and continuous improvement efforts across services and infrastructure.
This is an opportunity to shape SRE practices at scale while supporting engineers in their technical and professional growth.

  • Lead, mentor, and develop a small to medium-sized team of Site Reliability Engineers through regular 1:1s, performance reviews, career planning, and ongoing coaching.

  • Own hiring, onboarding, team capacity, resourcing, and workforce planning decisions to ensure the team can effectively support business and platform priorities.

  • Establish team objectives, prioritize the engineering backlog, coordinate planning, and ensure projects and operational tasks remain aligned with reliability goals.

  • Lead reliability initiatives across infrastructure and services, improving availability, scalability, resilience, security, and operational performance.

  • Drive incident response activities and facilitate blameless post-incident reviews, ensuring timely root-cause analyses and actionable follow-up.

  • Partner with Development, Security, Product, and other engineering teams to resolve cross-functional issues and strengthen collaboration.

  • Champion automation and operational excellence by reducing manual work, eliminating recurring toil, and introducing self-healing systems and infrastructure automation.

  • Support the design and evolution of scalable, secure, cloud-native environments and continuously identify opportunities to improve performance, reliability, and cost efficiency.

Requirements

~1 min read
  • Demonstrated experience in SRE, DevOps, infrastructure engineering, or a related discipline, including experience leading engineering teams.

  • Expert-level knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and large-scale troubleshooting.

  • Advanced experience with Terraform, including modular infrastructure-as-code design, state management, multi-environment provisioning, and policy-as-code.

  • Deep knowledge of Azure Cloud services, including compute, networking, identity and access management, storage, and cost optimization.

  • Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and automated rollback approaches.

  • Strong knowledge of observability platforms such as Prometheus, Grafana, and OpenTelemetry, along with experience managing SLOs, SLAs, and error budgets.

  • Strong automation capabilities and advanced proficiency in Python, Bash, and/or PowerShell for infrastructure tooling and operational automation.

  • Deep understanding of networking fundamentals, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking.

  • Proven experience leading incident response, conducting root-cause analysis, and implementing measurable reliability improvements.

  • Strong people leadership, communication, prioritization, and problem-solving skills, with the ability to support engineers while coordinating effectively across technical teams.

What We Offer

~2 min read
✓U.S. national base salary range of $118,300–$219,800, with geographic differentials potentially applying depending on location.
✓Eligibility for an annual incentive bonus.
✓Country- and location-specific employee benefits designed to support overall health and well-being.
✓Support for an accessible and inclusive hiring process, including reasonable accommodations where required.
✓Remote/home-based opportunities available in multiple U.S. locations, including Florida, Connecticut, New Jersey, New York, and Pennsylvania.
✓Opportunity to lead a technically sophisticated SRE function focused on cloud platforms, automation, resilience, and operational excellence.
✓Professional development opportunities through team leadership, cross-functional collaboration, and exposure to large-scale cloud environments.

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

First seen
October 2, 2026
Last seen
October 2, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
October 2, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Site Reliability Engineering Lead