adhoc
adhoc~25d ago
New

Senior Site Reliability Engineer

senior
EngineeringDevops Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

As a Senior Site Reliability Engineer, you will serve as an experienced individual contributor responsible for the availability, performance,

Requirements Summary

Defining and maintaining service level objectives (SLOs), service level indicators, and error budgets, and driving the platform toward them Designing and operating observability across metrics,

Technical Tools
EngineeringDevops Engineer

Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people’s needs and transform their experience of government.

Our collaborations have shaped some of the defining moments in public-sector service delivery. We’ve helped build products that connect Veterans to tailored services, help millions access affordable health care, and support important programs like Head Start. As we work with agencies to deliver critical services, we’re also changing how the government approaches technology.

Our culture, communications, and tools are built for remote work, enabling us to bring together top talent nationwide. At Ad Hoc, remote life empowers our teams to design work environments that fit their lives and that foster flexibility and collaboration to achieve positive outcomes for our customers.

Ad Hoc values acceptance, accountability, and humility. We aren’t heroes. We learn from our mistakes and improve the process for the next time. We build small, inclusive teams to collaborate closely with our partners to solve the right problems and deliver software that works.

The Veterans Affairs business unit helps transform the VA into a modern digital services organization where Veteran outcomes are at the center of every effort. We partner with the VA to design and deliver seamless user experiences for Veterans, their families and caregivers, and VA employees. By applying better practices in service design, product management, and technology, we enable the VA to increase the use, quality, and reliability of services and decrease the time Veterans spend waiting for outcomes.

Responsibilities

~1 min read

As a Senior Site Reliability Engineer, you will serve as an experienced individual contributor responsible for the availability, performance, and reliability of a large federal enterprise cloud platform that operates around the clock. With minimal oversight, you will help meet scope, schedule, and delivery requirements while shaping the platform's reliability strategy. Primary expectations of a Senior Site Reliability Engineer include:

  • Defining and maintaining service level objectives (SLOs), service level indicators, and error budgets, and driving the platform toward them
  • Designing and operating observability across metrics, logging, tracing, and alerting
  • Leading incident response and on-call practices, including escalation, mitigation, and time-to-recovery improvements
  • Driving blameless postmortems and systemic reliability improvements
  • Engineering automation to eliminate toil and improve operational efficiency
  • Self-directed design of reliable cloud infrastructure (AWS) and Kubernetes (Amazon EKS), including tradeoffs between cost, reliability, and efficiency
  • Building reusable modules and mentoring engineers on reliability practices
  • Presenting design documents and system diagrams to stakeholders
  • Participating in technical depth interviews with new candidates

Requirements

~1 min read
  • Bachelor's and 7+ years of experience; relevant experience may be substituted for education
  • Demonstrated experience owning reliability (SLOs, observability, incident response) for production systems
  • Expert-level knowledge of at least one infrastructure-as-code tool (Terraform preferred)
  • Deep command of cloud infrastructure, containerization, and networking
  • Must be able to obtain and maintain a U.S. Public Trust / suitability determination
  • Prior experience with the Department of Veterans Affairs
  • Kubernetes (Amazon EKS) and AWS at scale
  • Familiarity with FedRAMP, NIST 800-53, and zero-trust architecture
  • Relevant certifications (e.g., AWS, CKA/CKS)

What We Offer

~1 min read
Company-subsidized health, dental, and vision insurance
Flexible PTO
401K with employer match
Paid parental leave after one year of service
Employee Assistance Program

Location & Eligibility

Where is the job
Location terms not specified

Listing Details

First seen
July 28, 2026
Last seen
July 29, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
51%
Scored at
July 28, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

adhocSenior Site Reliability Engineer