Quick Summary
As a Site Reliability Engineer, you will help ensure the availability, performance, and reliability of a large federal enterprise cloud platform that operates around the clock.
Monitoring platform health and supporting service level objectives (SLOs), service level indicators, and error budgets Building and maintaining observability tooling, including metrics, logging,
Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people’s needs and transform their experience of government.
Our collaborations have shaped some of the defining moments in public-sector service delivery. We’ve helped build products that connect Veterans to tailored services, help millions access affordable health care, and support important programs like Head Start. As we work with agencies to deliver critical services, we’re also changing how the government approaches technology.
Our culture, communications, and tools are built for remote work, enabling us to bring together top talent nationwide. At Ad Hoc, remote life empowers our teams to design work environments that fit their lives and that foster flexibility and collaboration to achieve positive outcomes for our customers.
Ad Hoc values acceptance, accountability, and humility. We aren’t heroes. We learn from our mistakes and improve the process for the next time. We build small, inclusive teams to collaborate closely with our partners to solve the right problems and deliver software that works.
The Veterans Affairs business unit helps transform the VA into a modern digital services organization where Veteran outcomes are at the center of every effort. We partner with the VA to design and deliver seamless user experiences for Veterans, their families and caregivers, and VA employees. By applying better practices in service design, product management, and technology, we enable the VA to increase the use, quality, and reliability of services and decrease the time Veterans spend waiting for outcomes.
Responsibilities
~1 min readAs a Site Reliability Engineer, you will help ensure the availability, performance, and reliability of a large federal enterprise cloud platform that operates around the clock. With the support and guidance of senior engineers, you will help meet scope, schedule, and delivery requirements while improving the platform's reliability practices. Primary expectations of a Site Reliability Engineer include:
- →Monitoring platform health and supporting service level objectives (SLOs), service level indicators, and error budgets
- →Building and maintaining observability tooling, including metrics, logging, alerting, and dashboards
- →Participating in on-call rotations and incident response, helping restore service and reduce time to recovery
- →Contributing to blameless postmortems and driving follow-up actions
- →Automating repetitive operational tasks to reduce toil
- →Supporting capacity planning and performance tuning across cloud infrastructure (AWS) and Kubernetes (Amazon EKS)
- →Implementing reliability improvements as infrastructure as code (Terraform)
- →Working with government partners and application teams to meet security, SLA, and performance requirements
- →Supporting recruiting efforts by evaluating exercises and assisting with interviews
Requirements
~1 min read- Bachelor's and 5+ years of experience; relevant experience may be substituted for education
- Experience with monitoring and observability tooling and on-call operations
- Proficient with at least one infrastructure-as-code tool (Terraform preferred)
- Background in key DevOps concepts: containerization, networking, and cloud infrastructure
- Must be able to obtain and maintain a U.S. Public Trust / suitability determination
- Prior experience with the Department of Veterans Affairs
- Experience with Kubernetes (Amazon EKS) and AWS in production
- Familiarity with SLO-based reliability practices and error budgets
- Relevant certifications (e.g., AWS, Certified Kubernetes Administrator)
What We Offer
~1 min readLocation & Eligibility
Listing Details
- First seen
- July 28, 2026
- Last seen
- July 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 51%
- Scored at
- July 28, 2026
Signal breakdown
Please let adhoc know you found this job on Jobera.
4 other jobs at adhoc
View all →Explore open roles at adhoc.
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.