Site Reliability Engineer

United StatesUnited States·San Diegomid
EngineeringDevops Engineer
0 views0 saves0 applied

Quick Summary

Overview

Vannevar builds AI systems for the Department of War's most consequential missions.

Technical Tools
EngineeringDevops Engineer

Vannevar builds AI systems for the Department of War's most consequential missions. We have 125 deployments across every branch and combatant command spanning our core platform and five distinct products. An an example of how we work, when major combat operations started with Iran, we fielded a new product supporting 24/7 operations and 9,000+ users in three months.

How we build is our advantage: we deploy forward with the people who own the mission. Our engineers, product team, and CTO deploy forward to the point of friction, including visiting units in Ukraine. We focus on embedding directly with operational units supporting great power competition with China, combat operations with Iran, and counter-narcotics missions.

About the Role

~1 min read

We are looking for an Site Reliability Engineer to own the reliability, health, and deployment automation of the platform at Vannevar Labs. In this role you'll be the person watching the system's pulse — monitoring dashboards, catching health issues before they become incidents, and owning the debugging process from first alert to resolution. Your decisions today will have a large impact on the company's future.
We believe that simple systems are easier to understand, maintain, and scale. You will be making trade-offs as you work to ensure that our systems are prepared to operate reliably in high-side environments at scale. A strong sense of judgment matters here: knowing when to dig deeper into a problem yourself and when to pull in the right people to escalate. Clear, calm communication — during an incident and in day-to-day work — is a must.

Responsibilities

~1 min read
  • →Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks — often before anyone else notices them.
  • →Own the debugging and incident response process end to end, exercising good judgment about when to investigate more deeply and when to escalate.
  • →Build logging, monitoring, and observability tooling to visualize the state of the platform and continuously mature our SRE practices.
  • →Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.
  • →Understand and help improve the deployment process, and automate build & deployment pipelines.
  • →Identify bottlenecks in engineering workflows and drive improvements that make the whole team faster and more reliable.
  • →Develop self-service tools and automation to improve engineering efficiency.
  • →Play a critical part in implementing a secure, robust, high-availability delivery pipeline.
  • →Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.

Requirements

~2 min read
  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents — comfortable owning a debugging process and making the call on when to dig in versus escalate.
  • Excellent communication skills, especially the ability to stay clear and organized while troubleshooting live issues.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience participating in an on-call rotation and running or contributing to post-mortems.
  • Knowledge of AWS cloud technologies.
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience working in an agile scrum environment, with the ability to work independently.
  • Able to quickly learn new and existing technologies.
  • Strong attention to detail and analytical capabilities.
  • Willingness and ability to work on-site in San Diego, CA.
  • U.S. Citizenship status is required, as this position requires the ability to access U.S.-only data systems and export-controlled data.
  • TS/SCI Clearance required.
  • Experience defining and tracking SLOs/SLIs and error budgets.
  • Experience crafting CI/CD processes and automation.
  • Proficient with containerization technologies like Docker.
  • Experience working in AWS GovCloud.
  • Experience with modern web services architectures.
  • Experience with relational database systems, including SQL and relational design.
  • Experience working with Elasticsearch/OpenSearch.
  • Strong collaboration and negotiation skills, with the ability to work on cross-functional projects with internal partner engineering teams.

What We Offer

~2 min read
We’re proud to offer competitive benefits that support our employees. Some key highlights of our benefits package include:
✓Health, dental, and vision insurance
✓100% remote first culture. You can work from anywhere in the US and all full time employees have WeWork access
✓Unlimited PTO including competitive vacation and holiday schedules
✓Lifestyle stipends - Monthly mental health, wellness & fitness stipend, in-home office setup stipend and family planning assistance
✓Salary top-up during military reserve duty
✓Fully paid parental leave
✓Child and pet care reimbursement during travel

Location & Eligibility

Where is the job
San Diego, United States
On-site at the office
Who can apply
Open to applicants worldwide

Listing Details

Posted
September 16, 2026
First seen
September 20, 2026
Last seen
October 4, 2026

Posting Health

Days active
13
Repost count
0
Trust Level
28%
Scored at
October 4, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Site Reliability Engineer