specter
specter2mo ago
New

Site Reliability Engineer

United StatesUnited States·San Franciscofull-timemid
EngineeringDevops Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Reactive — Triage & Recovery Debug and triage issues across a live fleet of diverse Linux-based sensor nodes and edge appliances deployed at customer sites. SSH into field hardware to diagnose, patch,

Requirements Summary

Strong Linux systems administration — comfortable working over SSH in production, not just dev environments. Experience

Technical Tools
EngineeringDevops Engineer

Responsibilities

~1 min read
  • →

    Debug and triage issues across a live fleet of diverse Linux-based sensor nodes and edge appliances deployed at customer sites.

  • →

    SSH into field hardware to diagnose, patch, and recover systems — often with limited remote access and incomplete information.

  • →

    Own site bring-ups end to end; be the person who gets things back online.

  • Build and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.

  • Identify repeat fires and eliminate them — build tooling, pre-deployment checks, and root cause processes that prevent recurrence.

  • Automate toil relentlessly: if you're doing something twice, you should be scripting it.

  • Collaborate with embedded systems, and platform teams to define reliability and deployment requirements.

  • Design and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).

  • Surface and close telemetry gaps; build fleet-wide visibility that enables data-driven reliability decisions.

  • Develop runbooks, incident response procedures, and participate in on-call rotations.

Requirements

~1 min read
  • Strong Linux systems administration — comfortable working over SSH in production, not just dev environments.

  • Experience with edge or on-prem hardware alongside cloud infrastructure.

  • Solid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.

  • Scripting or programming in Python, Go, or Bash for operational tooling.

  • Familiarity with containerization (Docker, Kubernetes a plus).

  • Embedded systems experience — reading firmware logs, understanding hardware-software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.

  • Deeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.

  • Rust or C experience — we have firmware in both; being able to read and reason about low-level code accelerates triage significantly.

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
July 2, 2026
First seen
September 25, 2026
Last seen
September 25, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
14%
Scored at
September 25, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

specterSite Reliability Engineer