Sr. Reliability Engineer
Quick Summary
Produce the tooling that powers the metrics and narrative that go into bi-weekly VP-level OE reviews and the monthly cross-engineering OER, and drive the resulting action items to closure.
Experience building AI agents or LLM-backed automation for operations: incident triage, RCA generation, alert correlation, or observability data analysis,
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve.
Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.
About the Role
~1 min readWe're looking for a Senior Reliability Engineer to make Hims & Hers systems measurably more reliable: building the observability, tooling, and automation that catch problems before people do. Recent work at this level includes preparing our stack for 32x baseline traffic during high-stakes seasonal surges with zero customer-facing issues and improved P95 latency, turning databases into a paved platform with consistent observability and guardrails, and building AI agents that pull FireHydrant, Datadog, Jira, and Confluence into a single incident picture. We use AI first: every engineer gets a Claude Enterprise license, and we expect you to use it as a core part of how you investigate, build, and write, not as an afterthought.
Responsibilities
~1 min read- →
5+ years as a Software, SRE, Platform, or Infrastructure Engineer, with a track record of owning reliability outcomes for production systems that customers depend on.
- →
Strong software engineering fundamentals. You solve reliability problems by writing code and building tooling, and you're comfortable reading application code across the stack to find the real cause.
- →
Hands-on depth in observability and SLO engineering: golden signals, burn-rate alerting, journey-level monitors, and turning noisy alert streams into actionable pages (Datadog preferred; Prometheus/Grafana and OpenTelemetry welcome).
- →
Production experience with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora).
- →
Experience running or maturing incident management end to end: on-call design, escalation policies, incident command, blameless post-mortems, and action-item follow-through in a tool like FireHydrant or PagerDuty.
- →
Daily, practical use of AI coding and analysis tools (Claude, Cursor, or similar) to accelerate investigation, code, documentation, and reporting, and clear judgment about when to trust the output and when to verify it.
- →
Communication skills to explain risk, tradeoffs, and post-incident learnings to engineers and leadership alike, in writing and in review meetings.
Requirements
~1 min readExperience building AI agents or LLM-backed automation for operations: incident triage, RCA generation, alert correlation, or observability data analysis, ideally with MCP or tool-calling integrations against Datadog, FireHydrant, or Jira.
Load-testing and performance-engineering experience at meaningful scale (k6 or similar), including validating systems to well above expected peak.
Familiarity with service mesh (Istio) and its observability and traffic-management features.
Background in a regulated or healthcare environment, where reliability and data handling carry patient-safety and compliance weight.
Experience designing vendor and partner escalation frameworks with defined severities and response SLAs.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 7, 2026
- First seen
- October 7, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 65%
- Scored at
- October 7, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.