Rapidsos
Rapidsos11d ago
$160,000 – $195,000/yr

Senior Site Reliability Engineer

New York Or BostonRemotesenior
EngineeringSite Reliability EngineerDevops EngineerInfrastructure & Cloud
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Own performance and reliability outcomes: Ownership of how application-level decisions create system-level impact, including connection pooling, database architecture, traffic routing patterns,

Technical Tools
EngineeringSite Reliability EngineerDevops EngineerInfrastructure & Cloud

At RapidSOS, we are committed to using technology to build a safer, stronger future and working together to save lives. We’re in an exciting phase of growth, welcoming new members from across the globe to our mission-driven, ambitious, and inclusive team. Our work is founded on our values of elevating purpose, inventing tomorrow, delivering with urgency, serving with integrity, and winning together, all of which support a company culture where people can innovate, collaborate, grow, and, above all, make an impact. 

RapidSOS is ​​the leading public safety AI company that unlocks mission-critical intelligence for first responders and security teams – enabling faster, smarter and more accurate emergency response. Real-time data from the world’s largest safety network of 700M+ devices, 200+ global enterprises, and 23,000+ federal, state and local agencies fuels the RapidSOS HARMONY AI engine that delivers this intelligence to those who need it most. Learn more at www.RapidSOS.com.

Responsibilities

~1 min read
  • Own performance and reliability outcomes: Ownership of how application-level decisions create system-level impact, including connection pooling, database architecture, traffic routing patterns, and memory allocation. Collaboration with engineering teams that own specific domains, partnering directly to improve reliability and performance across their systems.
  • Design for system resilience: Responsibility for strengthening reliability through proactive design decisions, including safer deployment patterns, failover strategies, and redundancy approaches that improve system behavior under stress.
  • Build observability into system behavior: Proactively instrument services with structured logging, metrics, and alerting so systems are easier to understand and debug. The focus is on creating clear signals from production behavior before issues escalate.
  • Own incidents from signal to resolution: Ownership of production issues from first signal through resolution, including investigation across infrastructure and application layers, root cause identification, and implementation of fixes that restore stability and strengthen system behavior long term.
  • Work across the stack without a permission slip: You’ll work across infrastructure-as-code, container orchestration, CI/CD pipelines, and service-level application code. When issues come up, you don’t wait for a handoff—ownership is taken directly and driven through to resolution.
  • 5+ years of professional engineering experience with deep expertise in Python 
  • Real cloud infrastructure experience with AWS: networking, managed databases, cost implications of traffic routing decisions, IAM, DNS-based routing and failover
  • Hands-on kubernetes experience with containerized workloads in production across EKS, ECS, or Fargate, you can read events, understand resource limits, know when to drain vs. delete a node, and understand the tradeoffs between orchestration models
  • Strong understanding of distributed systems and how they fail, including resource exhaustion, replication lag, queue backpressure, and other common failure modes
  • Experience operating high-throughput messaging systems (RabbitMQ, Kafka, AWS SNS / SQS, etc.)  and the infrastructure around them, including infrastructure-as-code (e.g., Terraform) and CI/CD pipelines, with an emphasis on improving reliability and scalability
  • Experience building or improving observability through logging, metrics, and alerting
  • Demonstrable experience in using AI to safely and securely enhance velocity, improve reliability and recoverability of services
  • Strong communication and interpersonal skills; is a team player with a positive attitude 
  • Highly self-motivated; ability to adapt and learn quickly in a fast-paced environment with a strong sense of ownership
  • Strong proficiency in coding best practices – ability to write clean, maintainable, and testable code 
  • Demonstrated expertise in problem solving – comfortable working across both infrastructure and application layers to diagnose and resolve issues at the source
  • Ability and willingness to collaborate in-person a few times per quarter, or as needed
  • Experience supporting production systems in an on-call or similar capacity where reliability matters
  • Experience with observability and GitOps tooling; hands-on with Datadog (APM, alerting), Elasticsearch/OpenSearch, and ArgoCD-based GitOps deployments; comfortable modernizing legacy CI/CD pipelines (e.g., Concourse, Jenkins) toward cloud-native approaches

What We Offer

~1 min read
The chance to work with a passionate team on solving one of the largest challenges globally
Competitive salary and benefits and equity participation
A dynamic, flexible and fun start-up work environment with a highly talented team

Starting pay for a successful applicant will depend on a variety of job-related factors, which may include experience, relevant skills, training, education, location, business needs, or market demands. The salary range for this role is $160,000 - $195,000. This role will also be eligible to receive equity options. #LI-Remote

RapidSOS is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, or Veteran status. 

Interested in the role but you don’t meet 100% of the requirements? We’d love to hear from you! We encourage you to apply; we’d be excited to see if your unique skill set and experience could be a match.

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location
Listed under
Worldwide

Listing Details

Posted
April 21, 2026
First seen
April 21, 2026
Last seen
May 3, 2026

Posting Health

Days active
11
Repost count
0
Trust Level
56%
Scored at
May 3, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Rapidsos
Rapidsos
greenhouse

RapidSOS is an intelligent safety platform that securely links life-saving data from connected devices, apps, and sensors to 9-1-1 and first responders, empowering faster and more effective emergency response.

Employees
350
Founded
2012
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

RapidsosSenior Site Reliability Engineer$160k–$195k