soundhound
soundhound1mo ago
New

Staff Site Reliability Engineer

Toronto / CanadaRemotefull-timelead
OtherStaff Site Reliability Engineer
0 views0 saves0 applied

Quick Summary

Overview

The Opportunity We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability,

Technical Tools
OtherStaff Site Reliability Engineer
We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.

Responsibilities

~1 min read
  • →Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • →Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
  • →Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
  • →Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • →Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • →Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • →Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
  • →Lead department wide compliance (PCI, SOC) initiatives.


  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
  • Excellent communication skills and a demonstrated ability to mentor engineers.

Requirements

~1 min read
  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.



What We Offer

~1 min read
This role is available throughout Canada.

Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.

#LI-MQ1 #LI-REMOTE

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location

Listing Details

Posted
August 7, 2026
First seen
September 26, 2026
Last seen
September 26, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
23%
Scored at
September 26, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

soundhoundStaff Site Reliability Engineer