soundhound1mo ago
New
New
Staff Site Reliability Engineer
Toronto / CanadaRemotefull-timelead
OtherStaff Site Reliability Engineer
0 views0 saves0 applied
Quick Summary
Overview
The Opportunity We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability,
Technical Tools
OtherStaff Site Reliability Engineer
We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.
Responsibilities
~1 min read- →Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- →Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
- →Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
- →Partner with engineering teams to optimize performance, cost, and reliability of backend services.
- →Drive incident response, post-mortem analysis, and long-term remediation efforts.
- →Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
- →Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
- →Lead department wide compliance (PCI, SOC) initiatives.
- 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
- Experience designing and managing high-throughput, distributed systems.
- Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
- Excellent communication skills and a demonstrated ability to mentor engineers.
Requirements
~1 min read- Experience working in a high-velocity, customer-focused environment.
- Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
- Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
- Experience implementing security and compliance best practices in the cloud.
What We Offer
~1 min readThis role is available throughout Canada.
Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.
#LI-MQ1 #LI-REMOTE
Location & Eligibility
Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location
Listing Details
- Posted
- August 7, 2026
- First seen
- September 26, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 23%
- Scored at
- September 26, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Staff Site Reliability Engineer jobs
View all →Staff Site Reliability Engineer
Staff Site Reliability Engineer
Remote
Staff Site Reliability Engineer
USD 199750-270000
full-timeRemote
Staff Site Reliability Engineer
Staff Site Reliability Engineer, Environment Automation
Staff Site Reliability Engineer
full-timeRemote
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.