New

Graphite - Site Reliability Engineer (SRE)

MexicoMexico·Guadalajarafull-timemid
EngineeringDevops Engineer
1 views0 saves0 applied

Quick Summary

Key Responsibilities

Design, implement, and maintain robust monitoring and alerting systems (e.g., GCP Monitoring, Prometheus, Grafana, Traces,

Technical Tools
EngineeringDevops Engineer
We're looking for a passionate and hands-on Site Reliability Engineer (SRE) to join our team. This role is critical for ensuring the stability, performance, and scalability of our production services. You'll be the bridge between development and operations, with a strong focus on using code to manage infrastructure and eliminate toil.

Responsibilities

~1 min read
  • →Monitoring and Alerting: Design, implement, and maintain robust monitoring and alerting systems (e.g., GCP Monitoring, Prometheus, Grafana, Traces, Logs) to provide visibility into application performance and infrastructure health.
  • →Infrastructure Management: Build, provision, and maintain our core infrastructure, with a strong emphasis on Cloud environments and Kubernetes clusters.
  • →Automation and Tooling: Write and maintain scripts and automation workflows (e.g., Python, Bash, TypeScript (Pulumi)) to streamline deployment, scaling, and operational tasks, embracing the philosophy of "automating everything."
  • →Incident Response: Provide hands-on, real-time incident response and participate in an on-call rotation to quickly mitigate service disruptions and restore functionality.
  • →Production Debugging: Deeply debug and troubleshoot complex production problems across the entire stack, from network issues to application code defects.
  • →Process Improvement: Conduct blameless post-mortems for major incidents, implementing long-term solutions to prevent recurrence and continuously improve service reliability.

Requirements

~1 min read
  • Proven experience as an SRE, DevOps Engineer, or similar role.
  • Expertise in managing and scaling Kubernetes in a production environment.
  • Strong proficiency in a scripting or programming language (e.g., Python, Go, Bash).
  • Deep understanding of monitoring, logging, and alerting best practices.
  • Solid experience with at least one major Cloud provider (AWS, GCP, or Azure).
  • Experience with Infrastructure as Code (IaC) tools like Terraform or Pulumi is a plus.
A proactive, data-driven approach to reliability and a passion for managing complex systems at scale.


What We Offer

~1 min read
The base pay range for this role is $50,000 – $60,000 per year.

Location & Eligibility

Where is the job
Guadalajara, Mexico
Hybrid — some on-site time required
Who can apply
MX

Listing Details

Posted
October 10, 2025
First seen
September 26, 2026
Last seen
September 30, 2026

Posting Health

Days active
3
Repost count
0
Trust Level
23%
Scored at
September 30, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Graphite - Site Reliability Engineer (SRE)