Graphite - Site Reliability Engineer (SRE)
EngineeringDevops Engineer
1 views0 saves0 applied
Quick Summary
Key Responsibilities
Design, implement, and maintain robust monitoring and alerting systems (e.g., GCP Monitoring, Prometheus, Grafana, Traces,
Technical Tools
EngineeringDevops Engineer
We're looking for a passionate and hands-on Site Reliability Engineer (SRE) to join our team. This role is critical for ensuring the stability, performance, and scalability of our production services. You'll be the bridge between development and operations, with a strong focus on using code to manage infrastructure and eliminate toil.
Responsibilities
~1 min read- →Monitoring and Alerting: Design, implement, and maintain robust monitoring and alerting systems (e.g., GCP Monitoring, Prometheus, Grafana, Traces, Logs) to provide visibility into application performance and infrastructure health.
- →Infrastructure Management: Build, provision, and maintain our core infrastructure, with a strong emphasis on Cloud environments and Kubernetes clusters.
- →Automation and Tooling: Write and maintain scripts and automation workflows (e.g., Python, Bash, TypeScript (Pulumi)) to streamline deployment, scaling, and operational tasks, embracing the philosophy of "automating everything."
- →Incident Response: Provide hands-on, real-time incident response and participate in an on-call rotation to quickly mitigate service disruptions and restore functionality.
- →Production Debugging: Deeply debug and troubleshoot complex production problems across the entire stack, from network issues to application code defects.
- →Process Improvement: Conduct blameless post-mortems for major incidents, implementing long-term solutions to prevent recurrence and continuously improve service reliability.
Requirements
~1 min read- Proven experience as an SRE, DevOps Engineer, or similar role.
- Expertise in managing and scaling Kubernetes in a production environment.
- Strong proficiency in a scripting or programming language (e.g., Python, Go, Bash).
- Deep understanding of monitoring, logging, and alerting best practices.
- Solid experience with at least one major Cloud provider (AWS, GCP, or Azure).
- Experience with Infrastructure as Code (IaC) tools like Terraform or Pulumi is a plus.
A proactive, data-driven approach to reliability and a passion for managing complex systems at scale.
What We Offer
~1 min readThe base pay range for this role is $50,000 – $60,000 per year.
Location & Eligibility
Where is the job
Guadalajara, Mexico
Hybrid — some on-site time required
Who can apply
MX
Listing Details
- Posted
- October 10, 2025
- First seen
- September 26, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 3
- Repost count
- 0
- Trust Level
- 23%
- Scored at
- September 30, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Browse Similar Jobs
Fullstack Developer2.8kEngineering Manager2.3kSecurity2.1kQa Engineer2.1kSoftware Architect2.1kMechanical Engineer2kElectrical Engineer1.8kProject Engineer1.7kBackend Developer1.6kSecurity Engineer1.5kFrontend Developer1.2kDesign Engineer1.2kDevOps & Infrastructure1.2kQuality Engineer1kProcess Engineer904Mobile Developer787Product Engineer754Automation Engineer617Application Engineer613Embedded Engineer608
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.