Site Reliability Engineer
Quick Summary
Observability Environment Management: Design, build, and maintain our observability infrastructure, including monitoring tools, logging platforms, and distributed tracing systems (e.g., Prometheus,
Bachelors degree in computer science or a related field, or equivalent experience. 5+ years of experience as an SRE or in a similar role with a focus on observability.
We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our growing Observability team. The ideal candidate will have a strong background in building and maintaining robust observability environments, including monitoring, logging, and tracing systems. This role will focus on the design, implementation, and support of our observability infrastructure, ensuring the seamless onboarding of applications and providing critical support during incidents.
Responsibilities
~1 min read- →Observability Environment Management: Design, build, and maintain our observability infrastructure, including monitoring tools, logging platforms, and distributed tracing systems (e.g., Prometheus, Grafana, Elasticsearch, etc.). This includes capacity planning, performance tuning, and ensuring high availability.
- →Application Onboarding: Work with development teams to onboard applications to our observability platform, providing guidance on instrumentation best practices and ensuring data quality. This includes creating and maintaining documentation and training materials.
- →Incident Support: Provide timely and effective support during incidents, leveraging observability data to diagnose and resolve issues quickly. This includes contributing to post-incident reviews and implementing preventative measures.
- →Automation: Automate repetitive tasks and processes related to observability, improving efficiency and reducing manual effort. This may involve scripting, developing tools, or integrating with CI/CD pipelines.
- →Alerting and Monitoring: Develop and maintain effective alerting strategies, ensuring appropriate escalation procedures and minimizing noise. This includes creating dashboards and reports to visualize system health and performance.
Requirements
~1 min read- Bachelors degree in computer science or a related field, or equivalent experience.
- 5+ years of experience as an SRE or in a similar role with a focus on observability.
- Strong understanding of distributed systems and microservices architectures.
- Experience with any monitoring, logging, and tracing tools (e.g., Prometheus, Grafana, Jaeger, Elasticsearch, Fluentd, Datadog, Dynatrace, etc.).
- Proficiency in scripting languages such as Python, Go, or Bash.
- Strong problem-solving and analytical skills.
- Excellent communication and collaboration skills.
Nice to Have
~1 min read- Experience with cloud platforms.
- Experience with infrastructure-as-code tools (e.g., Terraform, Ansible)
Location & Eligibility
Listing Details
- First seen
- September 25, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 56%
- Scored at
- September 27, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.