Staff Site Reliability Engineer-Observability
Quick Summary
auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team Consult with partner dev teams on metrics, alert thresholds,
GEDE (Global Edge and Domains Engineering) keeps GoDaddy's domains, edge, and aftermarket platforms running for millions of customers. Within GEDE, the Reliability Engineering team (Domains Production Engineering) is the group that gets the call when something breaks — and, more importantly, builds the systems that mean it breaks less often. We own the monitoring, compliance, and patching toolchain for the whole Domains infrastructure footprint, provide advanced incident support, and are actively modernizing how the org detects, diagnoses, and even auto-remediates issues with AI-assisted tooling!
Responsibilities
~1 min read- →Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7 — closing gaps across the org.
- →Steer our cloud migration journey and bolster our efforts to keep the services reliable and performant.
- →Own patching compliance and vulnerability remediation at scale across a mixed on-prem + AWS fleet, hitting hard SLA targets. Operate and extend our multi-tenant Kubernetes/ArgoCD platform, including the migration of core services.
- →Contribute to our AI/automation initiatives: auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team
- →Consult with partner dev teams on metrics, alert thresholds, and monitoring standards — this is a platform-enablement role, not just a ticket queue. Mentor other engineers on the team and help mature our operational practices.
- 8+ years of hands-on experience working with AWS, both in cloud-native and agnostic capacities.
- 5+ years of experience with Kubernetes (EKS/AKS/GKE/Fargate, etc.).
- 4+ years of expertise in Linux administration.
- Strong coding skills in languages such as Go, Python, Ruby etc.
- Strong experience in coding infrastructure as code (Terraform, Ansible, CDK, etc.).
- Strong understanding of CI/CD concepts, version control systems, and testing tools (e.g., Jenkins, Gradle, Maven).
- Hands-on experience with SQL databases, especially MySQL or PostgreSQL.
- Verified background in developing and maintaining distributed infrastructure and systems.
- Proven understanding of networking principles and a strong dedication to cybersecurity guidelines.
- Ability to work independently and perform effectively under pressure.
We encourage you to apply even if your experience or skillset doesn’t align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth.
Requirements
~1 min readLocation & Eligibility
Listing Details
- Posted
- July 29, 2026
- First seen
- July 29, 2026
- Last seen
- July 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- July 29, 2026
Signal breakdown
GoDaddy helps the world easily start, confidently grow, and successfully run an online presence.
View company profilePlease let GoDaddy know you found this job on Jobera.
3 other jobs at GoDaddy
View all →Explore open roles at GoDaddy.
Similar Staff Site Reliability Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.