Quick Summary
Key Responsibilities
serve in the on-call rotation, coordinate response during outages, run blameless postmortems, and drive remediation to completion Define and track SLOs, SLIs, and error budgets
Technical Tools
Other
As a Site Reliability Engineer, you'll be responsible for the availability, scalability, and operational health of our AWS-hosted infrastructure. You'll lead incident response, build the automation that lets our engineering teams ship safely and often, and drive a culture of measurable reliability across the organization.
Responsibilities
~1 min read- →Own the reliability and performance of production services running in AWS, including capacity planning, cost optimization, and architecture reviews
- →Design, build, and maintain fully automated CI/CD pipelines that take code from commit to production with minimal manual intervention
- →Lead incident management: serve in the on-call rotation, coordinate response during outages, run blameless postmortems, and drive remediation to completion
- →Define and track SLOs, SLIs, and error budgets in partnership with product and engineering teams
- →Build and improve observability through monitoring, logging, alerting, and distributed tracing
- →Manage infrastructure as code and eliminate toil through automation
- →Partner with development teams to embed reliability best practices into system design and release processes
- →Contribute to disaster recovery planning, security hardening, and compliance efforts
- 4+ years in SRE, DevOps, or cloud operations roles supporting production systems
- Deep hands-on experience operating workloads in AWS (e.g., EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch)
- Proven experience with incident management: on-call ownership, incident command, root cause analysis, and postmortem processes
- Demonstrated track record of building fully automated CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline, or similar)
- Strong infrastructure-as-code skills with Terraform, CloudFormation, or CDK
- Proficiency in at least one scripting or programming language (Python, Go, Bash)
- Experience with containers and orchestration (Docker, Kubernetes)
- Familiarity with observability tooling such as Datadog, Prometheus/Grafana, or the ELK stack
- Clear communicator who stays calm under pressure and can explain complex issues to technical and non-technical audiences
Nice to Have
~1 min read- AWS certifications (Solutions Architect, DevOps Engineer)
- Experience with GitOps workflows (ArgoCD, Flux)
- Background in security operations, compliance frameworks (SOC 2, ISO 27001)
- Tech-First Culture: We believe in building smart, scalable systems—and we invest in them.
- Real-World Impact: Your work will touch thousands of users every day, improving workflows and outcomes.
- Autonomy + Collaboration: Own your space while being part of a highly connected, supportive team.
- Growth-Minded Environment: We prioritize learning, innovation, and pushing the limits of what’s possible.
What We Offer
~1 min read✓Competitive salary
✓Comprehensive health, dental, and vision coverage
✓401(k) with company match
✓Flexible PTO and hybrid work arrangement in Atlanta
Location & Eligibility
Where is the job
Atlanta, US
On-site at the office
Listing Details
- Posted
- September 22, 2026
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 24%
- Scored at
- September 30, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.