Senior Site Reliability Engineer
Quick Summary
Prometheus + Grafana
2K is headquartered in Novato, California and is a wholly owned label of Take-Two Interactive Software, Inc. (NASDAQ: TTWO). Founded in 2005, 2K Games is a global video game company, publishing titles developed by some of the most influential game development studios in the world. Our studios responsible for developing 2K’s portfolio of world-class games across multiple platforms, include Visual Concepts, Firaxis, Hangar 13, CatDaddy, Cloud Chamber, 31st Union, HB Studios, and 2K SportsLab. Our portfolio of titles is expanding due to our global strategic plan, building and acquiring exciting studios whose content continues to inspire all of us! 2K publishes titles in today’s most popular gaming genres, including sports, shooters, action, role-playing, strategy, casual, and family entertainment.
Our team of engineers, marketers, artists, writers, data scientists, producers, thinkers and doers, are the professional publishing stewards of 2K’s portfolio currently includes several AAA, sports and entertainment brands, including global powerhouse NBA®️ 2K, renowned BioShock®️, Borderlands®️, Mafia, Sid Meier’s Civilization®️ and XCOM®️ brands; popular WWE®️ 2K and WWE®️ SuperCard franchises, TopSpin 2K25, as well as the critically and commercially acclaimed PGA TOUR®️ 2K
At 2K, we pride ourselves on creating an inclusive work environment, which means encouraging our teams to Come as You Are and do your best work! We encourage ALL applicants to explore our global positions, even if they don’t meet every requirement for the role. If you're interested in the job and think you have what it takes to work at 2K, we encourage you to apply!
The 2K SRE team owns the infrastructure behind every player connection — All 2K game services, account platforms, CI/CD pipelines, and developer tooling spanning AWS, GCP, and on-premises data centers across multiple global regions. Global launch windows and live-service events push systems to their limits, and this team is expected to hold the line.
Post-mortems here focus on systems, not people. Automation is the default answer to repetitive work. The infrastructure keeps millions of players connected — and the team takes that seriously!
Requirements
~1 min readThe Senior SRE at 2K is a hands-on technical leader — shaping production infrastructure across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing the gap between what engineering ships and what players experience.
- Live-service game or large-scale consumer internet experience at millions of concurrent users
- Service mesh depth (Istio, Cilium) and advanced Kubernetes networking
- FinOps and managing resources at cloud scale
- Experience with AI and Agentic Development
- Cloud certifications (AWS Solutions Architect, GCP Professional Cloud Architect, CKA/CKS, or equivalent)
- Experience mentoring SREs or leading reliability working groups
Responsibilities
~1 min readDesign, build, and operate scalable multi-cloud and hybrid infrastructure using Terraform, Pulumi, and GitOps workflows (ArgoCD, Flux). Own Kubernetes platforms (EKS, GKE) end-to-end — cluster lifecycle, multi-tenancy, networking (Istio, Cilium), and autoscaling — and push progressive delivery patterns (blue/green, canary) across game service deployments.
- Build and run the full observability stack: Prometheus + Grafana + Datadog
- Define SLI/SLO/error budget policies and build alerting that cuts through the noise
- Lead chaos engineering exercises to surface failure modes before players encounter them
- Drive incident response and post-mortems with a focus on systemic fixes and real follow-through
Eliminate toil through self-service provisioning, automated remediation, and intelligent scaling. Harden CI/CD pipelines (GitHub Actions, Jenkins, ArgoCD) . Embed security at the platform layer through secrets management (PasswordState, 1Password, and AWS Secrets Manager), policy-as-code (OPA/Gatekeeper).
- Promote SRE practices across 2K studios through reliability reviews, runbooks, and embedded collaboration
- Shape architectural decisions and author engineering RFCs that move the platform forward
- 5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale
- Deep Kubernetes experience in cloud environments (EKS or GKE preferred) — networking, storage, multi-cluster patterns
- Strong IaC proficiency with Terraform and/or Pulumi; hands-on with Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions)
- Modern and Legacy Tech: AWS, GCP, VMware, and Bare metal servers
- Server Configuration using Ansible, Puppet, and AWS Systems Manager
- Observability stack experience: Datadog, Prometheus + Grafana, and OpenTelemetry,
- SLI/SLO/error budget fluency — including how to operationalize them inside engineering teams
- Production-quality code in Go, Python, or TypeScript: tools, automation, and internal libraries
- Linux internals, TCP/IP networking, DNS, and TLS — proven enough to debug at the system level
- Incident response and post-mortem leadership with a track record of systemic follow-through
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- August 17, 2026
- First seen
- August 17, 2026
- Last seen
- August 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 53%
- Scored at
- August 17, 2026
Signal breakdown
Please let 2K know you found this job on Jobera.
4 other jobs at 2K
View all →Explore open roles at 2K.
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.