T
Ttecdigital1mo ago
New
New
Site Reliability Engineer (SRE)
EngineeringDevops Engineer
5 views0 saves0 applied
Quick Summary
Overview
At TTEC Digital, we coach clients to ensure their employees feel valued, and fully supported, because an amazing customer experience is an employee first process. Our vision is the same,
Technical Tools
EngineeringDevops Engineer
At TTEC Digital, we coach clients to ensure their employees feel valued, and fully supported, because an amazing customer experience is an employee first process. Our vision is the same, a place where employees know they can thrive.
- Own production reliability for a real-time platform where uptime and latency ARE the product — voice, desktop, intelligence, and AI combined; an agent mid-call can't wait for a retry.
- First SRE hired immediately (Day 0–14) for production scaling and SLO ownership; a second joins at the start of Phase 3 for 24/7 coverage.
- Pairs with C1 Platform Foundation on observability and tenancy isolation.
- Startup environment: weekly deploys, 1-week sprints, fail fast, move forward — reliability engineering at that speed, not against it.
- SLOs and error budgets per tenant/service
- Incident response and blameless postmortems
- Production scaling and capacity
- Observability depth (p50/p95/p99 per event hop)
- Uptime as a personal mission
- On-call rotation with DevOps
- Your committed timelines.
- Self-starter, grit, show-me mentality — you prove reliability with dashboards and drills, not assertions.
- A ways-to-YES engineer: weekly deploys are the heartbeat and your job is making them safe, never slowing them.
- You love new technology, adapt fast when the stack changes under you, use AI tools daily to multiply velocity, and consider yourself exceptional.
- Calm in an incident, relentless after it.
- Team player who likes winning.
- 8+ years operating production systems at scale; owns SLOs, error budgets, incident command.
- Strong Go or Python — you automate reliability, you don't toil at it. Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero.
- Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production.
- Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.
- Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.
- GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus.
- Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.
- Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it.
- Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.
- AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.
- Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.
- A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.
Responsibilities
~2 min read- →8+ years operating production systems at scale; owns SLOs, error budgets, incident command.
- →Strong Go or Python — you automate reliability, you don't toil at it. Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero.
- →Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production.
- →Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.
- →Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.
- →GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus.
- →Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.
- →Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it.
- →Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.
- →AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.
- →Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.
- →A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.
Location & Eligibility
Where is the job
Hyderabad, India
Hybrid — some on-site time required
Who can apply
IN
Listing Details
- Posted
- July 17, 2026
- First seen
- July 17, 2026
- Last seen
- September 4, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 62%
- Scored at
- July 17, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on Ttecdigital's site
Please let Ttecdigital know you found this job on Jobera.
4 other jobs at Ttecdigital
View all →Explore open roles at Ttecdigital.
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Security2.2kFullstack Developer2.2kDevOps & Infrastructure2.1kEngineering Manager2.1kSoftware Architect1.8kQa Engineer1.7kBackend Developer1.6kMechanical Engineer1.5kSecurity Engineer1.4kFrontend Developer1.1kElectrical Engineer1.1kMobile Developer1kProject Engineer963Data Engineering956Backend Engineering937Design Engineer848Product Engineer527Embedded Engineer516Automation Engineer516Frontend Engineering496
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.
T
Site Reliability Engineer (SRE)