Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)
Quick Summary
Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools. - Support health-check frameworks including Airflow health-check
Responsibilities
~1 min read- Own platform reliability practices for availability, resilience, latency, and operational efficiency.
- Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.
- Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.
- Lead JFROG Helm chart automation and JFROG images/ACR migration work.
- Support microservices deployment enablement and platform/tooling upgrades.
- Own and optimize monitoring, alerting, observability, and logging stack components:
Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.
- Support health-check frameworks including Airflow health-check requirements.
- Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.
- Collaborate with architecture and delivery teams on reliability and scalability patterns.
- Lead cloud infrastructure creation, maintenance, governance, and access controls.
- Drive capacity planning, DR planning/exercises, and platform best-practice documentation.
- Support cost management, role enforcement, and license management governance.
- Maintain SOP documentation for established alerts and incident patterns.
Requirements
~1 min read- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
- Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.
- Advanced experience with CI/CD engineering and GitHub Actions.
- Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.
- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
- End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).
- Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).
- Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.
- Experience in governance controls: access management, role enforcement, and separation of duties.
- Proven high-severity incident leadership and post-incident reliability improvement execution.
- Postgres performance and reliability operations.
- Telecom-scale high-availability systems experience.
Senior to Lead IC (typically 10 to 17 years)
Onsite (Hyderabad / Bangalore or designated AT&T location)
What We Offer
~1 min read- Opportunity to define and scale platform reliability standards.
- High technical ownership and strong cross-functional influence.
- Enterprise-scale impact across observability, automation, and resilience engineering.
It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.
Location & Eligibility
Listing Details
- Posted
- July 30, 2026
- First seen
- August 4, 2026
- Last seen
- August 5, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 27%
- Scored at
- August 4, 2026
Signal breakdown
Please let att know you found this job on Jobera.
Similar Support jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.