Lead Software Engineer, Cloud Site Reliability (SRE)
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Software Engineer, Cloud Site Reliability (SRE) based in India.
This role leads cloud reliability and 24x7 site reliability operations for critical technology environments, with a strong focus on Azure infrastructure and cloud-native platforms. You will take ownership of major incidents, drive operational excellence, and ensure high availability and SLA adherence across production systems. The position combines hands-on cloud engineering with observability, automation, incident management, and reliability improvements. You will work extensively with Azure, AKS, Kubernetes, Docker, Datadog, and infrastructure-as-code technologies to build resilient and scalable environments. The role also provides an opportunity to advance proactive monitoring, anomaly detection, AIOps, and self-healing capabilities. As a technical leader, you will mentor engineers, collaborate with development teams, and communicate operational performance to stakeholders and leadership.
- Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements.
- Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution.
- Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services.
- Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements.
- Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor.
- Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users.
- Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies.
- Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices.
- Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance.
- Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders.
- Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function.
- Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate.
Requirements
~1 min read- 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments.
- Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services.
- Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning.
- Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis.
- Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting.
- Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm.
- Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies.
- Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards.
- Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles.
- Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams.
- Experience working in multi-cloud environments, particularly Azure and AWS, is a plus.
- Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable.
- Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous.
- Bachelor’s degree or equivalent technical education and professional experience.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 24, 2026
- First seen
- September 27, 2026
- Last seen
- September 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 27, 2026
Signal breakdown
Similar Lead Software Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.