J
Jobgether3d ago
New

Lead Software Engineer, Cloud Site Reliability (SRE)

IndiaIndiaRemoteFull-timelead
OtherLead Software Engineer
0 views0 saves0 applied

Quick Summary

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Software Engineer, Cloud Site Reliability (SRE) based in India.

Technical Tools
OtherLead Software Engineer

This role leads cloud reliability and 24x7 site reliability operations for critical technology environments, with a strong focus on Azure infrastructure and cloud-native platforms. You will take ownership of major incidents, drive operational excellence, and ensure high availability and SLA adherence across production systems. The position combines hands-on cloud engineering with observability, automation, incident management, and reliability improvements. You will work extensively with Azure, AKS, Kubernetes, Docker, Datadog, and infrastructure-as-code technologies to build resilient and scalable environments. The role also provides an opportunity to advance proactive monitoring, anomaly detection, AIOps, and self-healing capabilities. As a technical leader, you will mentor engineers, collaborate with development teams, and communicate operational performance to stakeholders and leadership.

  • Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements.
  • Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution.
  • Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services.
  • Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements.
  • Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor.
  • Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users.
  • Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies.
  • Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices.
  • Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance.
  • Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders.
  • Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function.
  • Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate.

Requirements

~1 min read
  • 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments.
  • Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services.
  • Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning.
  • Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis.
  • Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting.
  • Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm.
  • Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies.
  • Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards.
  • Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles.
  • Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams.
  • Experience working in multi-cloud environments, particularly Azure and AWS, is a plus.
  • Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable.
  • Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous.
  • Bachelor’s degree or equivalent technical education and professional experience.

What We Offer

~2 min read
✓Fully remote opportunity based in India, with a rotational shift structure supporting 24x7 operations.
✓Leadership responsibility across cloud reliability, infrastructure operations, incident management, and operational excellence.
✓Hands-on exposure to Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, and Infrastructure as Code.
✓Opportunity to develop advanced observability, AIOps, predictive monitoring, automation, and self-healing capabilities.
✓Collaboration with engineering teams on cloud-native architecture, deployment pipelines, scalability, and reliability initiatives.
✓Opportunities to mentor team members and influence operational standards and engineering practices.
✓Exposure to multi-cloud technologies and large-scale distributed systems.
✓An inclusive work environment focused on teamwork, openness, respect, fairness, and continuous improvement.
✓Support for professional development through exposure to modern cloud and reliability technologies and relevant certification paths.
✓Opportunities to participate in leadership reporting, business reviews, and cross-functional initiatives with broad organizational visibility.

Location & Eligibility

Where is the job
India
Remote within one country
Who can apply
IN

Listing Details

Posted
September 24, 2026
First seen
September 27, 2026
Last seen
September 27, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
September 27, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

J
Lead Software Engineer, Cloud Site Reliability (SRE)