Quick Summary
Responsible for overseeing cloud operations, driving operational excellence,
Why we're hiring:
The role is responsible for leading and overseeing end-to-end cloud operations, ensuring the availability, reliability, security, performance, and resilience of cloud platforms and services. It manages service monitoring, incident response, and incident resolution for production applications and cloud infrastructure, while ensuring operations teams are skilled and enabled to execute cloud-related requests with speed and diligence. The role also provides governance over a large third-party managed services organisation, ensuring delivery against agreed KPIs, SLAs, and operational commitments through effective service reviews, metrics reporting, performance management, and continuous service improvement.
What you'll be doing:
Responsible for overseeing cloud operations, driving operational excellence, and improving the operational landscape through automation and AI-driven solutions delivered by internal resources and third-party partnerships.
- Work with product and engineering teams to define operational support patterns for each cloud product.
- Collaborate with business, architecture, engineering, cybersecurity, and platform teams to ensure seamless service delivery and continuous improvement.
The role oversees cloud infrastructure monitoring and ensures support is provided when issues occur across the following domains:
- Multi-cloud (AWS, Azure & GCP)
- Databases
- Citrix
- CMDB & scope management
- FinOps
- Patching
- Vulnerability Management
- Backup
- Incident, Problem, and Change Management
- Compliance & Audit support
- Define and execute the Cloud Operations strategy and operating model, aligning operational objectives with business and technology priorities.
- Own the operational governance of a large third-party managed services provider, ensuring delivery against contractual commitments, KPIs, SLAs, OLAs, and service quality standards.
- Drive vendor performance through Management Service Reviews (MSRs), governance forums, performance scorecards, and continuous service improvement initiatives.
- Provide executive oversight of ITSM processes, including Incident, Major Incident, Problem, Change, and Service Request Management.
- Act as the senior operational escalation point for critical incidents, ensuring effective stakeholder communication and timely service restoration.
- Establish and monitor operational dashboards, KPIs, service health metrics, and executive reporting to drive data-driven decision-making.
- Ensure cloud operations comply with security, risk, audit, regulatory, and governance requirements.
- Lead operational readiness, capacity planning, resilience, disaster recovery, and business continuity for cloud services.
- Collaborate with business, architecture, engineering, cybersecurity, and platform teams to ensure seamless service delivery and continual improvement.
- Own performance reporting and provide clear, actionable insights to support operational governance and decision-making.
- Ensure cloud infrastructure is regularly checked to maintain availability, meet performance SLAs, and complete performance tuning and runbook activities when required.
What you'll need:
- Bachelor’s or Master’s degree in computer science, Information Technology, Engineering, or a related field.
- Solid experience in IT Infrastructure and Cloud Operations, with significant experience leading enterprise-scale cloud operations.
- Proven experience managing cloud platforms (Azure, Google, and/or AWS Cloud Platform) and hybrid cloud environments.
- Demonstrated experience leading large Cloud Operations teams, including governance of third-party managed service providers.
- Proven experience managing vendor performance through KPIs, SLAs, OLAs, contractual governance, Management Service Reviews (MSRs), and continuous service improvement.
- Strong understanding of cloud operations, infrastructure management, site reliability, operational resilience, and service lifecycle management.
- Experience driving operational transformation through automation, AIOps, AI-enabled operations, and self-service capabilities.
- Strong knowledge of ITSM processes, including Incident, Major Incident, Problem, Change, Request, Availability, Capacity, and Service Level Management.
- Strong understanding of enterprise IT operations and service management platforms, including ServiceNow (e.g., ITOM Pro), and other industry leading monitoring, observability, and automation tools.
- Good understanding of cloud security, risk management, compliance, business continuity, and disaster recovery principles, along with familiarity with vulnerability management and security tooling such as Qualys, Microsoft Defender for Endpoint (MDE), and Wiz, as well as cloud-native security and patching capabilities across Azure, Google & AWS Cloud Platform.
- Familiarity with cloud-native monitoring, observability, and AIOps platforms, as well as leading enterprise solutions such as Grafana, Prometheus, ELK Stack, Dynatrace, Splunk, or similar technologies.
- Strong understanding of DevOps, Infrastructure as Code (IaC), CI/CD and modern cloud operational practices.
- Deep understanding of cloud architecture with hands-on experience in designing High-Level Designs (HLDs) and Low-Level Designs (LLDs) for new technology initiatives, proof of concepts (PoCs), and enterprise solutions, including AIOps, MLOps, and automation platforms.
- Excellent leadership, people management, coaching, and organisational development skills.
- Strong stakeholder management skills, with the ability to engage effectively with business leaders, technology teams, vendors, and executive management.
- Excellent analytical, decision-making, and problem-solving skills, with the ability to manage complex operational environments.
- Strong communication, presentation, and influencing skills, with experience presenting operational performance and strategic initiatives to senior leadership.
- ITIL certification (Foundation/Managing Professional/Expert) or equivalent service management experience is preferred.
- Relevant cloud certifications (Microsoft Azure, Google Cloud & AWS) are added advantage.
- Hands-on experience in IT Operations and support with organisations operating in Mode 1 and 2
- Understanding of Enterprise IT management frameworks (e.g. ITIL v4 processes, DevOps, Agile)
- Strong, creative problem solving and analytical in nature
- Self-motivated, energetic and tenacious and working well under pressure
- Ability to multi-task and prioritise deadlines is a must
- Good verbal and written communication and documentation skills with close attention to detail and accuracy
- A keen interest in new technologies and willingness to research and self-study to keep technical skills relevant
- Skilled in building a strong rapport with customer and internal support teams in a managed services environment
Who you are:
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- August 17, 2026
- First seen
- August 17, 2026
- Last seen
- August 17, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- August 17, 2026
Signal breakdown
Please let Wpp know you found this job on Jobera.
3 other jobs at Wpp
View all →Explore open roles at Wpp.
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.
