~4d ago

Staff DevOps Engineer/SRE

IndiaIndia·Bangalorelead
EngineeringDevops Engineer
1 views0 saves0 applied

Quick Summary

Key Responsibilities

Design and evolve the infrastructure backbone for our AI and PaaS platform Build highly available, fault-tolerant, and scalable systems Define and drive SRE practices (SLIs, SLOs,

Technical Tools
EngineeringDevops Engineer

The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.

FlexAI is looking for a Staff DevOps / SRE Engineer to define our infrastructure strategy, establish SRE best practices, and build systems capable of running large-scale AI workloads across distributed, multi-cloud environments.


You’ll work closely with developers to ensure our platform is reliable, performant, and scalable — without slowing down product velocity.


Responsibilities

~1 min read
  • Design and evolve the infrastructure backbone for our AI and PaaS platform
  • Build highly available, fault-tolerant, and scalable systems
  • Define and drive SRE practices (SLIs, SLOs, error budgets)
  • Lead Infrastructure as Code using Pulumi
  • Own and scale Kubernetes clusters and containerized workloads
  • Standardize and automate infrastructure for global deployments
  • Design and scale CI/CD pipelines for fast, reliable releases
  • Build self-healing systems and automated remediation workflows
  • Drive GitOps and platform engineering practices
  • Implement end-to-end observability using VictoriaMetrics and Grafana (metrics, logs, traces)
  • Identify and resolve performance bottlenecks (latency, throughput, cost)
  • Lead incident response, root cause analysis, and postmortems
  • Partner with backend, AI, runtime, and security teams
  • Guide infrastructure decisions and scaling strategy
  • Mentor engineers and raise the bar on reliability and engineering standards
  • Embed security into infrastructure and deployment workflows
  • Design for resilience (disaster recovery, chaos testing, capacity planning)

  • 8+ years of experience in DevOps, SRE, or Infrastructure Engineering
  • Proven experience operating large-scale, distributed systems in production
  • Deep expertise in:
    • Kubernetes & container orchestration
    • Pulumi (or similar IaC tools)
    • Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
    • Observability stacks (Prometheus, Grafana, OpenTelemetry)
  • Strong experience with CI/CD, automation, and release engineering
  • Proficiency in Python, Go, or Bash
  • Strong systems thinking and debugging skills in high-scale environments
  • Experience defining and operating with SLOs / SLAs
  • Experience in startup environments
  • Comfortable leveraging AI coding tools and agents to move faster

Nice to Have

~1 min read
  • Experience with AI/ML infrastructure or GPU workloads
  • Familiarity with distributed or high-performance compute systems
  • Exposure to platform engineering / internal developer platforms
  • Experience scaling systems from Beta to production

  • Work on cutting-edge AI infrastructure
  • Build systems that power developers and enterprises
  • High ownership, fast execution, real impact
  • Collaborative, high-caliber team

Location & Eligibility

Where is the job
Bangalore, India
On-site at the office
Who can apply
IN

Listing Details

First seen
September 26, 2026
Last seen
September 30, 2026

Posting Health

Days active
4
Repost count
0
Trust Level
57%
Scored at
September 30, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Staff DevOps Engineer/SRE