Quick Summary
About FlexAI Build and Deploy AI the right way, anywhere. The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud,
The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.
FlexAI is looking for a Senior DevOps / SRE Engineer to build and operate the infrastructure powering our AI and PaaS platform.
Responsibilities
~1 min read- →Build and maintain infrastructure for our AI and PaaS platform
- →Deploy and operate Kubernetes clusters and containerized services
- →Implement Infrastructure as Code using Pulumi (or similar tools)
- Help define and implement SLIs, SLOs, and error budgets
- Improve system reliability, availability, and performance
- Participate in on-call rotations, incident response, and postmortems
- Build and improve CI/CD pipelines for reliable and fast releases
- Automate operational workflows and reduce manual toil
- Contribute to GitOps and platform engineering practices
- Implement and maintain observability using VictoriaMetrics, Grafana (metrics, logs, traces)
- Monitor systems and troubleshoot performance issues (latency, throughput, cost)
- Work closely with developers, platform, and AI teams to support production systems
- Help debug issues across infrastructure and application layers
- Contribute to improving engineering productivity and developer experience
What You’ll Need to Be Successful
- 4+ years of experience in DevOps, SRE, or Infrastructure Engineering
- Experience operating production systems at scale
- Hands-on experience with:
- Kubernetes & containers
- Infrastructure as Code (Pulumi, Terraform, etc.)
- Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
- Observability tools (Prometheus, Grafana, OpenTelemetry)
- Experience with CI/CD systems and automation
- Proficiency in Python, Go, or Bash
- Strong debugging and problem-solving skills
- Familiarity with SLOs and reliability practices
- Experience working in startup or fast-paced environments
- Comfortable leveraging AI coding tools and agents
Nice to Have
~1 min read- Experience with AI/ML infrastructure or GPU workloads
- Familiarity with distributed systems or compute platforms
- Exposure to platform engineering concepts
- Experience supporting systems from Beta to production
- Work on cutting-edge AI infrastructure
- Build systems that power developers and enterprises
- High ownership, fast execution, real impact
- Collaborative, high-caliber team
Location & Eligibility
Listing Details
- First seen
- September 26, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 4
- Repost count
- 0
- Trust Level
- 57%
- Scored at
- September 30, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.