Senior DevOps Engineer, Observability
Quick Summary
5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field.
This is a senior infrastructure-focused role supporting reliable, scalable SaaS and on-premise environments. You will design and operate cloud infrastructure, automation, CI/CD pipelines, and observability systems that enable engineering teams to deliver software efficiently. The position combines hands-on DevOps, SRE, and platform engineering with a strong focus on reliability, performance, security, and automation. You will help build internal platforms and self-service capabilities while improving monitoring, alerting, incident response, and service-level objectives. The role offers the opportunity to work across complex infrastructure challenges in a fast-moving B2B software environment. You will collaborate closely with product and engineering teams and contribute to systems that operate at scale. It is well suited to an experienced engineer who treats infrastructure as code and internal platforms as products.
- Design, build, and maintain infrastructure supporting SaaS and on-premise engineering environments.
- Operate and optimize AWS infrastructure with a focus on cost efficiency, security, scalability, reliability, and performance.
- Manage cloud services including EC2, VPC, IAM, RDS, EKS, and Kubernetes-based environments.
- Develop and improve infrastructure-as-code using Terraform and Helm.
- Build and maintain CI/CD pipelines and developer automation, particularly using GitHub Actions.
- Contribute to internal platform tooling and developer self-service capabilities.
- Enhance observability and incident response systems through effective monitoring, alerting, instrumentation, and SLO management.
- Collaborate with stream-aligned product teams to understand infrastructure needs and continuously improve internal platforms.
- Support security and compliance initiatives, including implementation and maintenance of SOC 2 controls.
- Create and maintain technical documentation, onboarding resources, and internal support processes.
- Participate in an on-call rotation and contribute to reliable incident response and operational practices.
- Contribute to infrastructure initiatives involving event streaming, Change Data Capture (CDC), or large-scale, multi-tenant observability systems.
Requirements
~2 min read- 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field.
- At least 2 years of experience working within a B2B software startup or similarly fast-paced technology environment.
- Strong hands-on AWS experience, including services such as EC2, VPC, IAM, RDS, and EKS.
- Solid Kubernetes experience combined with infrastructure-as-code expertise using Terraform, Helm, or comparable tools.
- Experience designing and maintaining CI/CD pipelines and automation tooling, ideally with GitHub Actions.
- Familiarity with observability platforms and technologies such as Prometheus, Grafana, Mimir, Loki, or similar solutions.
- Proficiency in at least one scripting or programming language, such as Python, Go, or shell scripting.
- Experience with either Change Data Capture (CDC) and event-streaming systems or scaling large, multi-tenant observability platforms covering ingestion, analysis, and alerting.
- Strong communication, collaboration, troubleshooting, and technical documentation skills.
- Ability to work effectively in a fast-paced environment with ambiguity and evolving priorities.
- Experience in security-focused or SOC 2-compliant environments is a plus.
- Familiarity with MQTT, AMQP, or other messaging technologies is advantageous.
- Experience with network automation or related infrastructure ecosystems is a plus.
- Open-source contributions or experience working with open-source technologies is valued.
- Familiarity with AI-assisted development tools such as Copilot, ChatGPT, or Cursor is beneficial.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 29, 2026
- First seen
- September 29, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 80%
- Scored at
- September 29, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.