Senior Site Reliability Engineer - Ceph / On-Prem Cloud Storage (m/f/d)
Quick Summary
Your Role You will help us build, operate and industrialize the storage foundation of our on-premise cloud platform. As part of a small, experienced team,
You will help us build, operate and industrialize the storage foundation of our on-premise cloud platform. As part of a small, experienced team, you will own the Ceph clusters behind our block, object and file storage services – essentially the storage every customer workload eventually lands on. The platform is still actively being built, which means you will have real influence on the storage architecture, hardware lifecycle, automation strategy and technical direction. The role goes beyond Ceph itself: you will regularly work across bare metal, networking, OpenStack, Kubernetes and GitOps. As a Senior Engineer, you take ownership of your area and help make storage predictable, durable and highly automated as we scale into the petabytes
· Ceph · OpenStack · Kubernetes · KVM · Linux · Bare Metal · Ansible · Terraform · Go · Cinder ·Neutron · Nova · RBD · Cilium
· RGW/ S3 · Prometheus/ Grafana · FluxCD / ArgoCD · Git· Python · Claude Code · Cursor · Agentic Coding Tooling
Responsibilities
~2 min read- →
Design, operate and evolve our Ceph clusters for block, object / S3 and file storage - from CRUSH topology and failure domains to pool design, placement groups, erasure coding, capacity headroom and performance.
- →
Own the storage hardware lifecycle end to end – from qualification and burn-in through firmware, controller and disk health to OSD add/drain/replace, node retirement and hardware refresh. A key part of the job is automating these processes without creating tenant-visible impact.
- →
Automate the provisioning of storage nodes on OpenStack-managed bare metal, working with Ironic, RAID and disk configuration, SEDs, provider networks and VLANs. You will also own the artefact path around images, package mirrors, certificates and other dependencies needed to bring up clusters in environments with tightly controlled external connectivity.
- →
Run upgrades, rebalancing and reconfigurations as routine production operations, evolve our Infrastructure as Code and GitOps setup with Ansible, Terraform and FluxCD / ArgoCD, and make sure storage integrates cleanly with both OpenStack and Kubernetes.
- →
Own capacity planning and growth forecasting, use telemetry to understand real consumption, and build out monitoring, testing and failure-injection around performance, durability and security.
- →
AI-assisted engineering is simply part of how we work today. You use LLMs and agentic tools where they genuinely help – across development, testing, reviews, incident triage, knowledge retrieval or automation - and help us push further into self-healing and auto-remediation.
- →
Act as a technical reference and sparring partner across Ceph, storage, automation and AI tooling, document technical decisions and share your knowledge with the team.
What We Offer
~3 min readLocation & Eligibility
Listing Details
- First seen
- October 1, 2026
- Last seen
- October 1, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 56%
- Scored at
- October 1, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.