Staff Software Engineer, Infrastructure
Quick Summary
Staff Software Engineer, Platform & Infrastructure The role We are hiring a Staff Software Engineer for our Infrastructure team.
We are hiring a Staff Software Engineer for our Infrastructure team.
Stream runs behind apps that more than a billion people use: billions of API requests a month, a 99.999% uptime promise, and a small team keeping it all running. Infrastructure here decides how fast our products feel, how well they scale, and how quickly we can launch new ones.
This year we're wrapping up a big move. We're simplifying where and how we host, and rebuilding our data layer on self-hosted, open-source systems that are faster and cheaper to run. That's the groundwork, not the job. Next comes building the infrastructure every Stream product runs on, including the ones we haven't built yet. That means real SLOs, clear costs, calmer on-call, and a platform that helps teams ship faster.
What you build will shape how much Stream spends and how fast it grows, and we'll trust you with big decisions from your first few weeks.
This is a Software Engineering role first. Attitude, ambition and a willingness to experiment across the stack matter more to us than years of Kubernetes or database operations. We expect AI to cover a good share of the DevOps-heavy work, and we expect you to use it that way. What a tool can't give us is system design judgement, strong fundamentals and good code, so that is what we look for.
Platform & Infrastructure is a small, senior team without the support structures of a large organisation. It moves fast and it gets hectic. We would rather say that now than have you find out in month two.
Runs Stream's global platform. Keeps the edge and backend that serve customers worldwide fast and up, at single-digit millisecond latency on a 99.999% uptime SLA.
Shapes Stream's margins. Our hosting and infrastructure decisions directly drive the company's profitability. You will decide where and how we run, and what we stop paying for.
Builds tools people actually want to use. Internal tools that change how engineers deploy and debug, and external ones that give customers a clear view of their own traffic and health.
Makes on-call boring. Builds a rotation people can sustain, and cuts alert noise until a page means something real is broken.
Makes cost visible. Gives every team a clear view of what it spends on infrastructure and why, so cost becomes part of every engineering decision.
Responsibilities
~1 min read- →
Define what "good" looks like for the platform. Set latency and uptime SLOs for every product team, and decide what we measure, from error rates and delivery times to how healthy our customers' apps really are.
- →
Make the big infrastructure calls. Lead the design reviews and decisions on where and how we run, working with backend, video and moderation engineers on tradeoffs that cross service boundaries.
- →
Build the platform tooling. Write the internal tools and services that help engineers ship, deploy and debug faster, plus the docs and AI skills that make launching a new product repeatable.
- →
Turn incidents into fixes that stick. Take part in on-call, and make sure every root cause ends in a durable fix, not a workaround.
- →
Make cost part of every design. Weigh cost alongside reliability in every decision, and give teams the numbers to do the same.
A software engineer who has built systems, not only configured them. You write production code in Go, Python or a similar language. Scripting-only backgrounds are not a fit.
Strong system design. You can reason about failure modes, capacity and cost in distributed systems, and explain why you chose what you chose.
Solid fundamentals in networking, storage and concurrency, and the habit of asking why a system behaves the way it does instead of accepting the default.
You have owned high-scale production systems and set technical direction beyond your own work, through design reviews, RFCs and decisions other engineers build on.
You experiment. When you hit Kubernetes, a database or a cloud service you have not run before, you dig in, use AI to close the gap, and ship.
You already use AI tools every day in your engineering work.
You treat cost and reliability as engineering problems, and you can put a number on an outcome you drove.
You write clearly. Much of this role is docs and standards that other teams will follow.
You are comfortable in a small team, setting direction and reviewing a PR in the same week.
Nice to Have
~1 min readKubernetes in production: cluster architecture, workload design or a migration you led.
PostgreSQL at scale, ideally self-hosted or with CloudNativePG: sharding, replication, partitioning tradeoffs.
Valkey or Redis at scale, ideally on Kubernetes.
Hands-on AWS and GCP, including multi-cloud setups or migrations between providers.
Cloud commitment and reservation strategy (committed use discounts, savings plans), or formal FinOps practice.
SLO design, alert hygiene and a Prometheus and Grafana based observability stack.
Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.
Open source contributions, or writing and talks on cloud, platform or distributed systems.
Go, gRPC, RocksDB, Python
PostgreSQL, RabbitMQ
GCP
Grafana, Prometheus, ELK (Elasticsearch and Kibana)
Jaeger and Tempo for distributed tracing, Datadog
Redis, Memcached
Claude Code, Cursor
You want platform and infrastructure problems at a scale most engineers never touch, and the autonomy to set the direction yourself.
You would rather try something, measure it and change course than wait for a perfect plan.
You ship fast and learn fast, including when it is hectic.
You are self-directed and comfortable working with a globally distributed team across time zones.
You want to stay a specialist in one layer of the stack.
You want tightly scoped tickets and step-by-step direction.
You need a calm, highly predictable environment.
What We Offer
~2 min readFor employees based in the Netherlands:
Location & Eligibility
Listing Details
- Posted
- October 5, 2026
- First seen
- October 5, 2026
- Last seen
- October 8, 2026
Posting Health
- Days active
- 2
- Repost count
- 1
- Trust Level
- 52%
- Scored at
- October 8, 2026
Signal breakdown
Similar Software Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.