Staff Software Engineer, Infrastructure
Quick Summary
8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering. Bachelor’s degree in Computer Science, Engineering,
This is a Staff-level infrastructure engineering role focused on building the platform foundations that enable engineering teams to move faster and operate more reliably.
You’ll combine hands-on software engineering with technical leadership, shaping self-service infrastructure, cloud platforms, deployment systems, and operational tooling.
A major focus will be replacing expert-driven provisioning and operational workflows with paved roads that provide clear ownership, safe defaults, strong guardrails, and measurable adoption.
You’ll help evolve multi-region, cross-account infrastructure and Kubernetes foundations while improving reliability, security, scalability, and cost efficiency.
The role also includes developing AI-assisted operational workflows that reduce toil while keeping production changes safe, auditable, and human-reviewed.
You’ll influence architecture across teams through RFCs, design reviews, technical standards, and pragmatic engineering decisions.
Working in a small, growing team, you’ll have substantial ownership and the opportunity to drive infrastructure investments through real production adoption.
- Turn ambiguous infrastructure challenges into clear technical proposals and drive them through RFCs, architecture reviews, and cross-team alignment.
- Design and build self-service platform capabilities and APIs, primarily in Go, covering onboarding, provisioning, deployment, observability defaults, and day-2 operations.
- Establish reliable delivery standards using Terraform, GitOps with Argo CD, progressive delivery, automated testing, and continuous deployment practices.
- Evolve multi-tenant EKS infrastructure to improve reliability, security, scalability, and cost efficiency.
- Develop and improve ingress and traffic-routing capabilities, including Envoy Gateway and multi-region, cross-account connectivity.
- Strengthen SLOs, alerting, incident response, and operational follow-up through Grafana Cloud and improved observability practices.
- Measure success through outcomes for consuming engineering teams, including faster provisioning and deployment, greater self-service, and improved operational reliability.
- Develop AI-assisted operational workflows such as alert enrichment, incident context gathering, runbook-assisted diagnosis, remediation recommendations, and onboarding assistants.
- Maintain appropriate human oversight for AI-assisted operational actions, with an emphasis on safety, auditability, and responsible automation.
- Participate in the on-call rotation after onboarding and shadowing, while helping improve the overall health of on-call through better alerts, runbooks, automation, and blameless postmortems.
- Lead strategic platform initiatives from initial design through production adoption and establish durable technical patterns across engineering teams.
Requirements
~2 min read- 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong software engineering expertise in Go or a comparable programming language, including system design, testing, debugging, code review, and long-term maintainability.
- Proven experience designing, delivering, and operating cloud services or infrastructure platforms in production.
- Deep expertise in at least one area such as Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms.
- Strong Linux, networking, and production operations fundamentals.
- Experience setting technical direction and leading initiatives that require alignment across multiple engineering teams.
- Strong written and verbal communication skills, particularly in remote environments and through RFCs, design documents, and incident writeups.
- Experience with EKS, ingress, CNI, or service-mesh technologies is valuable.
- Familiarity with OpenTelemetry, Prometheus, Grafana, CI/CD, progressive delivery, GitHub Actions, Argo CD, and canary deployments is a plus.
- Experience leading large-scale migrations, platform adoption programs, or cross-team infrastructure initiatives is beneficial.
- Strong systems judgment, curiosity, pragmatic decision-making, and the ability to develop deep expertise while navigating adjacent technical domains.
- Willingness to participate in an operational on-call rotation and contribute to improving its effectiveness.
- Visa sponsorship may be considered on a case-by-case basis depending on business needs.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 29, 2026
- First seen
- September 29, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 29, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.