Staff Engineer (Core & MLOps)
Quick Summary
10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands-on architecture with technical leadership, mentoring, and cross-functional alignment. In a globally distributed, remote-first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long-term platform strategy.
- Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
- Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
- Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
- Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
- Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
- Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
- Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
- Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements
~2 min read- 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
- Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
- Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
- Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
- Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
- Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
- Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
- Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
- A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
- Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
- MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
- Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
- Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
- Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 1, 2026
- First seen
- October 1, 2026
- Last seen
- October 1, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 62%
- Scored at
- October 1, 2026
Signal breakdown
Similar Staff Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.