15h ago
New

Staff Engineer - Distributed Systems

IndiaIndiaRemoteFull-timelead
OtherEngineer
2 views0 saves0 applied

Quick Summary

Requirements Summary

10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances,

Technical Tools
OtherEngineer

This role focuses on the architecture and resilience of large-scale distributed systems operating under significant production load. You will look across services, queues, databases, infrastructure, and deployment patterns to identify systemic risks before they become incidents. The role begins with a high-throughput automation platform and expands across communication and emerging product systems. You will influence critical-path architecture while remaining hands-on, spending meaningful time prototyping solutions, resolving complex failures, and shipping remediation. You will help establish capacity models, consistency guarantees, isolation boundaries, and failure-handling strategies that can withstand continued growth. This is a high-ownership environment where technical rigor, proactive problem solving, and clear cross-team communication are central to the role.

  • Own the architecture health of large-scale distributed systems, including failure modes, capacity constraints, consistency guarantees, resilience, and interactions across numerous services and deployments.
  • Review and influence critical-path technical designs, providing rigorous architectural guidance and making well-reasoned recommendations when teams face complex technical trade-offs.
  • Proactively identify systemic risks such as single points of failure, unbounded queues, missing idempotency, thundering-herd effects, capacity constraints, and potential data-loss scenarios, then drive remediation before incidents occur.
  • Build and ship solutions for complex architectural problems, including prototypes, reliability improvements, production remediation, and fixes for difficult cross-team issues.
  • Design resilience into systems through degradation strategies, backpressure mechanisms, isolation boundaries, capacity models, and failure-handling patterns capable of supporting sustained growth.
  • Investigate and resolve the most challenging distributed-system failures by understanding interactions across application services, messaging, databases, infrastructure, and deployment environments.
  • Work hands-on with Node.js/TypeScript and/or Go to prototype architectural solutions and implement critical fixes when required.
  • Establish and improve engineering practices through design reviews, post-mortems, architectural patterns, documentation, and technical guidance that can be adopted across multiple teams.
  • Help define how AI-assisted engineering can be used safely and effectively when developing and operating highly critical distributed systems.
  • Work with technologies including GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch across large-scale production environments.
  • Help engineering teams prepare systems for substantial traffic growth by identifying capacity limits, improving observability, and validating critical failure modes.
  • Influence engineers across teams through technical rigor, collaboration, mentorship, and clear communication rather than relying solely on organizational authority.

Requirements

~2 min read
  • 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances, or systems processing billions of events is strongly preferred.
  • Demonstrated experience carrying direct accountability for production systems through significant incidents, migrations, reliability challenges, and operational failures.
  • Deep expertise in queueing and asynchronous architectures, including delivery semantics, ordering, backpressure, idempotency, and the practical limitations of exactly-once processing.
  • Strong knowledge of multiple storage technologies across SQL and NoSQL environments, with the ability to reason about consistency models, indexing at scale, performance, and appropriate technology selection.
  • Expert-level experience with Redis or comparable in-memory systems, including behavior under memory pressure, network partitions, failover scenarios, and other failure conditions.
  • Strong production experience operating Kubernetes at scale, including resource limits, autoscaling, capacity planning, node failures, and workload behavior under infrastructure disruption.
  • Fluent in Node.js and/or Go, with sufficient hands-on ability to prototype technical proposals and implement production fixes on critical paths.
  • Exceptional technical communication skills, including the ability to produce design documents, architectural diagrams, and root-cause analyses that drive decisions across multiple engineering teams.
  • Strong systems-thinking ability, with an instinct for evaluating tail behavior, failure modes, network partitions, capacity constraints, and high-load scenarios.
  • Proven ability to influence engineering teams through technical depth, constructive disagreement, clear reasoning, and respect.
  • Strong ownership, curiosity, judgment, and problem-solving skills, particularly when dealing with ambiguous or cross-functional technical challenges.
  • Experience using AI-assisted development tools effectively on complex systems is an advantage, particularly when balancing development speed with code quality, reliability, and operational safety.
  • GCP-native experience with technologies such as Pub/Sub, Cloud Tasks, GKE, and Firestore is a strong advantage.
  • Previous experience in a Staff, Principal, Architect, or comparable systems-focused engineering capacity is beneficial.
  • Experience taking new systems from initial architecture through production while also hardening mature systems is a plus.

What We Offer

~2 min read
✓Full-time remote opportunity for engineers based in India.
✓Opportunity to work on distributed systems operating at substantial scale, including billions of automation actions and messages and tens of thousands of requests per second at peak.
✓Broad technical scope spanning application runtimes, messaging, asynchronous processing, databases, caching, Kubernetes, cloud infrastructure, and observability.
✓Opportunity to influence architecture across multiple engineering teams and critical production systems.
✓Significant hands-on ownership, with the ability to prototype, build, ship, and remediate solutions rather than working solely in an advisory architecture capacity.
✓Exposure to large-scale technologies including Node.js, TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch.
✓Opportunity to develop resilience, capacity planning, consistency, and failure-management strategies for systems experiencing continued growth.
✓Collaborative environment where technical rigor, clear communication, proactive problem solving, and engineering mentorship are valued.
✓Opportunity to establish engineering patterns and practices that influence a broad engineering organization.
✓Opportunity to work with AI-assisted engineering tools and help define safe, high-quality practices for their use in critical production systems.

Location & Eligibility

Where is the job
India
Remote within one country
Who can apply
IN

Listing Details

Posted
September 28, 2026
First seen
September 28, 2026
Last seen
September 28, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
September 28, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Staff Engineer - Distributed Systems