Senior SDE - Platform
Quick Summary
4+ years of backend engineering experience with deep, hands-on expertise in Node.js and/or Go. Strong systems understanding across the application, database, and infrastructure layers,
This role focuses on building and operating the platform infrastructure that powers large-scale CRM, automation, and communication systems. You will own critical components such as queues, storage layers, caching, schedulers, rate limiters, and infrastructure services. The environment handles billions of events and messages, thousands of deployments, and high-throughput production traffic. You will work across application code, databases, cloud infrastructure, and distributed systems rather than focusing on a single technical layer. The role requires strong ownership, from architecture and design through deployment, observability, and production health. You will also help identify failure modes and scalability gaps before they impact customers and engineering teams.
- Design and build core platform components supporting high-volume workflows and communication systems, including queueing pipelines, storage access layers, caching strategies, rate limiters, and schedulers.
- Own platform components end to end, covering architecture, technical design documents, implementation, testing, rollout, dashboards, monitoring, and production reliability.
- Work deeply with large-scale databases, including schema and index design, sharding, replication behavior, query optimization, and safe migrations across systems containing billions of records.
- Partner closely with infrastructure teams and work hands-on with Kubernetes/GKE autoscaling, resource tuning, messaging infrastructure, capacity planning, and cloud-based deployment architectures.
- Proactively identify reliability and scalability risks such as missing idempotency, unbounded queues, cache stampedes, hot shards, resource constraints, and other failure modes, then drive improvements.
- Investigate and resolve production incidents that span application, database, messaging, and infrastructure layers, identifying root causes across interconnected systems.
- Design and operate systems using Node.js/TypeScript and Go, alongside technologies such as GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch.
- Strengthen observability through metrics, tracing, alerting, dashboards, and other mechanisms that improve system visibility and operational readiness.
- Write clear design documents and rigorous root-cause analyses that communicate technical behavior, trade-offs, risks, and recommendations to engineering teams.
- Make effective use of AI-assisted engineering tools to accelerate development while maintaining high standards for accuracy, testing, code quality, and production reliability.
Requirements
~2 min read- 4+ years of backend engineering experience with deep, hands-on expertise in Node.js and/or Go.
- Strong systems understanding across the application, database, and infrastructure layers, ideally demonstrated through experience operating high-throughput production systems.
- Deep experience with queueing and asynchronous processing technologies such as GCP Pub/Sub, Kafka, RabbitMQ, Cloud Tasks, or similar systems, including delivery semantics, ordering, backpressure, and idempotency.
- Strong database expertise beyond basic CRUD operations, including indexing, sharding, replication, query performance, and safe migrations across SQL or NoSQL databases at scale.
- Hands-on experience with Redis or comparable in-memory data stores, including caching patterns, data structures, and failure modes.
- Production experience with cloud infrastructure and Kubernetes, including autoscaling, resource limits, capacity planning, and system behavior under heavy load; GCP experience is preferred.
- Strong ability to understand distributed systems in depth, including behavior under load, tail latency, network failures, and other real-world failure conditions.
- Experience troubleshooting complex production incidents and improving system designs based on lessons from operational failures.
- Ability to write clear, practical design documents that explain trade-offs, risks, and technical recommendations.
- Strong engineering judgment, curiosity, ownership, problem-solving ability, and communication skills.
- Demonstrated ability to use AI-assisted development tools effectively to produce accurate, tested, production-quality code.
- Experience operating systems at comparable scale, including billions of events or thousands of instances, is a strong advantage.
- Production experience with ClickHouse, Elasticsearch, or Firestore is considered a bonus, as is previous experience building platform infrastructure.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 28, 2026
- First seen
- September 28, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 28, 2026
Signal breakdown
Similar Platform jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.