Member of Technical Staff, Inference Systems
Quick Summary
Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
attention, KV cache, batching, scheduling, and where the real bottlenecks are. Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-ord
Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints.
Recently founded · Small founding team · Industry: AI infrastructure / LLM inference
You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice.
Responsibilities
~1 min read- →Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
- →Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token.
- →Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly.
- →Profile, benchmark, and ship performance improvements across the entire inference pipeline.
- →Work closely with the founding team on the core architectural decisions that define the platform.
Tech stack: Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
Requirements
~1 min read- 2–10 years of experience as a backend or distributed systems engineer.
- Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs.
- A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are.
- Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns.
- Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch). If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining.
Nice to Have
~1 min read- Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo.
- A CS or systems degree from a strong program.
- Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation.
What We Offer
~1 min read- Location: Palo Alto, CA
- Work policy: Full-time, on-site five days a week
- Compensation: $230K–$350K + equity
- Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new H-1B, TN)
- Employment type: Full-time
Location & Eligibility
Listing Details
- First seen
- September 25, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 2
- Repost count
- 0
- Trust Level
- 56%
- Scored at
- September 28, 2026
Signal breakdown
Similar Member Of Technical Staff jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.