1mo ago

Forward Deployed Engineer

United StatesUnited States·Palo Altofull-timemid
OtherForward Deployed Engineer
0 views0 saves0 applied

Quick Summary

Overview

About DeepInfra DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale.

Technical Tools
OtherForward Deployed Engineer

DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale. Our team has deep experience building large systems that serve hundreds of millions of users, and we're bringing that same level of rigor to a rapidly evolving AI inference space. Our mission is to make advanced AI available to people and teams everywhere.

We're an early, tight-knit team where you can influence product direction, try bold ideas, and drive meaningful work forward quickly. If you want to join a fast-growing company at a defining moment, we'd love to talk.

DeepInfra is backed by leading investors including A.Capital, Felicis, 500 Global, Georges Harik, Samsung Next, Supermicro, Upper90, Peak6, SVAngel and Nvidia.


As DeepInfra's enterprise pipeline grows, our customers need a technical partner who can run rigorous evals, defend benchmarks, and speak fluently to both engineering and procurement — someone who can own the technical win from first call through production.

This is a pioneering role. You'll work closely with Sales, our co-founders, and the engineering team on the deals that matter most. You'll own the technical win end to end: running head-to-head bake-offs against leading AI providers, tuning deployments on the latest hardware, and turning what you learn into reusable assets that make every future deal faster to close. As our first FDE, you'll also define what the function looks like as GTM scales.

Responsibilities

~1 min read

  • →Own the technical win and the POC timeline, working closely with Sales and Engineering, from call one.
  • →Design and run reproducible benchmark harnesses (TTFT, ITL, throughput/GPU, p95/p99) and quality-parity evals.
  • →Run head-to-head bake-offs against leading AI providers — and win them.
  • →Tune model-to-hardware deployments on B200/B300/GB300 NVL72.
  • →Build cost-per-token models and write migration plans.
  • →Handle enterprise security and compliance review, and get deployments to launch readiness.
  • →Own account health post-signature, driving usage reviews and expansion.
  • →Turn what you learn into reusable benchmark reports, reference architectures, and AE enablement material.


  • Customer-facing engineering with an owned technical outcome at an infrastructure or ML platform company.
  • Strong Python skills.
  • Dual-audience presence with commercial instinct — credible with a skeptical staff engineer, clear with a CFO, and able to tell a technical objection from a procurement one.

Nice to Have

~1 min read

  • Hands-on experience with inference internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math, quantization.
  • Experience with agentic or coding-assistant workloads at scale.
  • Prefix-cache-heavy long context workloads.
  • Diffusion image/video, ASR/TTS, or multi-LoRA serving.
  • Open-source contributions to vLLM or SGLang.
  • Deep NVLink/InfiniBand topology knowledge.


  • Define DeepInfra's Forward Deployed Engineering function from day one and have a direct impact on its direction.
  • Work directly with co-founders and the inference team on the deals that matter most.
  • Join a small, high-performing team where your work ships quickly and reaches customers around the world.
  • Help shape how enterprises adopt some of the world's leading open-source AI models.


Three traits define the people who thrive here, and this role leans on all three.

Initiative. We take ownership and step in where we can add value. Whether it’s starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.

Drive. We’re energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving — because solving meaningful challenges is what motivates us.

Grit. Things don’t always work on the first try — and that’s expected. We stay persistent, adapt quickly, and learn as we go. We take setbacks seriously, but not personally, and use them to get better.

What We Offer

~1 min read
The base pay range for this role is $150,000 – $195,000 per year.

Location & Eligibility

Where is the job
Palo Alto, United States
On-site at the office
Who can apply
US

Listing Details

Posted
August 14, 2026
First seen
September 26, 2026
Last seen
October 1, 2026

Posting Health

Days active
4
Repost count
0
Trust Level
20%
Scored at
October 1, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Forward Deployed Engineer