Founding Support Engineer
Quick Summary
About TrueFoundry Every production AI system whether it's powering customer support, writing code, analyzing financial data,
Every production AI system whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.
We are looking for a Founding Support Engineer to join the team.
Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.
The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.
You need a control plane that handles:
- Intelligent routing with observability, cost policies, and fallback logic
- Centralized tool and MCP server management with security and lifecycle controls
- Agent orchestration with governance and guardrails
- A unified compute layer to run self-hosted models, custom tools, and agents
About the Role
~1 min read- Be the first and most trusted line of response when enterprise customers hit issues in production, across severities from routine questions to P0 incidents.
- Get hands-on with real infrastructure problems, Gateway configuration, MCP proxy scaling, RBAC and SSO, guardrails, sizing, not surface-level troubleshooting.
- Own tickets to actual resolution, not just closure, and know the difference.
- Triage and file issues to engineering with clean, reproducible detail when a fix needs to go deeper than support.
- Build and contribute to the internal knowledge base and runbooks, so the next ticket like this one gets solved faster.
- Work inside the escalation and priority process, and hold the line on incident response coverage.
- Responsiveness and resolution: first response time against SLA tiered by severity (P0 through P3), time to resolution by severity, and keeping the backlog of aging tickets under control.
- Quality of resolution: high first contact resolution rate, low reopen rate, strong CSAT on resolved tickets, and an escalation rate that reflects real complexity rather than gaps in troubleshooting.
- Technical depth: a track record of finding actual root cause on infra issues rather than patching symptoms, clean and reproducible bug reports handed to engineering, and real contributions to the knowledge base and runbooks.
- Team and process health: consistent adherence to the escalation and tiering process, reliable on-call and incident coverage, and documentation that holds up under peer review.
- Someone who can read logs, reason about distributed systems, and stay calm inside a live incident. Comfortable with Kubernetes and at least one major cloud.
- Some exposure to LLMs, agents, or ML pipelines in production is a strong plus, you will pick up the rest fast if the fundamentals are there.
- Clear written communication matters as much as technical chops, since a lot of trust is built through how you write up a ticket.
Traits we are looking for: Ownership, ability to execute, hustle and think out of the box, data driven decision making, be comfortable with more unknowns than knowns.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- July 29, 2026
- First seen
- July 29, 2026
- Last seen
- July 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 52%
- Scored at
- July 29, 2026
Signal breakdown
Please let truefoundry know you found this job on Jobera.
3 other jobs at truefoundry
View all →Explore open roles at truefoundry.
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.