Staff Engineer, AI Operations & Governance, Workplace AI
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer, AI Operations & Governance,
This is a deeply technical individual contributor role focused on the operational reliability, security, and governance of workplace AI systems.
You will own critical technical operations across AI platforms, integrations, observability, access controls, and production changes.
The role combines hands-on engineering with AI governance, risk management, documentation, and enablement.
You will investigate incidents, troubleshoot complex production issues, and implement mitigations or fixes to restore services quickly.
The position also contributes to model evaluations, benchmarking, operational playbooks, and technical training.
You will work closely with Security, IT, Legal, HR, and business stakeholders to translate technical signals into actionable risk and reliability insights.
This is a high-ownership environment suited to an experienced engineer comfortable operating close to production and navigating evolving AI technologies.
- Manage day-to-day configurations for AI platforms and orchestration tools, including models, routing, guardrails, tenants, policies, role mappings, and prompt libraries.
- Design, configure, and maintain connectors and integrations with SaaS applications, data sources, and workflow tools while ensuring secure and reliable data access.
- Build and maintain observability across AI workflows, including logging, metrics, alerts, latency, errors, usage patterns, and potential drift indicators.
- Implement secure secrets management and access controls, including API key and credential rotation, permissions, and least-privilege practices in collaboration with Security and IT.
- Execute production changes such as model replacements, policy updates, prompt modifications, version upgrades, and controlled rollouts while maintaining change records and rollback procedures.
- Conduct technical evaluations and benchmarking of AI models, tools, and configurations, summarizing results to support governance and roadmap decisions.
- Translate technical logs, metrics, incidents, and operational findings into clear risk, reliability, compliance, and governance perspectives for technical and business stakeholders.
- Contribute to operational playbooks, runbooks, documentation, and training materials, and occasionally support user training or technical office hours.
- Serve as a technical incident responder by triaging failures, investigating root causes, proposing and implementing mitigations, documenting lessons learned, and communicating clearly with non-technical stakeholders.
Requirements
~2 min read- 4–7+ years of experience in platform engineering, DevOps/SRE, ML/AI operations, technical SaaS operations, or a similar discipline involving hands-on production systems.
- Strong knowledge of APIs, integrations, infrastructure-as-code concepts, and configuration environments using code, JSON, YAML, and automation.
- Hands-on experience with monitoring and observability tools, including logs, metrics, and alerting systems, with the ability to use operational data to diagnose and improve systems.
- Practical experience with at least one AI or automation platform, such as LLM providers, AI productivity platforms, RPA solutions, workflow engines, or similar technologies.
- Experience with secure secrets management, access controls, role-based permissions, key rotation, and least-privilege security practices.
- Strong technical documentation skills, with the ability to communicate engineering decisions and operational information to both technical and non-technical audiences.
- Strong ownership, initiative, problem-solving ability, and comfort working directly with production systems in a high-stakes environment.
- Experience in ML/AI operations, including model deployment, evaluation, or drift monitoring, is preferred.
- Familiarity with prompt engineering, LLM policies, guardrails, and configuration patterns is an advantage.
- Exposure to AI governance, compliance, or risk frameworks for data-driven and AI systems is beneficial.
- Experience collaborating with Security, Legal, and business stakeholders on technical risks and mitigation strategies is preferred.
- Prior experience with incident response, change management, or on-call rotations for critical systems is a plus.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 24, 2026
- First seen
- September 27, 2026
- Last seen
- September 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 27, 2026
Signal breakdown
Similar Staff Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.