AI Infrastructure Architect
Quick Summary
- Design evaluation strategies for agent behavior, task completion, retrieval quality, groundedness, safety, and reliability.- Build automated evaluation harnesses, regression suites,
15 years full time educationSummary: Build the evaluation and quality engineering capability for AI and agentic systems. Establish measurable quality gates that determine whether systems are safe,
Project Role Description : Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance, power consumption, cost and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation, selection and full stack integration.
Must have skills : Large Language Models (LLMs)
Good to have skills : NA
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
Build the evaluation and quality engineering capability for AI and agentic systems. Establish measurable quality gates that determine whether systems are safe, reliable, effective, and ready for production.
Expert use of LangSmith, Braintrust, Arize Phoenix, Weights & Biases Weave, MLflow and OpenAI Evals, supported by Codex, Claude Code or Cursor, to build automated regression suites, trajectory evaluations, red-team tests, trace analysis and CI/CD quality gates for production agents.
Must have built evaluation or quality systems for production AI, ML, search, or decision systems. Manual prompt testing and subjective review alone are insufficient.
Roles & Responsibilities:
- Design evaluation strategies for agent behavior, task completion, retrieval quality, groundedness, safety, and reliability.
- Build automated evaluation harnesses, regression suites, benchmark datasets, and production quality gates.
- Evaluate multi-step agent trajectories, tool use, planning, recovery, and human escalation.
- Combine deterministic tests, model-based evaluation, human review, and production telemetry.
- Perform failure analysis, red teaming, and root-cause investigation.
- Integrate evaluations into CI/CD, release, monitoring, and incident-management processes.
- Define scorecards for engineering, risk, product, and business stakeholders.
Professional & Technical Skills:
- Python, test automation, data analysis, statistics, and experimentation.
- LLM and RAG evaluation, agent trajectory analysis, benchmark design, and error taxonomy.
- Tracing, observability, adversarial testing, safety testing, and production monitoring.
- Distinguishing model, retrieval, prompt, tool, data, and orchestration failures.15 years full time education
Visit us at www.accenture.com
We believe that no one should be discriminated against because of their differences. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, military veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by applicable law. Our rich diversity makes us more innovative, more competitive, and more creative, which helps us better serve our clients and our communities.
Location & Eligibility
Listing Details
- Posted
- September 22, 2026
- First seen
- September 29, 2026
- Last seen
- October 2, 2026
Posting Health
- Days active
- 3
- Repost count
- 0
- Trust Level
- 32%
- Scored at
- October 3, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.