7h ago
New

Senior Lead AI Systems and Performance Architect – LLM Inference

United StatesUnited States·Prescottsenior
ArchitectConstruction & Real Estate
2 views0 saves0 applied

Quick Summary

Key Responsibilities

1.

Requirements Summary

Qualifications: Minimum Qualifications: • Education: Ph.D. or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related quantitative field. • Experience: Ph.D.

Technical Tools
ArchitectConstruction & Real Estate

About the Role and Team The Intel DCG AISOC organization is developing the future of high-performance accelerated AI platforms. Real-world inference performance, latency SLOs, and TCO are governed by full-system interplay: dynamic batching, distributed KV cache movement, scale-up fabrics, and heterogeneous host-accelerator memory hierarchies. We are seeking a hands-on, technically deep Senior Lead AI Systems and Performance Architect to drive the technical direction, development, and system-level analysis of our End-to-End LLM Inference Simulation Platform. In this role, you will lead the technical execution of the simulation platform, define modeling workflows, and conduct dynamic trace replays. Your mission is to identify full-system bottlenecks, quantify GPU optimization, evaluate host CPU offloading, size memory subsystems, and assess high-speed interconnects -translating simulation findings into actionable engineering insights to guide Intel's future AI inference SoC and platform architectures. ________________________________________ Key Responsibilities: 1. Simulation Platform and Workflow Leadership • Technical Ownership: Lead the technical direction and continuous evolution of a discrete-event simulation platform modeling distributed multi-accelerator inference systems. • Serving Stack Modeling: Model critical serving mechanisms: e.g.continuous batching, chunked prefill, prefill-decode disaggregation, and dynamic KV cache allocation. • Fabric and Pipeline Abstraction: Integrate behavioral models for scale-up interconnects (UALink, PCIe Gen 6/7), capturing latency, arbitration, and collective communications (All-to-All, All-Reduce). • Workflow Automation: Design automated simulation pipelines to ingest empirical accelerator performance tables, execute parameter sweeps, and analyze results. 2. Dynamic Trace Replay and Bottleneck Analysis • Production Trace Replay: Drive the ingestion and replay of realistic production traces, capturing arrival burstiness, diverse prompt/output distributions, and shared prefix contexts. • Agentic Workload Modeling: Keep the simulator aligned with emerging patterns, specifically Agentic workflows (multi-step tool calls, state branches) and long-context KV cache demands. • Root-Cause Isolation: Conduct sensitivity analyses to isolate bottlenecks across compute throughput, memory bandwidth, bus latency, scale-up fabric contention, or scheduling overheads. • Metric Evaluation: Evaluate trade-offs across SLOs (TTFT, TPOT, ITL, queue wait times) against cluster throughput, power, and cost. 3. HW/SW Co-Design and Platform Recommendations • Silicon Architecture Input: Translate findings into architectural proposals for the AI SoC team, defining balanced compute-to-memory ratios, SRAM sizes, and interface configurations. • Platform Optimization: Evaluate the host CPU's role in the inference pipeline to identify offloading opportunities (orchestration, tokenization, routing, tiered KV cache staging). • Memory Hierarchy Sizing: Provide workload-grounded capacity and bandwidth recommendations, evaluating HBM/LPDDR vs. tiered system memory. 4. Team Leadership and Collaboration • Mentorship and Quality: Guide team members on simulation methodologies, code quality, model calibration, and experimental rigor. • Cross-Team Alignment: Partner with silicon architects, software engineers, and platform teams to translate architectural questions into structured simulation studies.

Requirements

~1 min read
Qualifications: Minimum Qualifications: • Education: Ph.D. or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related quantitative field. • Experience: Ph.D. with 4-6+ years (or Master's with 6-8+ years, or Bachelor's with 8-10+ years) in system performance modeling, computer architecture, or AI systems engineering. • Simulation and Modeling: Hands-on experience developing or utilizing discrete-event, trace-driven, or analytical performance simulators. • Performance Analysis: Proven track record of identifying performance bottlenecks across compute, memory, and interconnects in multi-accelerator/heterogeneous environments. • Programming: Proficiency in Python and C++ (14/17/20), with experience in modular software design and automated data processing. Preferred Qualifications: • AI Serving Frameworks: Understanding of LLM serving runtimes (e.g., vLLM, TensorRT-LLM, SGLang) and mechanisms (PagedAttention, continuous batching, speculative decoding). • Interconnect and Fabrics: Knowledge of accelerator interconnects, particularly UALink, PCIe Gen 6/7, CXL, or high-performance networking (RoCEv2, Ultra Ethernet). • Heterogeneous Co-Design: Experience evaluating host-accelerator partitioning, CPU-assisted pipelines, or tiered memory architectures. • Workload Characterization: Experience profiling and replaying production trace telemetry (multi-turn conversations, MoE routing, Agentic workflows). • Technical Leadership: Track record of leading technical projects, driving team consensus, and mentoring engineers.

          

College Grad

Shift 1 (China)

PRC, Shanghai

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.
N/A

This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.

Location & Eligibility

Where is the job
Prescott, United States
On-site at the office
Who can apply
US

Listing Details

Posted
October 8, 2026
First seen
October 8, 2026
Last seen
October 8, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
56%
Scored at
October 8, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Senior Lead AI Systems and Performance Architect – LLM Inference