Quick Summary
1) Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs2) Deploy, configure,
15 years full time educationKey Responsibilities:1) Design and implement HPC and AI infrastructure solutions,
Project Role Description : Creates production and non-production cloud environments using the proper software tools such as a platform for a project or product. Deploys the automation pipeline and automates environment creation and configuration.
Must have skills : Linux Architecture
Good to have skills : Machine Learning (ML), Cloud Technology Architecture, Docker Kubernetes Administration, Edge Computing, Microsoft Agentic AI Architect
Minimum 7.5 year(s) of experience is required
Educational Qualification : 15 years full time education
Key Responsibilities:
1) Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs
2) Deploy, configure, and manage XPU-based clusters (CPU/GPU/accelerators) using schedulers, VM/K8s orchestration platforms, Slurm, and containerized platforms in scalable designs to provide Metal as a Service (MaaS), GPUaaS, AIaaS, and other offerings
3) Optimize cluster performance, scalability, energy, and cost efficiency across on-premises, cloud, and hybrid environments
4) Integrate AI and HPC platforms with existing IT systems, data pipelines, and security frameworks
5) Monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency
6) Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks
7) Provide Provide technical guidance and support to users, enabling efficient execution of HPC/AI workloads, large-scale models, and simulations
Required Skills and Qualifications:
1) Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
2) Primary Skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro)
3) Secondary Skills: Parallel programming: MPI, OpenMP, CUDA, SYCL, NVIDIA DGX SuperPOD & BCM, InfiniBand NDR 400Gb/s & RoCEv2 design, MLOps & HPC integration (Kubeflow), Containerization: Singularity, Kubernetes, DevOps: Ansible, Terraform for HPC, UFM management, RunAI, Azure ML integration: distributed training, MLflow, Terraform / Bicep IaC for Azure HPC, FinOps: Reserved Instances, Spot VM strategies, Hybrid cloud HPC: on-prem to Azure/AWS burst
4) Proven hands-on experience designing, deploying, and managing HPC and AI infrastructure across on-premises, cloud, and hybrid environments in 2 or more segments: hyperscaler, neocloud, large Enterprise, Telco/Mobile, supporting key industries such as Financial Services, Life Sciences, Manufacturing, and Retail
5) Deep knowledge of accelerated computing architectures (GPUs, XPUs, DPUs), high-performance fabrics (InfiniBand, Ethernet), SONiC, networking, and modern storage/data platforms (e.g. NVMe-oF, Lustre, GPFS, BeeGFS, VAST, DDN, Weka) to build robust solutions
6) Proficiency with cluster management and orchestration (e.g. Slurm, Run:ai, Kubernetes, Docker), real-time performance monitoring, and observability frameworks
7) Hands-on experience with cloud and virtualization platforms (e.g. AWS, Azure, GCP, VMware, Nutanix) and expertise in automation and optimization using scripting (Python, AI tools) with foundational Infrastructure-as-Code tools such as Terraform and Ansible.
Preferred Skills and Qualifications:
1) Experience managing the deployment of 1,000+ GPU clusters for HPC and AI workloads with various infrastructure services enabled
2) Experience with GPU computing libraries and accelerators (e.g., NVIDIA CUDA, Dynamo, AMD ROCm).
3) Experience with AI and HPC Networking (e.g., RoCE, InfiniBand, muti-planar/multi-rail designs, platform buffer architectures)
4) Knowledge of Machine Learning and AI frameworks (e.g., TensorFlow, PyTorch, JAX), Jupyter notebooks / Google Colab environments
5) Familiarity with DevOps practices and tools (e.g., Ansible, Terraform) for infrastructure automation
6) Experience with AgenticAI and associated technologies to leverage and build agents for workflow automation and observability
7) Industry certifications in NVIDIA infrastructure, public cloud providers, Data Science, etc. are a plus
8) Strong problem-solving, troubleshooting, communication, and collaboration skills to deliver reliable, scalable, and high-performance infrastructure solutions in fast-paced, dynamic environments that reward technical talent
15 years full time education
Visit us at www.accenture.com
We believe that no one should be discriminated against because of their differences. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, military veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by applicable law. Our rich diversity makes us more innovative, more competitive, and more creative, which helps us better serve our clients and our communities.
Location & Eligibility
Listing Details
- Posted
- October 8, 2026
- First seen
- October 8, 2026
- Last seen
- October 8, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 55%
- Scored at
- October 8, 2026
Signal breakdown
4 other jobs at
View all →Similar Technology jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.