AI Platform Engineer
Quick Summary
Job Description The Position The Advanced Scientific Compute (ASC) team is evolving. As we shift our focus toward hosting proprietary models and managing state-of-the-art,
Job Description
The Advanced Scientific Compute (ASC) team is evolving. As we shift our focus toward hosting proprietary models and managing state-of-the-art, on-premise GPU hardware for both AI model training and inference, we are looking for a motivated AI Platform Engineer.
In the modern AI landscape, we are looking for an engineer with a strong technical foundation and good understanding of AI concepts, who can also act as an orchestrator. You will help operate and support the foundational layer of our on-premise AI hosting environments—encompassing GPU/HPC infrastructure, containerization, automation, and observability. You will leverage AI agents to augment your technical skills, accelerating your workflows and allowing you to focus on problem-solving and collaboration with our research community across the US and EU.
- Manage and automate on-premise infrastructure for AI and HPC, focusing on high-performance GPU servers and clusters, while occasionally assisting with cloud-based resources.
- Deploy, manage, and orchestrate AI training and inference workloads, primarily utilizing Slurm, Nvidia Triton and Docker containers.
- Leverage AI and GenAI tooling to accelerate development, build CI/CD pipelines, and automate repetitive infrastructure tasks.
- Assist in performance tuning, capacity tracking, and scaling of compute clusters to ensure reliable and efficient services for researchers.
- Collaborate directly with researchers to understand workflows and translate non-technical needs into robust infrastructure solutions.
- Maintain monitoring, dashboards, logging, and alerting for critical components to support vital AI infrastructure.
- Maintain and evolve the existing ecosystem of HPC modules and scientific applications to ensure reproducible builds and consistent environments.
- Share expertise in leveraging agentic AI tools and champion modern, AI-assisted engineering practices among team members.
Requirements
~1 min read- 3+ years of software or infrastructure engineering experience, with at least 1 year of hands-on experience working with AI platforms, GPU environments, or similar modern compute workloads.
- Practical experience in Enterprise Linux administration, automated testing, CI/CD pipelines, and Infrastructure as Code tools (e.g., Terraform, Ansible).
- Strong experience with Docker and container runtimes, alongside familiarity with job schedulers like Slurm in an HPC or scientific computing environment.
- Practical experience with GPU servers or clusters, including a solid understanding of NVIDIA drivers, CUDA/cuDNN/NCCL, and the NVIDIA Container Toolkit.
- Solid Python programming skills and strong DevOps fundamentals, particularly utilizing version control (Git).
- Experience with monitoring and troubleshooting distributed systems or compute clusters.
- Strong communication skills and the ability to work flexibly to collaborate across EU and US timezones.
Nice to Have
~1 min read- Experience with Kubernetes and operating containerized platforms at scale.
- Background in pharma, biotech, or academic scientific computing environments.
- Full-stack development capabilities (web UIs, REST APIs) to build internal researcher-facing tooling.
- Familiarity with Python packaging and environment management tools (Conda/mamba, pip) alongside the scientific Python stack.
Current Employees apply HERE
Current Contingent Workers apply HERE
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- September 29, 2026
- First seen
- October 1, 2026
- Last seen
- October 1, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 51%
- Scored at
- October 2, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.