Quick Summary
The High-Performance Computing Storage Engineer is primarily responsible for the overall health and maintenance of storage technologies in our managed services customer's environments.
-
Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities
-
Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast
-
Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference
-
Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations
-
Plan and perform maintenance activities
-
Assess customer environments for performance and design issues and propose resolutions
-
Work across technical teams to troubleshoot complex infrastructure issues
-
Create and maintain detailed documentation
-
Serve as a subject matter expert and escalation point for storage technologies
-
Work with vendors to resolve storage issues
-
Communicate with customers and internal team with transparency
-
Support data movement workflows including ingest, replication, caching, tiering, and archiving
-
Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics
-
Partner with infrastructure, platform, and research teams to support production AI/HPC workloads
-
Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency
-
Communicate with customers and internal team with transparency
-
Participate in on-call rotation
-
5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering
-
Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education
-
Strong experience with Linux systems administration
-
Hands-on experience configuring, managing, and tuning distributed or parallel filesystems
-
Experience tuning storage for performance-sensitive workloads
-
Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes
-
Familiarity with high-speed interconnects such as InfiniBand or RDMA
-
Ability to troubleshoot complex issues across storage, compute, and networking layers
-
Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments
-
Experience with machine learning or data science workflows in HPC environments
-
Managed Services or consulting experience
-
Strong background with customer service
-
High level problem-solving and communication skills
-
Strong oral and written communications skills
-
Managed Services or consulting experience
-
Experience supporting storage solutions for GPU clusters and AI/ML workflows
-
Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway
-
Experience with Terraform, Ansible, Helm, or GitOps workflows
-
Knowledge of observability platforms such as Prometheus and Grafana
-
Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems
-
Experience with machine learning or data science workflows in HPC environments
-
Scripting or programming experience with Python and Bash
-
Related Storage certifications are a bonus
Location & Eligibility
Listing Details
- Posted
- August 21, 2026
- First seen
- August 21, 2026
- Last seen
- August 21, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 70%
- Scored at
- August 21, 2026
Signal breakdown
Please let Thinkahead know you found this job on Jobera.
3 other jobs at Thinkahead
View all →Explore open roles at Thinkahead.
Similar Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.