asrcfh1d ago
New
New
Senior Systems Administrator
senior
Systems AdministratorInfrastructure & Cloud
0 views0 saves0 applied
Quick Summary
Key Responsibilities
Maintains HPC clusters with over 2000+ nodes with InfiniBand, 100+ petabytes of data storage in production. Designs and develops scripts for system administration, monitoring and usage reporting.
Requirements Summary
The successful candidate will be an active supporting member of the ASRC Federal team reporting directly to the Manager of the HPC Computer Systems and Storage (CSS) group.
Technical Tools
Systems AdministratorInfrastructure & Cloud
ASRC Federal is searching for a Staff HPC Engineer to support InuTeq LLC out of NASA AMES, CA
ASRC Federal, InuTeq proudly supports NASA's High-Performance Computing Services program with our site in Mountain View, CA at the Ames Research Center. Make a DIFFERENCE on a program that supports 4 On-site Supercomputers totaling ~9,000 nodes and 25+ combined petaflops. This program provides High-Performance Computing services throughout the HPC lifecycle for computational requirements, architecture, acquisition, to our NASA customer. Our employees embrace innovation and are committed to a culture of continuous, standards-driven process improvement, and assimilation of industry best practices.
Summary: The successful candidate will be an active supporting member of the ASRC Federal team reporting directly to the Manager of the HPC Computer Systems and Storage (CSS) group.
An individual at this skill level should have demonstrated Linux administration experience as well as working knowledge of HPC operations while contributing to the support of users of HPC resources on the various issues they might have getting jobs to run efficiently. This individual will be expected to participate in all aspects of HPC System Administration that includes such activities as installing, maintaining, and upgrading HPC systems. The individual, along with the entire HPC team, will be engaged in the day-to-day operations and support of the HPC resources. Activities may include system patching, OS upgrades, deploying new systems, writing scripts, and troubleshooting system issues on the HPC system. The ability to interact with users to determine symptoms and then reproduce their issues to isolate the causes is critical skills for this work. There may also be activities supporting testing, benchmarking, tool scripting, and analyzing trouble tickets to find patterns indicating system or user education issues.
Duties and Responsibilities:
Maintains HPC clusters with over 2000+ nodes with InfiniBand, 100+ petabytes of data storage in production.
Designs and develops scripts for system administration, monitoring and usage reporting.
Modify existing software to correct errors and/or improve performance
Designs and develops scripts for system regression test and performance (file systems (Lustre), scheduler (PBS), interconnect (HDR/NDR, Slingshot, ), high availability, etc.).
Troubleshoots, isolates and resolves application, system and other technical problems (hardware, software, and network).
Understands research use cases, researches and deploys new technologies, defining cost, performance and other trade-offs.
Manages and maintains tools for configuration management (HPCM, Ansible & GIT), resource management, scheduling and all necessary aspects of HPC in accordance with best practices.
Researches, deploys and manages networking and security infrastructure, including development of policies and procedures.
Assists in developing and writing proposals and publications.
Creates and provides clear documentation.
After hours/weekend support as required
Moderate Supercomputing System Administration that contributes to:
Day-to-day operations of the Linux HPC clusters and storage systems
Proactive monitoring, analyzing, and correct system issues
Development of scripts to automate repetitive tasks or tools to enhance support of the HPC systems
System performance analysis and tuning
Building, installing, and supporting user-requested software
Supporting evaluation and assessment of new HPC technology
Resolving user report issues and managing support tickets requests in Remedy
Requirements:
Bachelor’s degree in computer science or related field
Strong computer science background with in-depth systems-level knowledge in operating systems and networking
A minimum of 5 years’ experience of administration of Linux systems
Strong ability to analyze, debug and maintain the integrity of an existing code base
Demonstrated equivalence of 5 years of Linux/UNIX user support experience and hands-on experience with administration of Linux systems
Superior scripting skills and excellent attention to detail; proficiency in at least Python, Perl, or Bash
Strong ability to interact with customers to understand needs, elicit requirements, and get feedback on prototype solutions
Excellent communication and people skills; excellent time management and organizational skills
Experience with system configuration management tools e.g. , ansible, puppet, chef
Experience with revision control software e.g. CVS, SVN, Git
Proficiency at technical writing
Preferred Skills (Requesting Manager Defines):
Proficiency with analysis and problem-solving skills for debugging and optimization of applications
Familiarity/proficiency with OpenMP and Message Passing Interface (MPI) programming
Experience with Lustre, and InfiniBand
Experience with cloud technologies (AWS, Azure, GCP), OpenStack or Kubernetes is a plus
Location & Eligibility
Where is the job
—
Location terms not specified
Listing Details
- Posted
- September 23, 2026
- First seen
- September 23, 2026
- Last seen
- September 24, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 55%
- Scored at
- September 23, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Systems Administrator jobs
View all →S
SystemstechnologyresearchClassified Systems Administrator
$105k–$125k/yr
B
BogaardgroupSecurity Systems Administrator
USD 120000–150000
Full-Time
Systems Administrator
RevOps Systems Administrator (Salesforce)
C
ComptechcomputertechnologiesSystems Administrator III
$75k–$80k/yr
Critical Systems Administrator – BMS/BAS & Video Management
Full Time
Browse Similar Jobs
Systems Engineer1.9kNetwork Engineer700Site Reliability Engineer217Platform Engineer198Infrastructure Engineer142Network Administrator85It Operations Engineer75Cloud Engineer66Azure Cloud Engineer56Storage Engineer46Cloud Operations Engineer46Infrastructure Architect37Observability Engineer35Network Operations Engineer32Kubernetes Engineer24Server Engineer23Cloud Architect23Azure DevOps Engineer23Linux Administrator22Build and Release Engineer22
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.