tensorwave
tensorwave13d ago
New

Senior Manager, Cluster Engineering & Deployment

Remotefull-timesenior
OtherDeployment
0 views0 saves0 applied

Quick Summary

Key Responsibilities

staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.

Requirements Summary

building and enforcing playbooks, gates, metrics, and blameless defect loops. Team leadership with schedule accounta

Technical Tools
OtherDeployment

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

About the Role

~1 min read

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

Responsibilities

~1 min read
  • →

    Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.

  • →

    Lead deployment engineering across concurrent cluster builds, through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling vendors.

  • →

    Drive deployment velocity engineering: cut bring-up time per cluster through tooling, pre-staging, and defect-source elimination, and set the targets the team is measured against.

  • →

    Own defect feedback loops to Network Engineering (design), Layer One (cabling quality), and vendors (hardware/optics RMA patterns), holding those partners accountable to resolution.

  • →

    Define spares, test equipment, and deployment tooling requirements per site, and standardize them across sites.

Requirements

~1 min read
  • 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting.

  • Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment).

  • Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops.

  • Team leadership with schedule accountability across multiple concurrent builds or sites.

  • GPU cluster validation experience (NCCL/RCCL benchmarking).

  • Automation skills (Python, Ansible) applied to deployment.

  • Optics/link-layer debugging depth.

  • Experience with acceptance testing as a commercial gate (revenue-linked).

  • Own cluster validation end to end: bandwidth/latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates and raise the bar on each as the fleet scales.

What We Offer

~1 min read
✓Stock Options
✓100% paid Medical, Dental, and Vision insurance for Employees
✓Company Health Savings Account Contributions
✓100% paid Short Term and Long Term Disability Insurance for Employees
✓Life and Voluntary Supplemental Insurance Options
✓Other Insurance Options, such as Pet & Legal Insurance
✓Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
✓Flexible Spending Account
✓401(k)
✓Employee Assistance Program
✓Flexible PTO
✓Paid Holidays
✓Parental Leave
✓Other In-Office Perks

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact accomodations@tensorwave.com.

All offers of employment are contingent upon verification of identity and authorization to work in the United States, as required by law.

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location

Listing Details

Posted
September 15, 2026
First seen
September 26, 2026
Last seen
September 27, 2026

Posting Health

Days active
1
Repost count
0
Trust Level
33%
Scored at
September 27, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

tensorwaveSenior Manager, Cluster Engineering & Deployment