davidjoseph-co
New

Causal Labs — Machine Learning Infrastructure Engineer

United StatesUnited States·San Franciscomid
OtherMachine Learning Infrastructure Engineer
8 views0 saves0 applied

Quick Summary

Key Responsibilities

Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).

Technical Tools
OtherMachine Learning Infrastructure Engineer

Responsibilities

~1 min read
  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
  • Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales
  • Profile and debug low-level GPU operations to optimize performance
  • Track new research and bring fresh ideas into the work

Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).

Requirements

~1 min read
  • 2–10 years building large-scale ML infrastructure for core foundation models
  • Hands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing models
  • A background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)
  • Deep, demonstrable expertise optimizing large-scale training and inference workloads
  • Working proficiency with distributed training frameworks such as FSDP or DeepSpeed
  • A clear pattern of intentional, mission-driven career decisions
  • Able to work on-site 5 days/week in San Francisco (relocation supported)

Nice to Have

~1 min read
  • Generalist experience spanning the full ML lifecycle
  • Low-level GPU performance optimization and debugging (CUDA, JAX)

What We Offer

~1 min read
Take a bet on a distinctive, non-consensus approach to building intelligence
Join early, with real ownership of the training and inference backbone
Work in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM training
Well-funded and building a strong, senior research and engineering team
  • Location: San Francisco, CA
  • Work policy: In-person 5 days/week (relocation supported)
  • Compensation: $200K–$400K + competitive early-stage equity
  • Visa sponsorship: Open to supporting work authorization for the right candidate
  • Employment type: Full-time

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

First seen
August 21, 2026
Last seen
September 4, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
51%
Scored at
August 21, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

davidjoseph-coCausal Labs — Machine Learning Infrastructure Engineer