davidjoseph-co~22d ago
New
New
Causal Labs — Machine Learning Infrastructure Engineer
OtherMachine Learning Infrastructure Engineer
8 views0 saves0 applied
Quick Summary
Key Responsibilities
Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).
Technical Tools
OtherMachine Learning Infrastructure Engineer
Responsibilities
~1 min read- →Design, deploy, and maintain large distributed ML training and inference clusters
- →Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
- →Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales
- →Profile and debug low-level GPU operations to optimize performance
- →Track new research and bring fresh ideas into the work
Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).
Requirements
~1 min read- 2–10 years building large-scale ML infrastructure for core foundation models
- Hands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing models
- A background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)
- Deep, demonstrable expertise optimizing large-scale training and inference workloads
- Working proficiency with distributed training frameworks such as FSDP or DeepSpeed
- A clear pattern of intentional, mission-driven career decisions
- Able to work on-site 5 days/week in San Francisco (relocation supported)
Nice to Have
~1 min read- Generalist experience spanning the full ML lifecycle
- Low-level GPU performance optimization and debugging (CUDA, JAX)
What We Offer
~1 min read✓Take a bet on a distinctive, non-consensus approach to building intelligence
✓Join early, with real ownership of the training and inference backbone
✓Work in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM training
✓Well-funded and building a strong, senior research and engineering team
- Location: San Francisco, CA
- Work policy: In-person 5 days/week (relocation supported)
- Compensation: $200K–$400K + competitive early-stage equity
- Visa sponsorship: Open to supporting work authorization for the right candidate
- Employment type: Full-time
Location & Eligibility
Where is the job
San Francisco, United States
On-site at the office
Who can apply
US
Listing Details
- First seen
- August 21, 2026
- Last seen
- September 4, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 51%
- Scored at
- August 21, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on davidjoseph-co's site
Please let davidjoseph-co know you found this job on Jobera.
4 other jobs at davidjoseph-co
View all →Explore open roles at davidjoseph-co.
Similar Machine Learning Infrastructure Engineer jobs
View all →Senior Machine Learning Infrastructure Engineer, Embedding Platform
USD 190800-267100
Remote
Staff Machine Learning Infrastructure Engineer, Embedding Platform
USD 253300-354600
Remote
Machine Learning Infrastructure Engineer
Remote
Machine Learning Infrastructure Engineer, Safeguards Research
USD 350000-500000
C
CssmergeStaff Machine Learning Infrastructure Engineer
$224k–$280k/yr
Machine Learning Infrastructure Engineer
$140k–$165k/yr
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.