Quick Summary
Build and operate scalable infrastructure for data generation and simulation workflows (job orchestration, scheduling, queues, retries, observability).
pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring.
📍 San Francisco | 🏢 5 Days Onsite
Location: Onsite in San Francisco
Compensation: Competitive Salary + Equity
Engineering simulation is one of the last major categories of software that AI hasn't rebuilt. The tools used to design aircraft, ships, reservoirs, and medical devices still run on numerical methods that are decades old, and an engineer can wait a full day for a single answer. UniversalAGI is building foundation models that learn physics directly from data, and they are already running in early deployments on real computational fluid dynamics and reservoir engineering problems for some of the largest industrial and defense organizations in the world.
We are a team of 25 researchers and engineers in San Francisco backed by Elad Gil (#1 Solo VC), Eric Schmidt (former Google CEO), Prith Banerjee (ANSYS CTO), Ion Stoica (Databricks Founder), Jared Kushner (former Senior Advisor to the President), David Patterson (Turing Award Winner), and Luis Videgaray (former Foreign and Finance Minister of Mexico).
About the Role
~1 min readUniversalAGI is hiring a ML Platform Engineer to build and own the execution platform powering our research and customer deployments: data generation + simulation orchestration + training/fine-tuning infrastructure + benchmarking pipelines + production deployments in customer environments.
You’ll work closely with the CEO and founding team to turn research into repeatable, scalable, reliable systems - internally and in customer infrastructure. This is a “ship outcomes” role: your work directly determines how fast we can iterate, how reproducible our results are, and how reliably we deliver in production.
Responsibilities
~1 min readRequirements
~1 min read-
Experience with workflow orchestration (e.g., Ray, Kubernetes, Slurm).
Experience with GPU infrastructure and distributed training systems.
Experience building evaluation/benchmarking frameworks with strong reproducibility guarantees.
Experience deploying into regulated / security-sensitive environments (gov/defense/enterprise).
Experience with simulation/HPC pipelines (CFD, meshing, batch workloads) is a plus but not required.
Experience in an FDE-style / delivery execution role (or similar “ship results fast” environments).
What We Offer
~1 min read“The credit belongs to the man who is actually in the arena, whose face is marred by dust and sweat and blood; who strives valiantly; who errs, who comes short again and again... who at the best knows in the end the triumph of high achievement, and who at the worst, if he fails, at least fails while daring greatly." - Teddy Roosevelt
At our core, we believe in being “in the arena.” We are builders, problem solvers, and risk-takers who show up every day ready to put in the work: to sweat, to struggle, and to push past our limits. We know that real progress comes with missteps, iteration, and resilience. We embrace that journey fully knowing that daring greatly is the only way to create something truly meaningful.
Location & Eligibility
Listing Details
- Posted
- September 24, 2026
- First seen
- September 25, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 57%
- Scored at
- September 26, 2026
Signal breakdown
Similar Ml Platform Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.