Software Engineer (Compute Efficiency), London
Quick Summary
Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads Have experience working in cloud compute infrastructure design,
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease.
The future is coming. A future enabled and enriched by the incredible power of machine learning. A future in which diseases are curtailed or cured starting with better and faster drug discovery.
Come and be part of an interdisciplinary team driving groundbreaking innovation and play a meaningful role in contributing towards us achieving our ambitious goals, while being a part of an inspiring and collaborative culture.
The world we want tomorrow is the one we’re building today. It starts with the culture at this company. It starts with you.
Isomorphic Labs (IsoLabs) was launched in 2021 to advance human health by building on and beyond the Nobel-winning AlphaFold system. Since then, our interdisciplinary team of drug discovery experts and machine learning specialists has built powerful new predictive and generative AI models that accelerate scientific discovery at digital speed.
Our name comes from the belief that there is an underlying symmetry between biology and information science. By harnessing AI’s powerful capabilities, we can use it to model complex biological phenomena to help design novel molecules, anticipate how drugs will perform and develop innovative medicines to treat and cure some of the world’s most devastating diseases.
We have built a world-leading drug design engine comprising AI models that are capable of working across multiple therapeutic areas and drug modalities. We are continually innovating on model architecture and developing cutting-edge capabilities to advance rational drug design.
Every day, and with each new breakthrough, we’re getting closer to the promise of digital biology, and achieving our ambitious mission to one day solve all disease with the help of AI.
We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a key role and work at a grand scale to deliver the foundations that make this happen. Joining the Compute Infrastructure team, you will ensure our planet-scale accelerator fleet operates at peak health and efficiency. By partnering with in-house machine learning platforms, performance & scaling team, and AI researchers, you will design the observability systems and automated recovery mechanisms that maximize scientific throughput on every paid GPU-hour.
Responsibilities
~1 min read- →Design, deploy, and scale robust observability systems and telemetry pipelines to monitor fleetwide compute efficiency, hardware health, and workload goodput across distributed clusters
- →Drive hardware efficiency and node reliability across our accelerator fleet, and integrating new hardware to leverage advancements
- →Identify compute waste and efficiency bottlenecks across the fleet, partnering with ML and platform teams to actively optimize accelerator utilization and improve workload goodput
- →Collaborate with teams in the AI org, and work closely with the ML Infrastructure team to identify, instrument, and improve canonical efficiency metrics across both training and inference runs
- →Contribute to the efforts for consistently improving the reliability of our ML runs.
- →Operate, maintain, and harden research, development, and production cloud infrastructure and cluster deployments
- →Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations
- →Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design)
Requirements
~1 min read- Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
- Have experience working in cloud compute infrastructure design, preferably GCP
- Possess strong programmings skills
- Have significant experience working and deploying in Kubernetes at scale
- Familiarity with the Nvidia GPU generations
- Proven track record of building production observability and telemetry stacks
Nice to Have
~1 min read- Have a background in either ML SWE or infrastructure SRE work to build on
- Conceptual understanding of ML workload efficiency paradigms
- Have experience leading and delivering projects to multidisciplinary stakeholders
- Familiarity with Google TPU generations
- Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking
We are guided by our shared values. It's not about finding people who think and act in the same way. These values help to guide our work and will continue to strengthen it.
Location & Eligibility
Listing Details
- Posted
- October 6, 2026
- First seen
- October 6, 2026
- Last seen
- October 6, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 53%
- Scored at
- October 7, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.