Machine Learning Engineer, Ops
Quick Summary
Design and maintain inference infrastructure for generative audio model architectures. Implement and manage high-performance inference engines. Orchestrate service deployments using Kubernetes (K8S),
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
About the Role
~1 min readWe are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.
Responsibilities
~1 min read- →
Design and maintain inference infrastructure for generative audio model architectures.
- →
Implement and manage high-performance inference engines.
- →
Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
- →
Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
- →
Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
- →
Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
- →
Optimize inference performance for both streaming and batch applications.
Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.
Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
Experience with GPU-accelerated inference and performance profiling techniques.
Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.
What We Offer
~1 min readThe anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Location & Eligibility
Listing Details
- Posted
- July 27, 2026
- First seen
- July 27, 2026
- Last seen
- July 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 61%
- Scored at
- July 27, 2026
Signal breakdown
Please let cantina know you found this job on Jobera.
3 other jobs at cantina
View all →Explore open roles at cantina.
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.