Cantina Labs Logo

Cantina Labs

Machine Learning Engineer, Ops

Posted 18 Days Ago
Remote
Hiring Remotely in Greece
125K-165K Annually
Mid level
Remote
Hiring Remotely in Greece
125K-165K Annually
Mid level
Design, deploy, and scale low-latency inference infrastructure for generative audio models (TTS, ASR, voice conversion). Build high-performance inference engines, Kubernetes-based autoscaling, CI/CD pipelines, observability, and GPU optimization to bridge research and production for streaming and batch workloads.
The summary above was generated by AI

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

 

About the Role:

We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.

What You’ll Do:

  • Design and maintain inference infrastructure for generative audio model architectures.

  • Implement and manage high-performance inference engines.

  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.

  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.

  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.

  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.

  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring:

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.

  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.

  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.

  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.

  • Experience with GPU-accelerated inference and performance profiling techniques.

  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation:

The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

 

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Similar Jobs

9 Days Ago
In-Office or Remote
Mid level
Mid level
Information Technology • Software
Build and operate production-grade model serving infrastructure, design deployment pipelines (blue/green, canary), implement autoscaling and multi-model serving, optimize GPU utilization and network throughput, set up observability and model registries, manage CI/CD for reproducible deployments, own full ML system lifecycle including on-call support and platform scalability.
Top Skills: Ci/CdCudaExperiment TrackingGpuHelmKubeaiKubeflowMlflowModel RegistryPythonRocmTerraformTgiTritonVllm
16 Days Ago
In-Office or Remote
Mid level
Mid level
Information Technology • Software
The ML Ops Engineer will build and operate scalable ML inference platforms, focusing on model serving infrastructure and deployment pipelines for AI applications.
Top Skills: CudaHelmPythonRocmTerraformTgiTritonVllm
7 Hours Ago
Remote
United States
192K-287K Annually
Senior level
192K-287K Annually
Senior level
Artificial Intelligence • Productivity • Software • Automation
As a Sr. Applied AI Engineer at Zapier, you will build and enhance AI platform capabilities, focusing on LLM Ops and ML Ops to support scalable AI development across teams.
Top Skills: Cloud InfrastructureLlm OpsMl OpsPythonTypescript

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account