BJAK Logo

BJAK

Machine Learning Platform Engineer

Posted One Month Ago
Remote
Hiring Remotely in United States
Mid level
Remote
Hiring Remotely in United States
Mid level
Design, build, and operate ML infrastructure for training, evaluation, deployment, and inference. Improve reliability, scalability, latency, throughput, and cost. Create pipelines, observability, benchmarking, and tooling to enable fast experimentation and productionization of models while diagnosing regressions and bottlenecks.
The summary above was generated by AI

About ActAI

There are over 5 billion users using basic applications today such email, notes, tasks, calendar and they're not AI-native. Our mission is to build proactive applications for anyone in the world, who are not used to complex prompting. We aim to bring intelligence to conversations, errands, organising and workflows, with minimal to no prompting.

Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. We believe products will greatly reduce hallucinations.

Our objective is to organise anyone's life, allowing us all to spend time on valuable and meaningful things.

About the Role

As an ML Platform Engineer, you will build the infrastructure and systems that power ActAI's AI capabilities.

You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement.

You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence.

Focus

  • Build and operate the ML infrastructure and platforms powering A1’s AI products

  • Design systems for model training, evaluation, deployment, inference, and experimentation

  • Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads

  • Improve reliability, scalability, latency, and cost efficiency of AI systems

  • Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement

  • Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster

  • Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions

  • Build production observability, monitoring, tracing, and alerting for AI/ML workloads

  • Improve AI systems across reliability, scalability, latency, throughput, and cost

  • Identify bottlenecks across the ML stack and continuously improve system performance

  • Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure

Tech Stack

  • Python

  • PyTorch / JAX

  • LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM

  • Cloud infrastructure

  • Distributed systems

  • ML/data pipelines and workflow orchestration

  • GPU infrastructure and performance tooling

  • Vector databases and retrieval infrastructure

Ideal Experience

  • Strong software engineering fundamentals and experience building production systems

  • Experience building ML infrastructure, platforms, or production machine learning systems

  • Experience with model deployment, inference, evaluation, or data pipelines

  • Strong understanding of distributed systems and system reliability

  • Ability to write clean, maintainable, production-quality code

  • Comfortable working in ambiguous, fast-moving environments

  • Bias toward ownership, experimentation, and continuous improvement

Outcomes

  • AI infrastructure reliably supports production workloads at scale

  • Models can be trained, evaluated, deployed, and improved efficiently

  • Inference systems deliver strong latency, throughput, reliability, and cost efficiency

  • ML pipelines are reproducible, observable, maintainable, and robust

  • Model and infrastructure regressions are detected quickly and diagnosed efficiently

  • Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product

  • The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge

Similar Jobs

4 Hours Ago
Remote or Hybrid
219K-335K Annually
Senior level
219K-335K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Lead development of scalable, reliable continuous integration infrastructure supporting autonomous-vehicle development, machine learning training, simulation, and remote builds. Specialize in Remote Build Execution and a FUSE-based file system, enabling source-code editing and developer workflows. Design and implement productivity improvements, evaluate technologies, influence technical roadmaps, establish engineering best practices, manage technical debt, and mentor engineers while balancing business and customer priorities.
Top Skills: DockerFuseGoGoogle Cloud Platform (Gcp)KubernetesNetworkingPythonRemote Build Execution (Rbe)SshUnix/Linux
7 Days Ago
Remote
United States
145K-250K Annually
Senior level
145K-250K Annually
Senior level
Cloud • Software
Build and operate a secure Kubernetes-based AI/ML platform for government test and evaluation teams. Responsibilities include GPU scheduling, workload orchestration, infrastructure-as-code, GitOps, observability, capacity planning, upgrades, security hardening, compliance support, and operational documentation. The role involves restricted and disconnected environments, automated security tooling, collaboration with government and engineering stakeholders, technical leadership, and mentoring.
Top Skills: Argo CdArtifact SigningAWSAzureCi/CdContainer ScanningDastGitopsGoGCPGpu SchedulingGrafanaHelmKserveKubernetesKubernetes OperatorsLlm-DMlopsNebariNist 800-171Nist 800-53OpentelemetryOpentofuPolicy EnforcementPrometheusPythonRisk Management FrameworkSastTerraformVllm
10 Days Ago
Remote
United States
164K-240K Annually
Senior level
164K-240K Annually
Senior level
Automotive • Insurance • Machine Learning • Mobile • Software
Lead development of Root’s machine learning platform for insurance pricing. Build and operate feature pipelines, model training and orchestration, registries, validation, serving, observability, and reproducibility tooling. Partner with researchers to convert experimentation into reliable production systems, explore LLM-enabled data science automation, drive technical roadmaps and execution, and mentor engineers while maintaining high standards for reliability and correctness.
Top Skills: Agentic SystemsAutomated ValidationData PipelinesDistributed SystemsFeature PipelinesFeature StoresLarge Language Models (Llms)Machine LearningModel RegistriesModel ServingModel Training OrchestrationModel VersioningProduction ObservabilityPython

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account