Skip to main content
Back to Jobs

Machine Learning Platform Engineer

Build and operate ML infrastructure for AI model training and deployment

As an ML Platform Engineer at Bjak, you will design and operate the systems behind A1's AI stack, including model training, evaluation, deployment, and inference. You will build platforms and tooling that enable AI engineers and researchers to experiment quickly and ship models reliably. Your work will focus on improving reliability, scalability, latency, and cost efficiency of AI systems through optimized serving infrastructure and production...

Why This Role?

Work on turning evolving model requirements into production-ready infrastructure with direct impact on AI reliability

Key Responsibilities

  • Design and operate ML infrastructure for model training, evaluation, deployment, and inference
  • Build and optimize model serving infrastructure for high-throughput and low-latency workloads
  • Develop reliable pipelines for data preparation, training, evaluation, and model release
  • Build production observability, monitoring, tracing, and alerting for AI/ML workloads
  • Develop evaluation and benchmarking infrastructure to measure model quality and performance
  • Identify bottlenecks across the ML stack and continuously improve system performance

Requirements

  • Python
  • PyTorch or JAX
  • Experience with LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-L
  • Experience building ML platforms and tooling for AI experimentation and deployment
  • Knowledge of model training, evaluation, and continuous improvement systems
  • Familiarity with production observability and monitoring for AI workloads

Required Skills

pythonpytorchmachine-learningml-infrastructuresystem-designai-engineeringML InfrastructureModel ServingMLOpsSystem Optimization
View Original Description from Ashby Job Boards

Original description from Ashby Job Boards

About A1 There are over 5 billion users using basic applications today such email, notes, tasks that are not AI-native. Our mission is to build a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organising and workflows, with minimal prompting. Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. The system must handle multi-step reasoning, interact with external tools, and remain reliable despite non-deterministic model behavior. Our objective is to help users complete tasks daily enjoyable with over ~90%* reduced time. About the Role As an ML Platform Engineer, you will build the infrastructure and systems that power A1's AI capabilities. You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence. Focus - Build and operate the ML infrastructure and platforms powering A1’s AI products - Design systems for model training, evaluation, deployment, inference, and experimentation - Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads - Improve reliability, scalability, latency, and cost efficiency of AI systems - Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement - Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster - Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions - Build production observability, monitoring, tracing, and alerting for AI/ML workloads - Improve AI systems across reliability, scalability, latency, throughput, and cost - Identify bottlenecks across the ML stack and continuously improve system performance - Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure Tech Stack - Python - PyTorch / JAX - LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM - Cloud infrastructure - Distributed systems - ML/data pipelines and workflow orchestration - GPU infrastructure and performance tooling - Vector databases and retrieval infrastructure Ideal Experience - Strong software engineering fundamentals and experience building production systems - Experience building ML infrastructure, platforms, or production machine learning systems - Experience with model deployment, inference, evaluation, or data pipelines - Strong understanding of distributed systems and system reliability - Ability to write clean, maintainable, production-quality code - Comfortable working in ambiguous, fast-moving environments - Bias toward ownership, experimentation, and continuous improvement Outcomes - AI infrastructure reliably supports production workloads at scale - Models can be trained, evaluated, deployed, and improved efficiently - Inference systems deliver strong latency, throughput, reliability, and cost efficiency - ML pipelines are reproducible, observable, maintainable, and robust - Model and infrastructure regressions are detected quickly and diagnosed efficiently - Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product - The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge

Salary Context

Similar Engineering roles on LokerDollar pay around $160k/yr (range $1.8k–999.999k/yr, n=643 active listings).

Hiring at Bjak

Bjak has 139 other active roles on LokerDollar and has been hiring here since May 7, 2026 — across Engineering, AI, Product, Data & Analytics.

View all Bjak openings →

Openness not stated by employer — check the listing

Company
Bjak
Source
Ashby Job Boards
Job Type
full time
Location
Remote
Category
Seniority
mid
PostedFreshNew & verified
Aug 11, 2026

Share this job

Help a friend find their next remote role.

Frequently asked questions

Is Machine Learning Platform Engineer at Bjak a remote job?
This role is based in Remote. See the listing for remote/onsite details.
What type of employment is Machine Learning Platform Engineer at Bjak?
This is a full time position.
How do I apply?
Click the "Apply" button on this page to go to the official application at Bjak.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.

From the blog