Machine Learning Platform Engineer
Build and operate ML infrastructure for AI model training and deployment
As an ML Platform Engineer at Bjak, you will design and operate the systems behind A1's AI stack, including model training, evaluation, deployment, and inference. You will build platforms and tooling that enable AI engineers and researchers to experiment quickly and ship models reliably. Your work will focus on improving reliability, scalability, latency, and cost efficiency of AI systems through optimized serving infrastructure and production...
Why This Role?
Work on turning evolving model requirements into production-ready infrastructure with direct impact on AI reliability
Key Responsibilities
- Design and operate ML infrastructure for model training, evaluation, deployment, and inference
- Build and optimize model serving infrastructure for high-throughput and low-latency workloads
- Develop reliable pipelines for data preparation, training, evaluation, and model release
- Build production observability, monitoring, tracing, and alerting for AI/ML workloads
- Develop evaluation and benchmarking infrastructure to measure model quality and performance
- Identify bottlenecks across the ML stack and continuously improve system performance
Requirements
- Python
- PyTorch or JAX
- Experience with LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-L
- Experience building ML platforms and tooling for AI experimentation and deployment
- Knowledge of model training, evaluation, and continuous improvement systems
- Familiarity with production observability and monitoring for AI workloads
Required Skills
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
About A1 There are over 5 billion users using basic applications today such email, notes, tasks that are not AI-native. Our mission is to build a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organising and workflows, with minimal prompting. Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. The system must handle multi-step reasoning, interact with external tools, and remain reliable despite non-deterministic model behavior. Our objective is to help users complete tasks daily enjoyable with over ~90%* reduced time. About the Role As an ML Platform Engineer, you will build the infrastructure and systems that power A1's AI capabilities. You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence. Focus - Build and operate the ML infrastructure and platforms powering A1’s AI products - Design systems for model training, evaluation, deployment, inference, and experimentation - Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads - Improve reliability, scalability, latency, and cost efficiency of AI systems - Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement - Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster - Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions - Build production observability, monitoring, tracing, and alerting for AI/ML workloads - Improve AI systems across reliability, scalability, latency, throughput, and cost - Identify bottlenecks across the ML stack and continuously improve system performance - Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure Tech Stack - Python - PyTorch / JAX - LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM - Cloud infrastructure - Distributed systems - ML/data pipelines and workflow orchestration - GPU infrastructure and performance tooling - Vector databases and retrieval infrastructure Ideal Experience - Strong software engineering fundamentals and experience building production systems - Experience building ML infrastructure, platforms, or production machine learning systems - Experience with model deployment, inference, evaluation, or data pipelines - Strong understanding of distributed systems and system reliability - Ability to write clean, maintainable, production-quality code - Comfortable working in ambiguous, fast-moving environments - Bias toward ownership, experimentation, and continuous improvement Outcomes - AI infrastructure reliably supports production workloads at scale - Models can be trained, evaluated, deployed, and improved efficiently - Inference systems deliver strong latency, throughput, reliability, and cost efficiency - ML pipelines are reproducible, observable, maintainable, and robust - Model and infrastructure regressions are detected quickly and diagnosed efficiently - Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product - The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge
Salary Context
Similar Engineering roles on LokerDollar pay around $160k/yr (range $1.8k–999.999k/yr, n=643 active listings).
Hiring at Bjak
Bjak has 139 other active roles on LokerDollar and has been hiring here since May 7, 2026 — across Engineering, AI, Product, Data & Analytics.
- Technical Lead, Machine Learning
- Staff Machine Learning Engineer
- Business Intelligence & Performance Manager
Openness not stated by employer — check the listing
Frequently asked questions
- Is Machine Learning Platform Engineer at Bjak a remote job?
- This role is based in Remote. See the listing for remote/onsite details.
- What type of employment is Machine Learning Platform Engineer at Bjak?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Bjak.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Interview Radiologist Remote: What You NeedBreaking down the remote radiology interview process for USD-paying jobs, from screening to offer letter.
- Remote Radiology Jobs: August 2026 UpdateDiscover the latest remote radiology jobs and salary trends for August 2026. Learn about the opportunities and challenges of remote work in radiology.
- Unlocking Remote USD Jobs: A GuideDiscover the best remote USD-paying jobs and learn how to succeed in the global job market.