Skip to main content
Back to Jobs

Capacity Operations Manager

Drive GPU fleet health and utilization across neocloud and bare metal environments

Own the operational and analytical supply side of Baseten's GPU fleet, focusing on lifecycle, health, observability, utilization monitoring, and remediation. Drive suppliers to keep the maximum amount of the GPU fleet online and healthy by maintaining live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity. Monitor SLA performance, file credit claims, and enforce remedies when suppliers fall short, while coordinatin...

Why This Role?

Help build the platform engineers turn to ship AI products at a rapidly growing AI infrastructure company

Key Responsibilities

  • Drive suppliers to maximize GPU fleet availability and health across neocloud and bare metal environments
  • Maintain live reconciliation of contracted, provisioned, healthy, and utilized GPU capacity by supplier and cluster
  • Own supplier-attributed fleet health accountability including replacement SLAs, MTTR, and RMA cycle times
  • Monitor SLA performance, file and pursue credit claims, and drive remediation plans for supplier underperformance
  • Coordinate internal communications to ensure stakeholders are aware of supplier maintenance impacting availability

Requirements

  • Experience driving supplier performance and managing vendor relationships
  • Strong analytical skills for monitoring utilization, health, and capacity metrics
  • Experience with observability and monitoring systems in infrastructure environments
  • Ability to drive cross-functional alignment and clear process execution
  • Ownership mindset with focus on execution and closing operational gaps

Required Skills

gpu managementcapacity planningvendor coordinationsla monitoringcommunicationvendor managementincident responsecross-functional coordinationdata reconciliation
View Original Description from Ashby Job Boards

Original description from Ashby Job Boards

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: - Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. - Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. - Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. - SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. - Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and Approach - Flexibility: this list covers the core of the role, not the limit of it. You'll be asked to take on adjacent work as the function evolves and as new gaps surface. - Ownership mindset: we need someone who treats "whatever it takes" as a genuine operating principle, not a line in a job posting. If something falls outside a defined lane but inside the overall goal of closing the capacity gap, it's yours to pick up. REQUIREMENTS - 5 to 10+ years within infrastructure working within the compute lifecycle to maximize functional compute, ideally in a hyperscale, cloud, or large-scale compute environment. - Direct experience managing GPU, server, or data center hardware supplier relationships. You understand fleet health, RMA processes, and how contracted capacity differs from delivered capacity. - Highly analytical. You should be comfortable pulling your own data, building your own reports, and generating insights without waiting on someone else to hand you a dashboard. - Comfortable with ambiguity. Part of the job is figuring out what should exist and building it. - Strong cross-functional collaboration skills. You'll work closely with finance, infrastructure/engineering, legal, and security on a regular basis. PREFERRED QUALIFICATIONS: - Experience at a hyperscaler, neo cloud provider, or AI infrastructure company - Familiarity with GPU hardware lifecycles (NVIDIA H100/H200/GB200 class systems), power/thermal constraints, and supply chain dynamics for compute. - Experience running formal supplier corrective actions. BENEFITS - Competitive compensation, including meaningful equity - 100% coverage of medical, dental, and vision insurance for employee and dependents - Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!) - Paid parental leave - Fertility and family-building stipend through Carrot - Company-facilitated 401(k) - Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities. Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you. At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Salary Context

Similar Engineering roles on LokerDollar pay around $218.4k/yr (range $41k–1000k/yr, n=621 active listings).

Hiring at Baseten

Baseten has 35 other active roles on LokerDollar and has been hiring here since Jun 23, 2026 — across Engineering, Data & Analytics.

View all Baseten openings →

Openness not stated by employer — check the listing

Company
Baseten
Salary
$225k–235k/yr
See remote (USD) vs local pay →
Job Type
full time
Location
San Francisco, USA · Remote
Category
Seniority
mid
PostedFreshNew & verified
Sep 5, 2026

Share this job

Help a friend find their next remote role.

Frequently asked questions

Is Capacity Operations Manager at Baseten a remote job?
This role is based in Remote. See the listing for remote/onsite details.
What is the salary for Capacity Operations Manager at Baseten?
The listed pay range for this role is $225k–235k/yr.
What type of employment is Capacity Operations Manager at Baseten?
This is a full time position.
How do I apply?
Click the "Apply" button on this page to go to the official application at Baseten.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.

From the blog