Skip to main content
Back to Jobs

Principal Operations Engineer Hardware — Data Center Operations

Lead site assessments and operational audits for hyperscale AI data center hardware

As Principal Operations Engineer, Hardware, you will serve as the most senior technical authority for the operational hardware fleet across Fluidstack's hyperscale AI data center portfolio. You will lead site assessments and operational audits, drive technical readiness of teams ahead of site activation, and review hardware platforms and integration designs from an operational lens. You will feed operational learnings back into hardware engine...

Why This Role?

Be the technical arm of senior operations leadership in the field, acting as a force multiplier across site hardware leads and deployment...

Key Responsibilities

  • Lead site assessments and operational audits for hyperscale AI data center hardware
  • Drive technical readiness of teams ahead of site activation
  • Review hardware platforms and integration designs from an operational perspective
  • Feed operational learnings back into hardware engineering, deployment, and supply chain teams
  • Act as a force multiplier across site hardware leads, deployment teams, and reliability engineers
  • Serve as connective tissue between hardware operations, hardware engineering, network, facilities, and customer-facing teams

Requirements

  • Career experience operating hardware at scale in hyperscale data centers or large HPC environments
  • Comfort diagnosing stubborn boot failures on the data center floor
  • Experience leading fleet-wide root cause investigations
  • Ability to push back on vendors regarding flawed RMA processes
  • Formal engineering credentials valued but not required

Required Skills

hardware operationsdata center managementai infrastructuresystem reliabilitytechnical leadershipdata center operationssite assessmentsoperational auditstechnical readinesshardware engineering collaboration

Indonesia Context

Working Hours Overlap:
Flexible — work your own hours
See remote (USD) vs local pay →

Keywords

Principal Operations Engineerhyperscale data centerAI infrastructurehardware fleet operationssite activation readinessoperational learnings feedback Fluidstackremote worldwide
View Original Description from RemoteOK

Original description from RemoteOK

About Fluidstack We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it. We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI. We hire people who care deeply about this problem space. If that is you, please apply! About the Role We are seeking a Principal Operations Engineer, Hardware to serve as the most senior technical authority for the operational hardware fleet across our hyperscale AI data center portfolio. AI infrastructure lives and dies on the reliability of the compute itself — this role exists to ensure that the GPU systems, servers, and supporting hardware we deploy at scale are operated, maintained, and continuously improved at the standard the workload demands. You will operate as the technical arm of senior operations leadership in the field — leading site assessments and operational audits, driving the technical readiness of teams ahead of site activation, reviewing hardware platforms and integration designs from an operational lens, and feeding operational learnings back into the hardware engineering, deployment, and supply chain organizations as we shift toward a productized, repeatable build model. You will be a force multiplier across our site hardware leads, deployment teams, and reliability engineers, and the connective tissue between hardware operations, hardware engineering, network, facilities, and customer-facing teams. The ideal candidate has spent a career operating hardware at scale — in hyperscale data centers, large HPC environments, or comparable 24/7 infrastructure — and is equally comfortable diagnosing a stubborn boot failure on the floor, leading a fleet-wide root cause investigation, and pushing back on a vendor on a flawed RMA process. Formal engineering credentials are valued but not required — practical depth, judgment under pressure, the ability to teach, and the discipline to keep critical infrastructure running through change are what define this role. Responsibilities 10+ years of hands-on experience operating mission-critical hardware infrastructure, with at least 5 years as the senior technical voice on a site, campus, or fleet. Data center operations experience strongly preferred; hyperscale, large HPC, cloud, or other mission-critical compute infrastructure experience considered. Deep working command of GPU systems, server platforms, storage infrastructure, firmware lifecycle management, and hardware diagnostics — earned in the field, not from a textbook. Demonstrated ability to author, approve, and execute high-risk MOPs and change records in live production environments. A track record of leading root cause analysis on significant hardware events and driving corrective actions to closure. A track record of holding OEMs, ODMs, service vendors, and deployment partners accountable — you know how to enforce a standard without burning the relationship. Strong written communication: operational health assessments, RCAs, procedure reviews, and design review feedback are second nature. Comfort operating as the senior technical voice across operations, hardware engineering, network, facilities, supply chain, and customer-facing teams. Willingness to travel extensively across the fleet. 50-75%. Preferred Qualifications Bachelor's degree in Computer En

Source site may be blocked by Indonesian ISPs

Some Indonesian ISPs (Telkomsel, Indihome) block RemoteOK. If the Apply button doesn't open, try mobile data or a VPN.

Tip: switch network or enable a VPN, then click Apply again.

Openness not stated by employer — check the listing

Company
Fluidstack
Source
RemoteOK
Salary
Job Type
full time
Location
Remote · Open worldwide
Category
Seniority
senior
Posted
Jun 9, 2026

Share this job

Help a friend find their next remote role.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.