Skip to main content
Back to Jobs

Staff Software Engineer Infrastructure

Build self-service infrastructure for multi-region platform adoption

As a Staff Software Engineer on Docker's infrastructure team, you will build the foundations for a multi-region, cross-account network architecture and continuous-deployment flow that enables teams to spin up new global regions or application environments in hours instead of days. You'll create self-service systems with clear ownership, safe defaults, and strong guardrails to replace expert-driven manual workflows. This role involves staying h...

Why This Role?

Set technical direction for Docker's internal platform as it scales to support hundreds of engineers and high-scale production workloads.

Key Responsibilities

  • Design and implement a multi-region, cross-account network architecture for internal platform services
  • Build and maintain a trusted testing and continuous-deployment flow for engineering teams
  • Develop self-service systems with clear ownership, safe defaults, and measurable adoption
  • Stay hands-on in the codebase while setting technical direction for a growing infrastructure team
  • Align multiple development teams on platform standards and adoption goals

Requirements

  • Experience designing multi-region, cross-account network architectures
  • Background in building continuous-deployment and testing flows for large engineering teams
  • Proven ability to create self-service platforms with strong guardrails and measurable adoption
  • Hands-on software engineering skills in infrastructure or platform systems
  • Experience aligning cross-functional teams on technical direction and production adoption

Required Skills

software engineeringinfrastructureplatform engineeringmulti-region architecturecontinuous deploymentinfrastructure engineeringnetwork architectureplatform developmenttechnical leadership

Indonesia Context

Working Hours Overlap:
Flexible — work your own hours
See remote (USD) vs local pay →

Keywords

Staff EngineerInfrastructureMulti-regionSelf-service platformDockerPlatform engineeringContinuous deploymentNetwork architecture
View Original Description from RemoteOK

Original description from RemoteOK

Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout. We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default. Docker is shipping a wave of new products this year, with R&D initiatives likely to lead to more, and we're investing heavily in the platform underneath all of it. That platform supports hundreds of engineers across many development teams and carries high-scale production traffic and data transfer every day. It has grown faster than its foundations, and this year is about closing that gap. Today, much of that work still leans on a handful of experts unblocking the same provisioning and operational workflows by hand. The top priority for this role is moving that work from expert-driven support to paved roads : self-service systems with clear ownership, safe defaults, strong guardrails, and adoption we can measure. The goal is a platform teams trust enough to stop thinking about it, one that just works, so they can focus on their own products instead of ours. The concrete version sits on this year's roadmap: spinning up a new global region or application environment should take hours, not days. Right now it takes days. Getting there means building the foundations underneath it. We need a real multi-region, cross-account network architecture and a testing and continuous-deployment flow teams can trust, then a self-service layer on top. We're the container company building our own internal platform, so the bar for "the easy path is also the safe path" is high. You'd be joining a team of four, growing to seven this year (this is one of those hires), and we're looking for a Staff engineer to set technical direction and lead it through real production adoption. Responsibilities This is a Staff-level role, so success is measured by leverage rather than just your own commits. On a team this size you'll stay hands-on in the codebase while also setting direction, aligning teams on pragmatic standards, and carrying platform investments through to adoption. Concretely, you will: Take ambiguous infrastructure problems and turn them into proposals the org can rally around, then drive them through RFCs and architecture reviews across teams. Design self-service capabilities and platform APIs (primarily in Go ) for onboarding, provisioning, deployment, observability defaults, and day-2 operations, with contracts and docs teams actually use. Set delivery standards using Terraform , GitOps with Argo CD , progressive rollout, and good testing, including building the continuous-deployment flow we're missing today. Evolve the multi-tenant EKS foundations toward better reliability, security, scale, and cost: Envoy Gateway ingress, traffic routing, and the multi-region, cross-account connectivity we need. Improve SLOs, alerting, and incident follow-up on Grafana Cloud so production gets safer and less dependent on heroics. We judge this work by outcomes the consuming teams feel: how fast they can provision and ship, how much they can do without us, and how reliably it all runs. AI-assisted operations We're actively investing in AI-assisted and agentic workflows to cut operational toil. We care that they stay safe, auditable, and human-reviewed. You'll help shape where these earn their place and where they don't. Early targets include: Alert enrichment and incident context-gathering : assembling the relevant signals, history, and runbook so the on-call engineer st

Source site may be blocked by Indonesian ISPs

Some Indonesian ISPs (Telkomsel, Indihome) block RemoteOK. If the Apply button doesn't open, try mobile data or a VPN.

Tip: switch network or enable a VPN, then click Apply again.

Openness not stated by employer — check the listing

Company
Docker
Source
RemoteOK
Job Type
full time
Location
Remote · Open worldwide
Category
Seniority
senior
Posted
Jun 8, 2026

Share this job

Help a friend find their next remote role.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.