Senior Incident Manager
Lead critical incident response for AI cloud infrastructure
Lead the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services. Coordinate rapid resolution of service-impacting events, improve operational resilience, and drive incident management best practices across infrastructure, networking, platform engineering, and data center operations.
Why This Role?
Direct founder access, real impact from day one
Key Responsibilities
- Lead the response to critical incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations
- Coordinate engineering, networking, facilities, and vendor teams during major outages
- Conduct post-incident analysis to identify patterns and trends for improvement in response and systems reliability
Requirements
- Deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms
- Strong leadership and communication skills
- Experience in incident management and operational resilience
Required Skills
Indonesia Context
- Working Hours Overlap:
- Flexible — work your own hours
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU. If you'd like to build the world's best AI cloud, join us. We are seeking a Senior Incident Manager to lead critical incident response across our AI data center infrastructure. This role is responsible for coordinating rapid resolution of service-impacting events, improving operational resilience, and driving incident management best practices across infrastructure, networking, platform engineering, and data center operations. Role Overview The Senior Incident Manager is responsible for leading the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services. This individual acts as the central command point during major incidents, ensuring rapid triage, cross-team coordination, effective communication, and structured post-incident analysis. This role requires deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms, along with strong leadership and communication skills. What You’ll Do Incident Leadership - Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations. - Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams. - Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries. - Establish clear incident timelines, triage actions, and resolution plans. Incident Management Operations - Own the incident response lifecycle including: - Assisting Technical Triage - Escalation - Coordination - Resolution Post-incident review - Ensure timely and accurate communication with internal stakeholders and leadership. - Maintain incident response documentation and operational playbooks. - Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability. - Work in an On-Call Rotation to respond to, lead, and coordinate incidents Cross-Functional Coordination - Work closely with: - Data center operations - Infrastructure engineering & operations - Network engineering - Platform reliability engineering - Security operations - Hardware and facility vendors - Drive alignment during outages involving multiple infrastructure layers. Post-Incident Analysis & Continuous Improvement - Lead post-incident reviews (PIRs) and root cause analysis. Identify systemic reliability gaps and implement corrective actions. - Track incident metrics including MTTR, MTTD, and incident recurrence rates. Operational Excellence - Improve incident response processes, escalation paths, and tooling by working with technical support and engineering teams.. - Contribute to runbooks, operational standards, and reliability frameworks. - Support implementation of automation and observability improvements. Communication & Reporting - Provide executive-level incident summaries and reports. - Deliver clear, concise updates during active incidents. - Maintain incident dashboards and operational health reporting. You - 8+ years experience in incident management, site reliability engineering, or infrastructure operations - Experience managing incidents in large-scale distributed infrastructure environments - Strong understanding of: - Data center operations - GPU compute clusters Networking and storage infrastructure - Cloud or hybrid infrastructure platforms - Proven ability to lead high-pressure incident response situations - Experience with incident management frameworks (ITIL, SRE, or equivalent) - Excellent communication and stakeholder management skills - Experience with incident tracking and monitoring tools such as: - PagerDuty - ServiceNow - Jira - Datadog - Prometheus / Grafana Nice to Have - Experience operating AI or HPC infrastructure - Background in SRE, infrastructure engineering, or data center operations - Familiarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks) - Experience with hyperscale or colocation data center environments - Knowledge of automation and incident response tooling - Knowledge of and experience with Incident command system (ICS) - Experience in leading and developing incident command from stractch Key Competencies - Incident Command & Leadership - Operational Decision Making - Cross-Team Coordination - Root Cause Analysis - Crisis Communication - Infrastructure Reliability What Success Looks Like in This Role - Reduced Mean Time to Resolution (MTTR) for critical incidents - Improved cross-team incident coordination - High-quality post-incident reviews and corrective actions - Increased infrastructure reliability and operational maturity Salary Range Information The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description. About Lambda - Founded in 2012, with 500+ employees, and growing fast - Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove - We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG - Our values are publicly available: https://lambda.ai/careers - We offer generous cash & equity compensation - Health, dental, and vision coverage for you and your dependents - Wellness and commuter stipends for select roles - 401k Plan with 2% company match (USA employees) - Flexible paid time off plan that we all actually use A Final Note: You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills. Equal Opportunity Employer Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Salary Context
Similar Engineering roles on LokerDollar pay around $206.75k/yr (range $19.575k–2846.004k/yr, n=458 active listings).
Hiring at Lambda
Lambda has 8 other active roles on LokerDollar and has been hiring here since Jun 23, 2026 — across Engineering.
- Senior Site Reliability Engineer - Fleet
- Senior Software Engineer - Managed Kubernetes
- Senior Site Reliability Engineer - Managed Kubernetes
Openness not stated by employer — check the listing
Frequently asked questions
- Is Senior Incident Manager at Lambda a remote job?
- Yes. Senior Incident Manager at Lambda is a fully remote role open to candidates worldwide.
- What is the salary for Senior Incident Manager at Lambda?
- The listed pay range for this role is $125k–195k/yr.
- What type of employment is Senior Incident Manager at Lambda?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Lambda.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Remote Radiology Jobs: August 2026 UpdateDiscover the latest remote radiology jobs and salary trends for August 2026. Learn about the opportunities and challenges of remote work in radiology.
- Unlocking Remote USD Jobs: A GuideDiscover the best remote USD-paying jobs and learn how to succeed in the global job market.
- Funding Down 43%, But Global Remote Jobs AreGlobal startup funding dipped 43% in H1 2026. Yet, top-tier global companies are aggressively hiring remote talent worldwide, paying in USD.