Member of Technical Staff, Site Reliablity Engineer
Build reliability culture and improve call completion
Join the oncall rotation, analyze stability-gap incidents, and define SLOs for the call-completion path. Stand up error budgets and SLO-based alerting, run load tests, and tune autoscaling. Ship a platform service like capacity forecaster or auto-remediation in Go or TypeScript.
Why This Role?
Direct founder access, real impact from day one
Key Responsibilities
- Join the oncall rotation and analyze stability-gap incidents to create a reliability backlog
- Define the first set of SLOs for the call-completion path
- Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus
- Run load tests against provider rate limits and per-org concurrency
- Tune autoscaling for wscaler and workerpool-cron-scaler
- Ship a real platform service in Go or TypeScript
Requirements
- Experience running incident command and postmortem discipline at scale
- Experience operating SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog
- Experience in capacity planning and load testing for production systems
- Fluency in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown
- Knowledge of backpressure and autoscaling patterns: KEDA, custom metrics scaling
Required Skills
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
Voice AI that resolves, not transfers. Most phone systems trap callers in menus and scripts. Vapi is the platform for deploying voice agents that know your business and can listen, adapt, and resolve in minutes. - The numbers: 1 billion calls. 1 million developers. 10x enterprise ARR growth - The customers: Amazon Ring, ServiceTitan, New York Life, Intuit, Kavak, and thousands more, from YC startups to the Fortune 500 - The news: a $50M Series B led by Peak XV Partners, with Bessemer Venture Partners, Kleiner Perkins, M12 (Microsoft's Venture Fund), Y Combinator, and our earlier backers. Total raised: $72M WHY WE’RE HIRING THIS ROLE: - 99.99% call completion is the number this role drives. Vapi runs live phone calls — a p99 spike means callers drop. We’ve had 15 stability-gap outages worth learning from, and we need someone who runs incident command, owns SLOs and error budgets, and builds the reliability culture from scratch. - This is not a bash-and-YAML role. You’ll ship code (Go or TypeScript) for services that monitor and manage the platform: auto-remediation, capacity forecasters, oncall tooling. Capacity planning, load testing, and KEDA-based autoscaling for Vapi’s wscaler and workerpool-cron-scaler are on your plate. WHAT YOU’LL DO: - 30 Day: Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path. - 60 Day: Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler. - 90 Day: Ship a real platform service — capacity forecaster, auto-remediation, or oncall tooling — in Go or TypeScript. Own the postmortem process. Drive a measurable improvement in p99 call completion or MTTR. WHO YOU ARE: Must-haves - You’ve run incident command and postmortem discipline at scale on a real oncall rotation. - You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog. - You’ve done capacity planning and load testing for production systems with real users. - You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown. - You know backpressure and autoscaling patterns — KEDA, custom metrics scaling. Nice-to-haves - You ship code, not just scripts. You can build platform services in Go or TypeScript (matches Vapi’s cluster-manager, database-health, wscaler, incidentManager). - Real-time / latency-sensitive product background where degraded means a dropped call, not a slow dashboard. Tech stack you’ll work in - Languages: Go and TypeScript (you ship code, not just scripts), Bash. - Observability: Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry. - Orchestration: Kubernetes on EKS — production ops (HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown, pod crash diagnosis). - Autoscaling and backpressure: KEDA, custom metrics scaling (matches Vapi’s wscaler and workerpool-cron-scaler). - Load testing: script-based load testing, provider rate-limit auditing, per-org concurrency auditing. - Vapi services you’ll touch or build: cluster-manager, database-health, wscaler, incidentManager. Where you likely come from - A real-time / latency-sensitive product (Discord, Zoom, Mux, Twitch, Twilio, LiveKit, Cloudflare, a trading firm, a gaming backend), or a FAANG SRE / Production Engineer (Google, Uber, Twitter/X, Meta) who misses being hands-on. - Weak fit: SRE from analytics or CRM backends where “degraded” means a slow dashboard, not a dropped call. Anyone uncomfortable reading or writing code. WHY VAPI: - Generational impact: Build the human interface for every business - Ownership culture: 70% of the company are previous founders - Kind team: The founders, Jordan and Nikhil, are Canadians - Tier-1 Investors: YC, KP seed, Bessemer Series A WHAT WE OFFER: - Real stake: We offer a competitive salary and excellent equity ownership - Comprehensive health coverage: medical, dental, and vision plans - Team love: We love hanging out, and we do quarterly off-sites - Flexible time off: take what you need More: catered meals, transportation, gym, and a $10k annual L&D budget
Salary Context
Similar Engineering roles on LokerDollar pay around $207.5k/yr (range $19.575k–2846.004k/yr, n=455 active listings).
Hiring at Vapi
Vapi has 18 other active roles on LokerDollar and has been hiring here since Jun 23, 2026 — across Engineering, Operations, Product.
View all Vapi openings →Openness not stated by employer — check the listing
Frequently asked questions
- Is Member of Technical Staff, Site Reliablity Engineer at Vapi a remote job?
- This role is based in Remote. See the listing for remote/onsite details.
- What is the salary for Member of Technical Staff, Site Reliablity Engineer at Vapi?
- The listed pay range for this role is $200k–270k/yr.
- What type of employment is Member of Technical Staff, Site Reliablity Engineer at Vapi?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Vapi.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Remote Radiology Jobs: August 2026 UpdateDiscover the latest remote radiology jobs and salary trends for August 2026. Learn about the opportunities and challenges of remote work in radiology.
- Unlocking Remote USD Jobs: A GuideDiscover the best remote USD-paying jobs and learn how to succeed in the global job market.
- Funding Down 43%, But Global Remote Jobs AreGlobal startup funding dipped 43% in H1 2026. Yet, top-tier global companies are aggressively hiring remote talent worldwide, paying in USD.