Senior Site Reliability Engineer - India
Define reliability targets and build systems to keep Syfe's platform stable
As a Senior Site Reliability Engineer at Syfe, you will own the end-to-end reliability of the production platform, defining measurable reliability targets through SLIs/SLOs and error budgets. You will build and mature the observability stack using Datadog, Grafana, VictoriaMetrics, and ClickHouse, while reducing operational toil through automation and self-service tooling. You will also lead incident response, on-call programs, and resilience ...
Why This Role?
Own and lead the on-call and incident-response program for a regulated digital wealth platform
Key Responsibilities
- Define and drive SLIs/SLOs and error budgets across critical services in partnership with product and engineering teams
- Own the on-call rotation, escalation policies, and paging strategy, establishing incident command and running blameless postmortems
- Build and mature the observability stack using Datadog for APM/RUM and Grafana, VictoriaMetrics, and ClickHouse for metrics and logs
- Lead capacity planning, scalability analysis, failure-mode analysis, disaster recovery, and chaos exercises across regions
- Identify and eliminate operational toil through automation and self-service tooling to improve platform reliability
Requirements
- Experience with Kubernetes-native platforms, specifically EKS-based deployments
- Proficiency in GitOps delivery using ArgoCD and Helm-based release configuration
- Experience with Infrastructure as Code using Terraform or OpenTofu on AWS
- Background in observability tools including Datadog, Grafana, VictoriaMetrics, and ClickHouse
- Experience defining SLIs/SLOs, managing error budgets, and driving reliability improvements
Required Skills
Indonesia Context
- Working Hours Overlap:
- Flexible — work your own hours
Keywords
View Original Description from Manatal Career Pages
Original description from Manatal Career Pages
Syfe is APAC's largest and fastest-growing digital wealth platform , trusted with over US$10 billion in assets. We are fundamentally changing how hundreds of thousands of people across Asia-Pacific build wealth through a holistic approach to managing money rather than just pushing investment products. Backed by world-class investors and recognised as a leader in wealthtech, we are a team of passionate builders creating the future of wealth management. About the Role : We are looking for a Senior Site Reliability Engineer to own the reliability of Syfe's production platform end-to-end. Syfe runs a Kubernetes-native, multi-region platform (Singapore, Hong Kong, Sydney) serving a regulated digital wealth- management product, where availability, latency, and trust are first-order product features. This is a senior individual-contributor role, not a managerial one. You will define what "reliable" means in measurable terms, build the systems and automation that keep us there, and own and lead our on-call and incident-response program. You'll spend your time engineering reliability into the platform — through SLOs, observability, and automation — rather than firefighting, and you'll raise the bar for how the whole engineering org operates production What You'll Own Reliability targets. Define and drive SLIs/SLOs and error budgets across critical services; partner with product and engineering teams to make error-budget-based decisions that balance velocity and stability. On-call & incident program. Own the on-call rotation, escalation policies, and paging strategy. Establish incident command, run blameless postmortems, and turn RCAs into tracked, completed reliability work. Drive down MTTD and MTTR. Platform & Kubernetes reliability. Own the reliability of our EKS-based deployment platform — GitOps delivery (ArgoCD), Helm-based release configuration, and Infrastructure as Code (Terraform/OpenTofu) on AWS. Make deployments safe, progressive, and reversible. Observability. Build and mature the observability stack (Datadog for production APM/RUM; Grafana, VictoriaMetrics, and ClickHouse for metrics and logs). Make systems debuggable: meaningful dashboards, actionable alerts, and low alert noise. Resilience. Lead capacity planning, scalability, failure-mode analysis, disaster-recovery and business-continuity planning, and game-day / chaos exercises across regions. Toil reduction. Identify operational toil and eliminate it with automation and self-service tooling, so reliability scales with the platform rather than with headcount. Production safety. Strengthen rollout/rollback paths, deployment guardrails, and secrets handling (HashiCorp Vault), and partner with engineering teams to harden services before they reach production. Engineering influence. Lead by example with hands-on engineering — design reviews, production-readiness reviews, runbooks, documentation, and mentoring — embedding SRE practices across the org Qualifications (Must-have): Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 4–8 years in SRE, platform engineering, or DevOps, with a strong senior IC track record of owning production systems. Production-grade expertise with Kubernetes and containers, and a cloud platform (AWS preferred) in a distributed- systems environment. Hands-on experience defining and operating against SLIs/SLOs and error budgets, and leading incident response and blameless postmortems. Strong with observability tooling — metrics, logging, tracing, dashboards, and alerting (e.g., Datadog, Prometheus/Grafana, or equivalents). Proficient writing automation and infrastructure code (e.g., Python, Go, Shell, Terraform) and comfortable with GitOps / CI/CD delivery. Solid grasp of Linux/Unix internals, networking, and cloud-native security fundamentals. Strong operational rigor, ownership, and clear written/verbal communication, including during high-pressure incidents. Nice to Have: Experience operating in regulated or high-trust environments (fintech, payments, etc.). Experience running multi-region / multi-cluster Kubernetes at scale. Familiarity with ArgoCD, Helm/Helmfile, OpenTofu, Vault, Cloudflare, or comparable tooling. Experience building self-service developer platforms or internal reliability tooling. Contributions to open-source projects or public technical content (GitHub, blogs, talks). Relevant cloud or Kubernetes certifications (e.g., AWS, CKA/CKS) Come As You Are We believe in the power of diversity and are dedicated to creating a welcoming and innovative environment for all our employees. So we embrace and encourage applications from candidates of all backgrounds and provide equal employment opportunities for all. Due to the volume of applications, we regret that only shortlisted candidates will be notified.
Salary Context
Similar Engineering roles on LokerDollar pay around $195k/yr (range $36k–300k/yr, n=258 active listings).
Hiring at Syfe
Syfe has 21 other active roles on LokerDollar and has been hiring here since Aug 10, 2026 — across Engineering, Marketing, Product.
- Senior Wealth Advisor - Singapore
- Product Manager (Brokerage) - India
- Senior Product Manager – Australia
Frequently asked questions
- Is Senior Site Reliability Engineer - India at Syfe a remote job?
- This role is based in Remote. See the listing for remote/onsite details.
- What type of employment is Senior Site Reliability Engineer - India at Syfe?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Syfe.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Remote Jobs with Clear Salaries in 2026Explore remote job opportunities with transparent salaries in 2026, including positions at Automation Anywhere, Binance, Bree, and more.
- Freelance Remote USD 2026: Apa yang PerluCari kerja remote USD? Simak seluk-beluk gig economy, rate terkini, dan tips jadi freelancer sukses di Agustus 2026.
- EnableComp: Remote DRG VP & the RiseEnableComp is hiring a remote VP DRG. Salary is undisclosed, but what does this mean for the growing trend of overemployment? We dive in.