Skip to main content
Back to Jobs

G13 - Operations Support Engineer

Design service observability usage model ensuring metrics, logs, traces flow into Elastic Cloud

Design and own the service observability usage model by ensuring all service metrics, logs, and traces flow into Elastic Cloud as the authoritative source. Maintain dashboards and SLOs, evaluate supplemental tooling like CloudWatch and AWS Managed Prometheus/Grafana, and build proactive, noise-reduced alerting and incident response playbooks. Drive post-incident RCA and remediation tracking with closure SLAs, optimize service performance throu...

Why This Role?

Mentor engineers on production readiness and observability patterns while evolving on-call playbooks

Key Responsibilities

  • Design and maintain service observability in Elastic Cloud ensuring metrics, logs, and traces are properly ingested
  • Build noise-reduced alerting and incident response playbooks, drive RCA and remediation tracking
  • Optimize service performance via profiling, caching, autoscaling heuristics, and concurrency tuning
  • Implement secure supply chain and runtime controls including image scanning, SBOM, secrets management, and TLS/mTLS
  • Curate operational runbooks, golden dashboards, and reliability/production readiness checklists
  • Integrate model/guardrail service telemetry into unified Elastic Cloud views and support compliance evidence collection

Requirements

  • 4+ years of experience in SRE, Production Ops, Platform, or Reliability for SaaS or high-throughput services
  • Working knowledge of AWS and Kubernetes for deployment, troubleshooting, and networking concepts
  • Familiarity with Infrastructure as Code and GitOps tools like Terraform and Argo for consuming modules and reviewing changes
  • Observability implementation and usage with Elastic Cloud, including metrics, logs, traces, and Prometheus/OpenTelemetry concepts
  • Proven on-call and incident management experience with triage, MTTR reduction, and RCA authorship
  • Scripting and automation skills in Python, Bash, or Go for ops tooling

Required Skills

awskuberneteselastic-cloudobservabilitypythonbashgoincident responseperformance optimizationsecurity compliancescripting
View Original Description from Manatal Career Pages

Original description from Manatal Career Pages

Responsibilities: Design & own service observability usage model: ensure all service metrics, logs, traces flow into Elastic Cloud (authoritative); maintain dashboards & SLOs; evaluate pragmatic use of CloudWatch, AWS Managed Prometheus / Grafana for supplemental or fallback views. Build proactive, noise‑reduced alerting and incident response playbooks; drive post‑incident RCA & remediation tracking (closure SLA). Optimize service performance (profiling, caching layers, autoscaling heuristics, concurrency tuning) meeting latency & throughput targets. Implement secure supply chain & runtime controls (image scanning, SBOM consumption, secrets management, TLS / mTLS) leveraging shared platform tooling. Curate operational runbooks, golden dashboards, reliability readiness + production readiness checklists. Integrate model / guardrail service telemetry (latency, queue depth, GPU/CPU utilization) into unified Elastic Cloud views. Support compliance & audit evidence collection (access logs, config lineage, change histories) via automated evidence capture fed into Elastic. Introduce configuration drift detection & policy-as-code guardrails (OPA / Kyverno) at the workload / namespace layer to enforce baseline controls. Mentor engineers on production readiness, observability patterns, and operational excellence; evolve on-call playbooks. Participate in (and improve) an equitable on-call rotation focusing on sustainable alert volumes & burnout prevention. Requirements 4+ years (or equivalent impact) in SRE / Production Ops / Platform / Reliability for SaaS or high-throughput services. Working knowledge of AWS & Kubernetes (deployment, troubleshooting, networking concepts) sufficient to collaborate effectively with platform owners (not necessarily owning cluster upgrade orchestration). Familiarity with Infrastructure as Code & GitOps (Terraform, Argo, etc.) to consume modules, review changes, and enforce policy. Observability implementation & usage (metrics, logs, traces, profiling) with Elastic Cloud; understanding of Prometheus / OpenTelemetry concepts. Proven on-call & incident management experience (triage, MTTR reduction, RCA authorship). Scripting / automation in Python, Bash, or Go for ops tooling. Security & compliance aware: vulnerability management, image scanning, supply chain controls. Clear, concise communication of operational risk & trade-offs to technical + non-technical stakeholders.

Salary Context

Similar Engineering roles on LokerDollar pay around $160k/yr (range $1.8k–999.999k/yr, n=629 active listings).

Hiring at FPT Asia Pacific

FPT Asia Pacific has 108 other active roles on LokerDollar and has been hiring here since Aug 10, 2026 — across Engineering, Operations, Design.

View all FPT Asia Pacific openings →

We couldn't verify this apply link yet. You can still try applying — check back later if it doesn't open.

Openness not stated by employer — check the listing

Source
Manatal Career Pages
Job Type
full time
Location
Singapore, Singapore · Remote
Category
Seniority
senior
PostedFreshNew & verified
Aug 10, 2026

Share this job

Help a friend find their next remote role.

Frequently asked questions

Is G13 - Operations Support Engineer at FPT Asia Pacific a remote job?
This role is based in Remote. See the listing for remote/onsite details.
What type of employment is G13 - Operations Support Engineer at FPT Asia Pacific?
This is a full time position.
How do I apply?
Click the "Apply" button on this page to go to the official application at FPT Asia Pacific.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.

From the blog