G13 - Operations Support Engineer
Design service observability usage model ensuring metrics, logs, traces flow into Elastic Cloud
Design and own the service observability usage model by ensuring all service metrics, logs, and traces flow into Elastic Cloud as the authoritative source. Maintain dashboards and SLOs, evaluate supplemental tooling like CloudWatch and AWS Managed Prometheus/Grafana, and build proactive, noise-reduced alerting and incident response playbooks. Drive post-incident RCA and remediation tracking with closure SLAs, optimize service performance throu...
Why This Role?
Mentor engineers on production readiness and observability patterns while evolving on-call playbooks
Key Responsibilities
- Design and maintain service observability in Elastic Cloud ensuring metrics, logs, and traces are properly ingested
- Build noise-reduced alerting and incident response playbooks, drive RCA and remediation tracking
- Optimize service performance via profiling, caching, autoscaling heuristics, and concurrency tuning
- Implement secure supply chain and runtime controls including image scanning, SBOM, secrets management, and TLS/mTLS
- Curate operational runbooks, golden dashboards, and reliability/production readiness checklists
- Integrate model/guardrail service telemetry into unified Elastic Cloud views and support compliance evidence collection
Requirements
- 4+ years of experience in SRE, Production Ops, Platform, or Reliability for SaaS or high-throughput services
- Working knowledge of AWS and Kubernetes for deployment, troubleshooting, and networking concepts
- Familiarity with Infrastructure as Code and GitOps tools like Terraform and Argo for consuming modules and reviewing changes
- Observability implementation and usage with Elastic Cloud, including metrics, logs, traces, and Prometheus/OpenTelemetry concepts
- Proven on-call and incident management experience with triage, MTTR reduction, and RCA authorship
- Scripting and automation skills in Python, Bash, or Go for ops tooling
Required Skills
View Excerpt — Full Description at Manatal Career PagesShow more
Read the full description at Manatal Career Pages
Responsibilities: Design & own service observability usage model: ensure all service metrics, logs, traces flow into Elastic Cloud (authoritative); maintain dashboards & SLOs; evaluate pragmatic use of CloudWatch, AWS Managed Prometheus / Grafana for supplemental or fallback views.…
Openness not stated by employer — check the listing
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- The Compliance Layer of the AI Hiring Stack (2026)6,349 remote listings: 77.2% never state who may apply. Methodology and a CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.