Applicantz

Site Reliability Engineer

Applicantz

Singapore · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
há 3 horas
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

The Site Reliability Engineer (SRE) is responsible for maintaining the dependability, availability, and efficiency of system and platform services by integrating engineering practices with operational expertise. The SRE leverages software engineering skills to automate, monitor, and analyze data, thereby enhancing reliability while supporting swift development cycles.

Key Responsibilities

  • Take full end-to-end accountability for system reliability, availability, and performance, guided by clearly defined service level agreements, objectives, and indicators, utilizing continuous monitoring and proactive improvements to uphold service health.
  • Develop and implement error budget policies collaboratively with engineering leadership to strike a balance between release speed and system reliability, using these budgets to guide priorities and release readiness assessments.
  • Lead response efforts during significant and complex incidents, engage actively in customer-impacting situations, and facilitate blameless postmortems to ensure rapid implementation of systemic corrective measures.
  • Standardize and upgrade observability across different environments by establishing strong monitoring, logging, and tracing frameworks through tools like Dynatrace, CloudWatch, and OpenTelemetry.

Technical Expertise Required

  • Advanced proficiency with core AWS services including EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, and IAM.
  • Competence in designing architectures that are highly available across multiple availability zones and regions.
  • Strong knowledge of AWS networking concepts such as subnets, routing tables, NAT, security groups, network ACLs, VPC peering, and PrivateLink.
  • Experience applying well-architected framework pillars focusing on reliability, security, and cost efficiency.
  • Design and operate fault-tolerant, horizontally scalable systems.
  • Expertise with Infrastructure as Code tools like Terraform, CloudFormation, or CDK including modular design patterns and best practices for state management.
  • Hands-on experience with monitoring and observability tools including CloudWatch, Prometheus, Grafana, Datadog, Dynatrace, or OpenTelemetry.

Organizational Context

SREs act as reliability custodians and domain experts, collaborating with both platform and product engineering teams on SRE and DevOps functions. This structure is supported by a Senior Principal SRE who aligns organizational goals, standardizes practices, and ensures consistency across teams.

Tools & software

Amazon DynamoDB required Amazon Elastic Compute Cloud EC2 required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help