R

Principal Software Engineer - DevOps / Site Reliability Engineer

Riot Games

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
2 तासपूर्वी
Work mode
In office
Education
Bachelor's degree in Computer Science or related field or equivalent experience
Eligibility
Qualified professionals with relevant education and experience; gaming passion is an advantage but not mandatory.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Riot Games and Role Overview

Founded in 2006 by passionate gamers, Riot Games focuses on player-centric game development, culminating in titles like League of Legends, which currently boasts over 100 million monthly players. We seek driven and sharp professionals who are passionate about gaming and innovation. This role is for a Principal DevOps / Site Reliability Engineer on the AI Efficiency team, responsible for building and maintaining the operational platforms that facilitate safe and scalable AI use across the company’s products and workflows.

Key Responsibilities

  • Maintain and enhance the reliability, scalability, availability, and performance of AI-related platforms and internal tools.
  • Build and manage infrastructure, deployment pipelines, and operational fundamentals to support expanding AI services.
  • Optimize CI/CD systems, release engineering, and deployment automation for rapid and secure software delivery.
  • Set production readiness standards including monitoring, alerting, ownership, documentation, rollback, and support methodologies.
  • Define and use service health metrics such as SLIs, SLOs, error budgets for guiding engineering priorities.
  • Create observability frameworks using metrics, logs, tracing, dashboards, synthetic monitoring, and alerts.
  • Develop and maintain sustainable on-call, incident management, and escalation processes.
  • Lead or assist incident diagnosis and resolution, coordinating post-incident reviews and improvements.
  • Develop automation to reduce manual toil and detect and recover from failures swiftly and reliably.
  • Implement safe deployment techniques including canary releases, feature flags, rollback strategies, and environment promotion.
  • Conduct capacity planning, load testing, performance evaluations, and forecasting for growing user adoption.
  • Design resilience and disaster recovery strategies for crucial services and infrastructure.
  • Identify and mitigate single points of failure and systemic operational risks across services and infrastructure.
  • Enhance developer productivity through self-service workflows, infrastructure reuse, dev and test environments, and clear operational documentation.
  • Implement infrastructure-as-code, configuration and secrets management, and environment governance for safe and repeatable changes.
  • Collaborate with cross-functional teams to embed reliability, security, and maintainability in system design.
  • Troubleshoot complex production issues involving applications, APIs, distributed services, containers, cloud infrastructure, authentication, and third-party dependencies.
  • Partner with ML platform engineers to standardize model serving and inference within platform reliability practices.
  • Work closely with infrastructure, security, IT, platform, and compliance teams to ensure operational and security compliance.
  • Evaluate and deploy AI-driven operational automation like anomaly analysis and runbook automation.
  • Define governance, approvals, audit trails, and escalation mechanisms for automated operational systems interacting with production.
  • Provide technical leadership by mentoring, creating standards, conducting architecture reviews, and building tools to promote operational excellence.

Required Qualifications

  • Bachelor’s degree in Computer Science or a related field, or equivalent experience.
  • Minimum 5 years in roles such as Site Reliability, DevOps, Infrastructure, Platform, or Production Engineering supporting live systems.
  • Proficiency in programming and automation with languages like Python, Go, JavaScript, or TypeScript.
  • Experience managing cloud production environments (AWS, GCP, Azure, or others).
  • Building and maintaining CI/CD pipelines, deployment automation, and environment management workflows.
  • Deep understanding of observability including monitoring, logging, distributed tracing, dashboarding, and alerting.
  • Experience with incident response, on-call duties, root cause analysis, and continuous improvement.
  • Expertise improving reliability, scalability, and performance of distributed systems, service-oriented architectures, or web platforms.
  • Strong knowledge of containerization and orchestration tools like Docker, Kubernetes, or ECS.
  • Hands-on with infrastructure-as-code and config management tools like Terraform, Pulumi, or CloudFormation.
  • Familiarity with Linux, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security basics.
  • Ability to recognize systemic risks and implement lasting operational improvements across multiple engineering teams.
  • Excellent collaboration skills, technical influence, and clear communication under normal and crisis conditions.
  • Experience leading engineers, mentoring, and establishing engineering norms and practices at team or organizational level.

Desired Qualifications

  • Background supporting AI/ML platform services, inference systems, GPU workloads, or compute-intensive pipelines.
  • Expertise defining and leveraging SLIs, SLOs, and error budgets to balance engineering priorities.
  • Experience creating or enhancing internal developer platforms, self-service infrastructure, or standardized engineering workflows.
  • Track record establishing sustainable on-call rotations and operational support for multi-team services.
  • Familiar with advanced deployment mechanisms like progressive delivery, canary and blue-green deployments, feature flags, and automated rollbacks.
  • Experience with resilience testing techniques including chaos engineering, fault injection, and failure analysis.
  • Architecting backup, disaster recovery, business continuity, and regional failover solutions.
  • Operating databases, caches, queues, object stores, service meshes, and API gateways effectively.
  • Improving system security through access controls, secrets management, dependency verification, vulnerability handling, network segmentation, and infrastructure hardening.
  • Balancing availability, latency, velocity, infrastructure efficiency, and cost at scale.
  • Knowledge of browser automation and end-to-end testing with tools like Playwright for validating critical user flows and regression detection.
  • Experience assessing and integrating AI-powered tools for incident management, anomaly detection, code review, test generation, or automated remediation.
  • Understanding risks and governance when AI agents interact with source control, CI/CD, cloud infrastructure, or production systems.
  • Proficiency establishing controls such as approvals, audit trails, and quality checks for automated operational tools.
  • Experience transitioning experimental technology into robust, supported production services.

Eligibility

Open to qualified professionals meeting the education and experience criteria; enthusiasm for gaming is a bonus but not required.

Benefits

  • Full relocation assistance provided.
  • Comprehensive health coverage inclusive of employee, spouse, and children.
  • Flexible paid time off policies.
  • Retirement plans with company matching contributions.
  • Life insurance, parental leave, and both short-term and long-term disability coverage.
  • Access to Play Fund to engage deeper with gaming community and products.
  • Company doubles employee charitable donations of time and money.

Level

Lead

Minimum education

Bachelor's Degree

How they work

Communication Teamwork & Collaboration Problem Solving Leadership

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help