P

AIDC Delivery & Operations Director

PaleBlueDot AI

Singapore · Full Time

Be the first to apply

Experience
10+ yrs
Salary
Openings
1
Posted
2 કલાક પેહલા
Work mode
In office
Education
Bachelor's degree or higher
Resume
Required to apply

Where you'll work

Job description

Overview

PaleBlueDot AI seeks an experienced leader to oversee the design, deployment, and operation of large-scale AI compute clusters overseas. This position is responsible for managing end-to-end delivery of 10,000+ GPU AI infrastructure, including architectural planning, execution of deployments, operational readiness, customer technical collaboration, and ongoing performance optimization.

Key Responsibilities

  • Lead architecture design for international AI clusters covering compute, networking, storage, and liquid cooling systems.
  • Convert customer AI workload requirements into scalable, reliable, and high-performance solutions.
  • Manage the full deployment lifecycle: planning, procurement, acceptance, installation, cabling, commissioning, and launching into production.
  • Coordinate execution among data center facility teams, hardware vendors, and cooling partners.
  • Build and direct a high-performing overseas operations team, including 24/7 operation protocols, standard operating procedures, and incident response frameworks.
  • Implement monitoring, alerting, logging, and performance analysis tools to ensure comprehensive system observability and stability.
  • Serve as the senior technical escalation contact for critical incidents, leading troubleshooting, root cause analyses, and continuous improvements.
  • Engage directly with customer technical teams to present solutions, support proofs of concept, and conduct in-depth technical discussions.
  • Ensure services consistently achieve or surpass defined SLA standards.
  • Continuously enhance cluster efficiency focusing on power usage effectiveness (PUE), water usage effectiveness (WUE), compute utilization, and operating expenses.

Qualifications

  • Bachelor’s degree or higher in Computer Science, Electronic Engineering, or related fields.
  • At least 10 years of experience in operations and management of large-scale data centers, HPC, or AI cluster environments.
  • Demonstrated success in delivering advanced AI compute clusters or hyperscale data center projects internationally.
  • Deep technical expertise with AI cluster architectures such as NVIDIA DGX, SuperPOD, or GPU-as-a-Service models.
  • Strong grasp of networking technologies including InfiniBand, RoCE, distributed storage systems, and large-scale infrastructure design.
  • Practical experience deploying or managing immersion or cold plate liquid cooling systems.
  • Proficient in Linux, Slurm workload manager, Kubernetes orchestration, Prometheus, Grafana monitoring, Ansible automation, and Python programming.
  • Minimum 5 years leadership experience overseeing technical teams, adept at managing cross-cultural groups.
  • Excellent communication skills with strong customer interaction abilities; fluency in English preferred.

Why Join Us

  • Opportunity to create state-of-the-art AI infrastructure at a global level.
  • Contribute to critical AI computing environments with significant impact.
  • Become part of a dynamic team dedicated to advancing AI infrastructure innovation.

Work styles they’re looking for

Leadership Problem Solving Customer Communication Operational Excellence Cross-cultural Collaboration

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help