- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 घंटे पहले
- Work mode
- In office
- Education
- Bachelor's degree or higher
- Resume
- Required to apply
Where you'll work
Job description
Overview
PaleBlueDot AI seeks an experienced leader to oversee the design, deployment, and operation of large-scale AI compute clusters overseas. This position is responsible for managing end-to-end delivery of 10,000+ GPU AI infrastructure, including architectural planning, execution of deployments, operational readiness, customer technical collaboration, and ongoing performance optimization.
Key Responsibilities
- Lead architecture design for international AI clusters covering compute, networking, storage, and liquid cooling systems.
- Convert customer AI workload requirements into scalable, reliable, and high-performance solutions.
- Manage the full deployment lifecycle: planning, procurement, acceptance, installation, cabling, commissioning, and launching into production.
- Coordinate execution among data center facility teams, hardware vendors, and cooling partners.
- Build and direct a high-performing overseas operations team, including 24/7 operation protocols, standard operating procedures, and incident response frameworks.
- Implement monitoring, alerting, logging, and performance analysis tools to ensure comprehensive system observability and stability.
- Serve as the senior technical escalation contact for critical incidents, leading troubleshooting, root cause analyses, and continuous improvements.
- Engage directly with customer technical teams to present solutions, support proofs of concept, and conduct in-depth technical discussions.
- Ensure services consistently achieve or surpass defined SLA standards.
- Continuously enhance cluster efficiency focusing on power usage effectiveness (PUE), water usage effectiveness (WUE), compute utilization, and operating expenses.
Qualifications
- Bachelor’s degree or higher in Computer Science, Electronic Engineering, or related fields.
- At least 10 years of experience in operations and management of large-scale data centers, HPC, or AI cluster environments.
- Demonstrated success in delivering advanced AI compute clusters or hyperscale data center projects internationally.
- Deep technical expertise with AI cluster architectures such as NVIDIA DGX, SuperPOD, or GPU-as-a-Service models.
- Strong grasp of networking technologies including InfiniBand, RoCE, distributed storage systems, and large-scale infrastructure design.
- Practical experience deploying or managing immersion or cold plate liquid cooling systems.
- Proficient in Linux, Slurm workload manager, Kubernetes orchestration, Prometheus, Grafana monitoring, Ansible automation, and Python programming.
- Minimum 5 years leadership experience overseeing technical teams, adept at managing cross-cultural groups.
- Excellent communication skills with strong customer interaction abilities; fluency in English preferred.
Why Join Us
- Opportunity to create state-of-the-art AI infrastructure at a global level.
- Contribute to critical AI computing environments with significant impact.
- Become part of a dynamic team dedicated to advancing AI infrastructure innovation.
Skills
Work styles they’re looking for
Leadership
Problem Solving
Customer Communication
Operational Excellence
Cross-cultural Collaboration