KAUST (King Abdullah University of Science and Technology)

Lead HPC System Administrator

KAUST (King Abdullah University of Science and Technology)

Thuwal, Makkah Province, Saudi Arabia · Full Time

Be the first to apply

Experience
5–10 yrs
Salary
Openings
1
Posted
11 jam yang lalu
Work mode
In office
Education
Bachelor's degree or higher in relevant field
Resume
Required to apply

Where you'll work

Job description

Overview

We are looking for a Lead System Administrator to manage the Linux cluster comprising over 300 GPU/CPU compute nodes, parallel filesystems, and high-performance networks at KAUST. This role is a blend of technical expertise and leadership, requiring supervision of a team of 3-4 seasoned HPC administrators. The lead will be responsible for developing, implementing, and overseeing standard operating procedures for both the computing environment and the team.

Key Responsibilities

  • Plan and execute system operations and upgrades to align with research laboratory and client needs.
  • Develop and enforce workload scheduler policies.
  • Maintain and support high-performance parallel filesystems.
  • Manage network infrastructure encompassing TCP/IP and HPC-specific networking technologies.
  • Automate node configuration and management using scripting languages.
  • Coordinate hardware maintenance, including managing failures and spare parts inventory.
  • Foster strong collaboration with staff, faculty, and students through core laboratory partnerships.
  • Manage multiple projects using advanced project planning methods.
  • Schedule and oversee detailed project phases and moderate-scope total projects.
  • Identify and facilitate technical training for team members.
  • Respond as a key resource for security and safety incident management.
  • Drive innovation by developing new technical methods and expanding applications of existing technologies.
  • Provide strategic contributions and maintain customer relationships.
  • Initiate and propose new project concepts; engage with prospective clients through presentations.
  • Supervise scientific and technical staff, contributing significantly to team staffing and development.
  • Evaluate candidates for open roles and mentor staff in technical proficiency, project management, and business development.

Required Competencies

  • In-depth knowledge of SLURM workload management with GPU scheduling capabilities.
  • Experience with parallel filesystem technologies such as Weka IO and Lustre.
  • Proficiency in TCP/IP and high-performance networks including Infiniband.
  • Skilled in scripting languages like Bash, Python, and Ruby for automation and configuration.
  • Familiar with configuration management tools such as Puppet.
  • Strong documentation and communication skills in English, both written and oral.
  • Ability to work closely with users and suppliers, demonstrating systematic problem-solving approaches.
  • Experience collaborating within multi-disciplinary research environments.
  • Ability to manage personal and team workloads effectively under deadline-driven conditions.
  • Discretion and versatility in handling complex and non-routine technical assignments.
  • Maintains expert-level knowledge in HPC systems administration, storage, or networking.

Qualifications and Experience

  • Bachelor's degree or equivalent in relevant field with a minimum of 10 years experience, OR
  • Master's degree with at least 7 years relevant experience, OR
  • PhD with a minimum of 5 years experience in a related area.

Work styles they’re looking for

Communication Leadership Problem Solving Team Collaboration Initiative

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help