KAUST (King Abdullah University of Science and Technology)

Lead HPC System Administrator

KAUST (King Abdullah University of Science and Technology)

Thuwal, Makkah Province, Saudi Arabia · Full Time

Be the first to apply

Experience
5–10 yrs
Salary
Openings
1
Posted
4 કલાક પેહલા
Work mode
In office
Education
Bachelor's degree or higher in relevant field
Resume
Required to apply

Where you'll work

Job description

Overview

We are looking for a Lead System Administrator to manage the Linux cluster comprising over 300 GPU/CPU compute nodes, parallel filesystems, and high-performance networks at KAUST. This role is a blend of technical expertise and leadership, requiring supervision of a team of 3-4 seasoned HPC administrators. The lead will be responsible for developing, implementing, and overseeing standard operating procedures for both the computing environment and the team.

Key Responsibilities

  • Plan and execute system operations and upgrades to align with research laboratory and client needs.
  • Develop and enforce workload scheduler policies.
  • Maintain and support high-performance parallel filesystems.
  • Manage network infrastructure encompassing TCP/IP and HPC-specific networking technologies.
  • Automate node configuration and management using scripting languages.
  • Coordinate hardware maintenance, including managing failures and spare parts inventory.
  • Foster strong collaboration with staff, faculty, and students through core laboratory partnerships.
  • Manage multiple projects using advanced project planning methods.
  • Schedule and oversee detailed project phases and moderate-scope total projects.
  • Identify and facilitate technical training for team members.
  • Respond as a key resource for security and safety incident management.
  • Drive innovation by developing new technical methods and expanding applications of existing technologies.
  • Provide strategic contributions and maintain customer relationships.
  • Initiate and propose new project concepts; engage with prospective clients through presentations.
  • Supervise scientific and technical staff, contributing significantly to team staffing and development.
  • Evaluate candidates for open roles and mentor staff in technical proficiency, project management, and business development.

Required Competencies

  • In-depth knowledge of SLURM workload management with GPU scheduling capabilities.
  • Experience with parallel filesystem technologies such as Weka IO and Lustre.
  • Proficiency in TCP/IP and high-performance networks including Infiniband.
  • Skilled in scripting languages like Bash, Python, and Ruby for automation and configuration.
  • Familiar with configuration management tools such as Puppet.
  • Strong documentation and communication skills in English, both written and oral.
  • Ability to work closely with users and suppliers, demonstrating systematic problem-solving approaches.
  • Experience collaborating within multi-disciplinary research environments.
  • Ability to manage personal and team workloads effectively under deadline-driven conditions.
  • Discretion and versatility in handling complex and non-routine technical assignments.
  • Maintains expert-level knowledge in HPC systems administration, storage, or networking.

Qualifications and Experience

  • Bachelor's degree or equivalent in relevant field with a minimum of 10 years experience, OR
  • Master's degree with at least 7 years relevant experience, OR
  • PhD with a minimum of 5 years experience in a related area.

Work styles they’re looking for

Communication Leadership Problem Solving Team Collaboration Initiative

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help