Lead HPC System Administrator
KAUST (King Abdullah University of Science and Technology)
Thuwal, Makkah Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- 5–10 yrs
- Salary
- —
- Openings
- 1
- Posted
- منذ 4 ساعات
- Work mode
- In office
- Education
- Bachelor's degree or higher in relevant field
- Resume
- Required to apply
Where you'll work
Job description
Overview
We are looking for a Lead System Administrator to manage the Linux cluster comprising over 300 GPU/CPU compute nodes, parallel filesystems, and high-performance networks at KAUST. This role is a blend of technical expertise and leadership, requiring supervision of a team of 3-4 seasoned HPC administrators. The lead will be responsible for developing, implementing, and overseeing standard operating procedures for both the computing environment and the team.
Key Responsibilities
- Plan and execute system operations and upgrades to align with research laboratory and client needs.
- Develop and enforce workload scheduler policies.
- Maintain and support high-performance parallel filesystems.
- Manage network infrastructure encompassing TCP/IP and HPC-specific networking technologies.
- Automate node configuration and management using scripting languages.
- Coordinate hardware maintenance, including managing failures and spare parts inventory.
- Foster strong collaboration with staff, faculty, and students through core laboratory partnerships.
- Manage multiple projects using advanced project planning methods.
- Schedule and oversee detailed project phases and moderate-scope total projects.
- Identify and facilitate technical training for team members.
- Respond as a key resource for security and safety incident management.
- Drive innovation by developing new technical methods and expanding applications of existing technologies.
- Provide strategic contributions and maintain customer relationships.
- Initiate and propose new project concepts; engage with prospective clients through presentations.
- Supervise scientific and technical staff, contributing significantly to team staffing and development.
- Evaluate candidates for open roles and mentor staff in technical proficiency, project management, and business development.
Required Competencies
- In-depth knowledge of SLURM workload management with GPU scheduling capabilities.
- Experience with parallel filesystem technologies such as Weka IO and Lustre.
- Proficiency in TCP/IP and high-performance networks including Infiniband.
- Skilled in scripting languages like Bash, Python, and Ruby for automation and configuration.
- Familiar with configuration management tools such as Puppet.
- Strong documentation and communication skills in English, both written and oral.
- Ability to work closely with users and suppliers, demonstrating systematic problem-solving approaches.
- Experience collaborating within multi-disciplinary research environments.
- Ability to manage personal and team workloads effectively under deadline-driven conditions.
- Discretion and versatility in handling complex and non-routine technical assignments.
- Maintains expert-level knowledge in HPC systems administration, storage, or networking.
Qualifications and Experience
- Bachelor's degree or equivalent in relevant field with a minimum of 10 years experience, OR
- Master's degree with at least 7 years relevant experience, OR
- PhD with a minimum of 5 years experience in a related area.
Skills
Work styles they’re looking for
Communication
Leadership
Problem Solving
Team Collaboration
Initiative