This page was automatically translated and may contain errors. View in English.
A

AI/HPC Level 1 Support Engineer

AIHostingHub

Dubai, United Arab Emirates · Jornada completa

Sé el primero en postularte

Experiencia
1–3 años
Salario
Vacantes
1
Al corriente
hace 21 horas
Modo de trabajo
En la oficina
Reanudar
Se requiere solicitud

Dónde trabajarás

Descripción del trabajo

About AIHostingHub

AIHostingHub is a premier provider of advanced AI and High-Performance Computing (HPC) infrastructure in the UAE. We specialize in constructing expansive AI data centers and offering GPU-as-a-Service solutions, ranging from small setups to massive clusters exceeding 16,000 GPUs. Our partnerships with global leaders such as Supermicro and VAST Data enable us to deliver cutting-edge technology, expert knowledge, and comprehensive support to empower ambitious AI projects across the GCC region.

Our service portfolio includes:

  • AI/HPC Data Centers tailored for demanding AI workloads with scalable, customized environments.
  • GPU as a Service delivering on-demand access to extensive GPU clusters.
  • Cybersecurity Managed Security Service Provider (MSSP) powered by Fortinet and AttackIQ, providing round-the-clock protection for critical data and infrastructure.
  • Professional Services offering complete lifecycle support from design through deployment and continuous optimization, all operated by local GCC partners.

Role Overview

We are searching for a full-time, onsite AI/HPC Level 1 Support Engineer based in Dubai. The main objective is to deliver first-level technical and customer support, focusing on AI and HPC infrastructure. The role includes monitoring system health, addressing operational issues, aiding client inquiries, and maintaining accurate documentation. Collaboration with internal teams and escalation of complex problems to higher-tier engineers is also essential.

Primary Duties

  • Observe dashboard metrics (such as Grafana) and ticketing systems to monitor GPU node status, link integrity (InfiniBand), temperatures, and power irregularities.
  • Record, classify, and set priority levels for incidents ranging from P1 to P4, adhering strictly to SLA targets (1 hour for urgent issues, 2 hours for high priority).
  • Perform onsite smart-hands tasks including cable patching, hardware replacements, fiber optic cleaning, and visual system inspections.
  • Run validation scripts post-repair (such as CUDA P2P, NCCL local, DCGMI, and Stream) following RMA activities.
  • Liaise with vendors like Nvidia and Supermicro to facilitate warranty-based replacements.
  • Escalate unresolved problems promptly to Level 2 AI/HPC Engineers.
  • Keep detailed operational logs, asset inventories, and maintenance records up to date.

Required Expertise

  • 1 to 3 years of experience supporting datacenters, network operations centers (NOC), or HPC environments.
  • Working knowledge of GPU server hardware, InfiniBand and Ethernet cabling, and fiber optics.
  • Proficiency with basic Linux command-line tools including dmesg, nvidia-smi, grep, and uptime.
  • Understanding of incident management processes and ability to meet SLA response and resolution times.
  • Willingness and ability to work in 24/7 rotating shifts, including weekends.
  • Strong communication skills and ability to create clear documentation.

Preferred Qualifications

  • Experience using monitoring and management tools such as DCGM, Grafana, Jira, or ServiceNow.
  • Familiarity with liquid cooling distribution units or management platforms like Proxmox and Ceph is advantageous.

Opportunities and Benefits

  • Defined career progression paths toward Level 2 and Level 3 Engineering roles.
  • Hands-on training for HGX platforms and AI validation frameworks.

Déjelo si desea una respuesta; no lo utilizaremos para ningún otro fin.

Haz clic para navegar, arrastrar y soltar, o pasta una captura de pantalla

PNG, JPG, GIF, MP4, WebM, MOV · Máximo 20 MB cada uno · Hasta 5 archivos

🤖
En línea · Ayuda instantánea con IA