P

HPC Architect

Paradigm Operations LP-AL

Remote · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
8시간 전
Work mode
Work from home
Resume
Required to apply

Job description

About Andromeda Cluster

Founded by Nat Friedman and Daniel Gross, Andromeda Cluster aims to provide early-stage startups with access to high-scale AI infrastructure, previously exclusive to hyperscalers. Initially launching with a single managed cluster that quickly reached capacity, Andromeda has since developed the underlying systems, networks, and orchestration layers to enhance accessibility to AI infrastructure globally.

The platform collaborates with top AI labs, data centers, and cloud providers to enable seamless routing of training and inference workloads across a global supply chain, promoting flexibility and efficiency in a rapidly expanding market. Andromeda envisions creating a liquidity layer for worldwide AI compute resources, expanding its reach to attract brilliant minds in AI infrastructure, research, and engineering.

Role Overview

The HPC Architect will serve as the technical lead on provider relationships. This role is responsible for evaluating potential providers’ clusters to determine their suitability for inclusion in Andromeda’s network and assisting providers in meeting required standards when necessary. Responsibilities include vetting providers based on quality metrics, collaborating closely with procurement for sourcing, and maintaining ongoing technical collaboration to prevent performance issues.

Key Responsibilities

  • Evaluate potential compute providers by assessing their cluster architectures, GPU hardware, networking fabric, storage solutions, and orchestration against established quality benchmarks.
  • Establish and continuously refine the qualification criteria, acceptance test suites, benchmarking methodologies, and quality thresholds from the ground up, shifting these from informal knowledge to formal standards.
  • Perform hands-on validation activities critical to quality assurance such as burn-in testing, fabric inspections (InfiniBand/RoCE), NCCL and application benchmarking, and storage performance assessments. For repeat validations, develop methodologies and oversee result evaluations.
  • Support providers through the technical onboarding process, working directly with engineers to close gaps, optimize configurations, and bring clusters up to compliance efficiently.
  • Partner with procurement teams to identify and technically vet new providers, accurately gauging the remediation effort needed prior to commitment.
  • Maintain strong technical relationships with existing providers by serving as the primary engineering contact, detecting architectural or quality deviations proactively.
  • Continuously update provider-facing documentation and internal standards to improve the speed and consistency of qualification procedures.

Qualifications and Skills

  • Extensive experience in designing, building, or managing GPU-based HPC clusters at significant scale, with deep understanding of performance, stability, and debugging.
  • Strong expertise in network fabric technologies, particularly InfiniBand and RoCE, capable of analyzing topology and diagnosing performance issues beyond surface symptoms.
  • Proficiency with distributed orchestration systems and HPC software stacks, including Slurm, Kubernetes, OpenMPI, and underlying Linux system engineering.
  • Practical knowledge of data center infrastructure including power, cooling, physical cabling, and hardware environment, enabling effective onsite facility assessments.
  • Astute benchmarking judgment, understanding critical metrics for large-scale AI training, the potential for manipulation, and design of robust, manipulation-resistant tests.
  • Strong ability to formalize and write precise, actionable technical standards usable by external engineering teams without direct supervision.
  • Effective communicator who can maintain positive relationships with external providers even when delivering critical feedback or failing evaluations.
  • Comfortable working within undefined frameworks, capable of creating new qualification systems where none currently exist.

Preferred Additional Experience

  • Background in hyperscalers, neocloud services, colocation, or data center providers.
  • In-depth understanding of NVIDIA data center GPU platforms such as DGX/HGX, NVLink/NVSwitch architectures, and large-scale GPU health and reliability metrics.
  • Experience supporting AI research environments or large training customers, with awareness of typical workload sensitivities to issues like stragglers, network jitter, and storage bottlenecks.
  • Demonstrated success in cluster acceptance testing or site commissioning, including transitioning clusters from delivery to production phases.

Expected Impact Within One Year

  • Create a documented, trusted qualification bar including acceptance criteria, benchmark suites, and quality thresholds, moving away from tribal knowledge to standardized procedures.
  • Enable providers to meet clear technical requirements upfront, reducing time from sourcing to qualification consistently.
  • Develop repeatable and uniform onboarding processes that minimize surprises during customer training workloads by catching issues early in acceptance phases.
  • Establish trusted technical partnerships with provider engineering teams, becoming the primary consultant on infrastructure decisions.
  • Create a procurement pipeline deeply informed by technical due diligence, influencing buying decisions based on realistic remediation estimates.

Why Join Andromeda?

  • Join a fast-growing company at the forefront of the expanding AI infrastructure market.
  • Be the inaugural HPC Architect on the solutions engineering team, shaping the role and function from its inception.
  • Competitive salary coupled with significant equity participation.
  • Comprehensive employee benefits including healthcare, dental, vision, 401(k), and unlimited paid time off for employees and their dependents.

Diversity and Inclusion

Andromeda Cluster is committed to equal opportunity employment, valuing diversity and fostering an inclusive work environment free from discrimination based on race, religion, color, nationality, gender, sexual orientation, age, marital or veteran status, or disability.

Work styles they’re looking for

Adaptability Problem Solving Attention to Detail Technical Communication

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help