K

Senior AI Compute Infrastructure Engineer

Kraken

Ireland, England, United Kingdom · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
12 മണിക്കൂർ മുൻപ്
Work mode
In office
Resume
Required to apply

Where you'll work

Job description

About the Company

Kraken, operated by Payward—the organization behind multiple financial technology platforms—has developed one of the most advanced and globally reachable financial infrastructure networks over the past 15 years. The company is passionate about advancing an open and global financial system.

Established in 2011, Kraken is among the longest-standing cryptocurrency platforms worldwide, trusted by over 10 million individuals and institutional clients. Its offerings encompass spot trading, margin, futures, staking, and OTC services tailored to diverse investor types.

Role Overview

Kraken is forming a specialized AI Compute and Infrastructure team responsible for enabling next-generation AI model training, inference, evaluation, and experimentation throughout the exchange. This team handles all layers of GPU and accelerator infrastructure, cluster operations, scheduling, model deployment, observability, capacity planning, and cost optimization, thus providing the foundation for in-house AI workloads where privacy, latency, reliability, cost-effectiveness, and product differentiation are vital.

Key Responsibilities

  • Manage and operate GPU and accelerator clusters utilized for training, inference, evaluation, and experimentation, including software components like drivers, runtimes, device plugins, and workload isolation.
  • Design infrastructure that facilitates local GPU model execution by various Kraken teams to minimize external dependency and control compute expenses.
  • Develop and refine scheduling, orchestration, placement, quota management, and utilization systems in diverse accelerator environments.
  • Enhance inference pipelines optimizing latency, throughput, reliability, memory usage, and cost through frameworks such as vLLM, Triton Inference Server, TensorRT, or similar.
  • Collaborate with machine learning engineers and researchers to eliminate constraints in training, evaluation, inference workflows, deployment, and production troubleshooting.
  • Implement observability solutions monitoring GPU utilization, memory, queue depths, performance metrics, failures, capacity pressure, and related costs.
  • Lead initiatives for reliability, including incident response, alerting mechanisms, runbooks, and continuous improvement in maintaining 24/7 AI compute infrastructure.
  • Evaluate and incorporate emerging hardware, cloud instances, accelerators, runtimes, schedulers, and serving frameworks as technology evolves.
  • Create tools that simplify GPU consumption visibility and accountability for internal users without requiring deep infrastructure expertise.
  • Contribute to long-term architectural decisions balancing scalability, performance, cost, operational simplicity, and production safety.

Candidate Profile

  • A minimum of 5 years experience in infrastructure engineering focused on GPU compute, machine learning infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
  • Proven operational management of GPU or accelerator infrastructure in production or similar environments, covering scheduling, orchestration, monitoring, and cost efficiency.
  • Robust systems engineering knowledge in Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production-level debugging.
  • Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalents.
  • Proficient in Python for automating infrastructure tasks, tooling, debugging, and operational integration.
  • Understanding of performance trade-offs including batching, concurrency, memory management, GPU utilization, latency, throughput, availability, and cost considerations.
  • Demonstrated success in optimizing compute costs while ensuring performance, reliability, and high availability.
  • Expertise building observability systems with metrics, logs, traces, dashboards, alerts, and incident management workflows.
  • Ability to perform in high-pressure, always-on environments that demand uptime, throughput, accuracy, and disciplined operations.
  • Strong communication skills to articulate infrastructure trade-offs effectively with researchers, product teams, engineering leadership, security, and platform engineers.

Desirable Qualifications

  • Experience at leading AI research labs, hyperscalers, high-frequency trading firms, or platforms managing large scale ML workloads.
  • Familiarity with custom silicon or specialized accelerators like TPUs, AWS Trainium, Gaudi, or similar technologies.
  • Background in capacity planning, procurement strategies, reserved capacity management, cloud accelerator economics, or GPU fleet cost oversight.
  • Knowledge of distributed training platforms such as DeepSpeed, Megatron-LM, FSDP, Ray, or comparable systems.
  • Skills in debugging low-level issues in CUDA, NCCL, kernel, driver, runtime, or networking performance bottlenecks.
  • Proficiency in systems programming languages including Rust, C++, Go, CUDA used in performance-critical infrastructure development.
  • Experience within crypto, financial services, trading infrastructure, or securely managed production environments.

Additional Information

Applications are reviewed continuously unless a specific deadline is provided. Candidates may redact personal information such as age or graduation dates from their resume for privacy reasons. The company considers applicants with criminal records complying with applicable regulations.

Payward values diversity and inclusion, hires based on merit, and encourages candidates passionate about crypto to apply, even if they do not meet all criteria. Job-related assessments may be administered consistently across applicants and used alongside interviews and experience in hiring decisions. Equal opportunity employment policies prohibit discrimination or harassment of any kind.

Work styles they’re looking for

Communication Collaboration Operational Discipline reliability focus

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help