f

Software Engineer, Distributed Systems

fal

Remote · Full Time

Be the first to apply

Experience
3+ yrs
Salary
USD 180,000 – USD 250,000 / year
Openings
1
Posted
7 hours ago
Work mode
Work from home
Resume
Required to apply

Job description

About fal

fal is at the forefront of the generative media ecosystem, enabling the development of next-generation AI products. We create the infrastructure, tools, and model access needed for teams to transition from concept to production efficiently and at scale, without compromises. By providing a unified platform that combines high-performance inference, orchestration, and monitoring, fal facilitates the creation of innovative AI-powered products for both developers and enterprises. As generative media transforms a multibillion-dollar market over the coming decade, fal stands as the essential ecosystem for ambitious teams.

Role Overview

We are seeking a seasoned software engineer proficient in constructing large-scale computing platforms. This role requires deep expertise in managing complex distributed systems that handle heavy traffic and data volumes. You will focus on achieving system reliability and scalability while minimizing operational complexity.

Key Responsibilities

  • Develop our core platform using Python and Rust, including components such as request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing systems.
  • Design forward-looking platform enhancements to accommodate a 100-fold increase in traffic and ensure low-latency service globally.
  • Utilize AI extensively to automate routine aspects of building complex, dependable systems.
  • Profile and optimize system CPU and memory performance at a low level.

Requirements

  • Minimum of three years building distributed compute and orchestration platforms with Python or Rust.
  • Solid grasp of distributed systems principles, including consensus algorithms, scheduling, fault tolerance, and capacity planning.
  • Deep knowledge of computational complexity and memory management.
  • Proven experience designing scalable systems that perform effectively under production workloads.
  • Expertise in employing observability tools to guide performance tuning and reliability improvements.
  • Strong communication skills with the capability to lead technical decisions across multiple teams.
  • A proactive self-starter who takes ownership, executes promptly, and continuously pursues improvements.

Preferred Qualifications

  • Experience with AI/ML inference or training infrastructure.
  • Background in high-performance systems programming, including asynchronous runtimes, zero-copy techniques, and memory-safe concurrency.
  • Familiarity with multi-tenant compute platform development.
  • Understanding of networking fundamentals and their impact on system performance.
  • Knowledge of GPU workload behavior and scheduling constraints.

Compensation

  • Annual salary range of $180,000 to $250,000, supplemented with equity and benefits. This range applies across Mid, Senior, and Staff levels.

Location & Remote Work

  • Position based in San Francisco, CA, with remote work options for Senior and Staff engineers.

What We Offer

  • Engaging and challenging projects.
  • Ample opportunities for personal and professional growth.
  • Relocation support for candidates moving to San Francisco.
  • Comprehensive health, dental, and vision insurance for US employees.
  • Regular team events and offsites to foster collaboration and culture.

Work styles they’re looking for

Communication Continuous Improvement Proactive Mindset

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help