[Job Radar](https://jobradar.live/) / [Jobs](https://jobradar.live/ilanlar) / Senior Distributed Systems Engineer

Active
On-site
Sunnyvale, CA
Posted · 03.03.2026
Lever (US)

# Senior Distributed Systems Engineer

Institute of Foundation Models

About the Institute of Foundation Models
The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.
This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.

The Mission
We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.
This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design.
• Design and optimize expert-parallel and hybrid-parallel communication patterns
• Drive high-performance hierarchical collectives for MoE workloads
• Co-design runtime orchestration with communication topology awareness
• Reduce tail latency and improve determinism across thousands of GPUs
• Architect fault-tolerant distributed execution under real-world cluster failures
Core Technical Scope
• Communication-compute overlap and topology-aware collective optimization
• Deep debugging of NCCL, RDMA, and custom communication layers
• Hybrid expert parallel strategies in modern large-scale MoE systems
• Elastic and resilient distributed job orchestration concepts
• Congestion analysis and routing optimization across InfiniBand/RoCE fabrics
• Microbenchmarking and performance modeling for communication-heavy workloads
Expected Technical Depth
• Hybrid expert parallel communication for Mixture-of-Experts training
• Scaling behavior under network pressure
• Distributed orchestration for elastic, large-scale training
• Fault detection and recovery in distributed GPU workloads
• Cross-layer bottlenecks: GPU ↔ NIC ↔ PCIe ↔ NVSwitch ↔ Fabric ↔ Scheduler
Required Background
• Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth)
• Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA
• Deep familiarity with NCCL and/or UCX internals
• Strong systems programming ability (C/C++, Rust, or Go)
• Strong familiarity with modern model training frameworks such as PyTorch
• Ability to troubleshoot and profile training performance issues related to communication bottlenecks
• Ability to translate research ideas into production-grade optimizations
• Experience debugging distributed hangs, desynchronization, and performance regressions
What We Mean by "Hardcore"
• You can explain why an communication degrades at scale and how to fix it
• You have improved real cluster throughput via communication redesign
• You can trace a distributed hang across ranks and identify the root cause
• You are comfortable working at the boundary between hardware and runtime
Application Requirements
• Include a link to your GitHub (required)
• Provide links to relevant distributed systems, HPC, or large-scale training projects
• Include a list of publications and/or public technical reports (if applicable)
• Describe the hardest distributed debugging problem you solved
• Include measurable performance improvements you have delivered
Academic Qualifications
Master’s, or Bachelor’s + 1 year of relevant experience.

This job was verified from Lever (US). Applications are completed on the original source.

[Apply on the original listing ↗](https://jobradar.live/ilan/83e74030-96c8-4de2-93b6-b22d90f4d30c/git)

## Stop searching one by one for roles like this.

Upload your resume or enter your target roles to see your first 3 matches for free.

[Find jobs for me →](https://jobradar.live/uye/kayit)
Your resume is never shared with employers; it is processed only for matching.

Something wrong with this job?

## Similar jobs

[Institute of Foundation Models Jobs](https://jobradar.live/company/institute-of-foundation-models) · [Jobs in California](https://jobradar.live/jobs/california)
· [Jobs by location](https://jobradar.live/jobs) · [Jobs by company](https://jobradar.live/company)

- [Local Truck Driver](https://jobradar.live/ilan/ce13e327-ea4c-4e74-883a-eae936db4f2b) J.B. Hunt Transport Services, Inc. · Glendale, CA

- [Local Truck Driver](https://jobradar.live/ilan/188cbcd8-3e05-4dae-b70d-6909a55757ab) J.B. Hunt Transport Services, Inc. · Fontana, CA

- [AC Hotel Santa Clara - Front Desk Agent](https://jobradar.live/ilan/99d85205-6f08-4a98-bc4d-b3fc5d5f0dbb) Evolution Hospitality · Santa Clara, CA

- [Supervisory Program Specialist (Human Capital)](https://jobradar.live/ilan/ba050de9-dc69-4830-9522-ee5a16e09677) Federal Emergency Management Agency · Oakland, California

- [Local Truck Driver](https://jobradar.live/ilan/1126f15c-b152-434c-87f7-de15681db887) J.B. Hunt Transport Services, Inc. · Hesperia, CA

- [Regional Truck Driver](https://jobradar.live/ilan/e334b2cc-6c89-4932-b9f8-0aef539ffa89) J.B. Hunt Transport Services, Inc. · Anaheim, CA

- [Registered Nurse Full Time Nights](https://jobradar.live/ilan/2dd8c9d7-aa79-468b-a372-efa372f08682) Kindred · Gardena, CA

- [Java Backend Engineer](https://jobradar.live/ilan/9688c0d3-f269-4cda-8a53-9cf54736fb7c) UST · Sunnyvale, CA
