[Job Radar](https://jobradar.live/) / [Jobs](https://jobradar.live/ilanlar) / Staff AI Infrastructure Engineer

Active
Hybrid
Redwood City, CA, California, United States
Posted · 25.07.2026
Ashby (US)

# Staff AI Infrastructure Engineer

Luma

You'll own the reliability of Luma's 10k+ GPU fleet: the scheduling, efficiency, and resilience that research and products depend on. As a Staff AI Infrastructure Engineer, you'll be a technical authority who turns deep systems knowledge into repeatable, company-wide reliability, and a leader other strong engineers want to work with.
This is close-to-the-metal work — kernels, containers, schedulers, networking, storage, GPU behavior — under demand hard enough that yesterday's solutions break regularly. It's also a technical-leadership role: you'll set the bar and grow the team. If most of your experience has been inside highly abstracted internal platforms where others owned the underlying machinery, this likely isn't a match.
What You'll Own
• Architect and operate large, heterogeneous GPU environments under extreme demand, improving utilization and performance where small gains change company outcomes.
• Resolve failures spanning hardware, OS, runtimes, and orchestration, and eliminate whole classes of instability.
• Define how infrastructure and workloads evolve as cluster size and concurrency grow — scheduling, placement, resource management.
• Work directly with research to build the systems new model capabilities require, and scale inference without sacrificing reliability or latency.
• Hire and develop exceptional systems and reliability engineers, and set the bar for depth, judgment, and production ownership.
• Shape product and research architecture early through strong partnerships.
First 90 Days
One way the first 90 could unfold.
• Days 1–30 — Immerse & Diagnose: Learn the fleet, its failure modes, and the biggest reliability and utilization gaps.
• Days 30–60 — Ship & Validate: Eliminate a recurring class of instability or land a utilization or performance win that moves company outcomes.
• Days 60–90 — Scale & Systemize: Set the reliability direction, redesign ahead of where today's abstractions will fail, and begin building the team.
What You Bring
• Deep expertise in Linux and distributed systems.
• Experience operating GPU or accelerator clusters in real production environments.
• Strong fluency in Kubernetes and modern open-source infrastructure.
• Comfort debugging across hardware, kernel, runtime, and orchestration, and understanding how systems behave under contention and at scale.
• You write code and build automation, and think in bottlenecks, failure modes, and trade-offs.
• Judgment engineers trust, especially when things break.
Nice to Have
• You raise reliability standards company-wide and influence product and research architecture early.
• You build partnerships rather than ticket queues, and attract and level up strong engineers.
• Curiosity for how models use infrastructure, because improving systems expands what becomes possible.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

This job was verified from Ashby (US). Applications are completed on the original source.

[Apply on the original listing ↗](https://jobradar.live/ilan/499bbead-2007-4934-8c83-5832164125cb/git)

## Stop searching one by one for roles like this.

Upload your resume or enter your target roles to see your first 3 matches for free.

[Find jobs for me →](https://jobradar.live/uye/kayit)
Your resume is never shared with employers; it is processed only for matching.

Something wrong with this job?

## Similar jobs

[Luma Jobs](https://jobradar.live/company/luma) · [Jobs in California](https://jobradar.live/jobs/california)
· [DevOps Engineer Jobs in California](https://jobradar.live/jobs/california/devops-engineer) · [Jobs by location](https://jobradar.live/jobs) · [Jobs by company](https://jobradar.live/company)

- [Software Engineer - Azure DevOps](https://jobradar.live/ilan/6f663a79-1033-4e93-b742-73d859423da5) DLR Group · Los Angeles, California, United States

- [Staff Platform Engineer, Security](https://jobradar.live/ilan/9c2b626c-4317-4d46-9e49-ec3a01a0b482) True Anomaly · Denver, CO or Long Beach, CA or SF Bay Area

- [Staff Platform Engineer, Infrastructure](https://jobradar.live/ilan/0957225e-a7dd-43e2-8f54-9c818611f628) True Anomaly · Denver, CO or Long Beach, CA

- [Staff Platform Engineer, AI](https://jobradar.live/ilan/af1611de-7b09-4a49-8046-1f2149ccb132) True Anomaly · Denver, CO or Long Beach, CA

- [Staff Mission Cloud Engineer](https://jobradar.live/ilan/3b7d78ba-4be2-4bc9-878e-5ace6cdcd65e) True Anomaly · Denver, CO or Long Beach, CA

- [Senior Platform Engineer, Infrastructure](https://jobradar.live/ilan/a51bf5b1-9abc-44c4-a73b-895a8e286933) True Anomaly · Denver, CO or Long Beach, CA

- [Senior Platform Engineer, AI](https://jobradar.live/ilan/543feb81-cf12-440a-9149-181e6fcbc05b) True Anomaly · Denver, CO or Long Beach, CA

- [Senior DevOps Engineer](https://jobradar.live/ilan/9e422b87-351b-4dbc-887b-62c0abc163ee) True Anomaly · Denver, CO or Long Beach, CA
