[Job Radar](https://jobradar.live/) / [Jobs](https://jobradar.live/ilanlar) / Senior Infrastructure/Site Reliability Engineer On-Call

Active
On-site
San Francisco, California, United States
Posted · 01.10.2026
Ashby (US)

# Senior Infrastructure/Site Reliability Engineer On-Call

Clera

ABOUT THE ROLE
This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will be the technical expert enterprise clients depend on when critical systems fail. You will own incident response across Kubernetes clusters, Ceph storage, and bare metal servers, keeping high-value AI workloads running at all times. Your calm judgment and deep distributed systems expertise will have a direct and immediate impact on client operations.
WHAT YOU'LL DO
• Respond to and resolve production incidents across client infrastructure spanning Kubernetes, Ceph, and bare metal environments.
• Troubleshoot complex distributed systems problems including pod scheduling failures, CNI networking issues, storage performance degradation, and hardware faults.
• Handle escalations requiring deep expertise in etcd clusters, Ceph RGW authentication, Cilium networking, and bare metal load balancers.
• Communicate directly with enterprise clients during incidents, providing clear status updates and resolution timelines.
• Participate in a follow-the-sun on-call rotation with engineers across multiple time zones.
• Document incidents thoroughly and improve runbooks based on recurring patterns.
• Collaborate with the infrastructure team on long-term reliability improvements and automation to reduce incident frequency.
WHAT WE'RE LOOKING FOR
• 5 or more years of production experience with Kubernetes in enterprise environments, including cluster operations, bare metal troubleshooting, admission controllers, and control plane architecture.
• Deep production experience with distributed storage systems, particularly Ceph, or equivalent platforms such as Weka or VAST.
• Production experience with at least one CNI plugin, preferably Cilium or Calico.
• Strong modern Linux systems administration skills and comfort with bare metal infrastructure, IPMI, hardware troubleshooting, and networking.
• Demonstrated ability to systematically diagnose and resolve complex distributed systems issues under pressure during live outages.
• Production experience with etcd cluster management, including backup and restore procedures.
• Experience with GPU infrastructure for AI/ML workloads, including the NVIDIA Kubernetes operator.
• Familiarity with infrastructure-as-code tools such as Ansible, Kubespray, or similar orchestration frameworks.
• Clear, composed communication with both technical and non-technical audiences during incidents.
• Background in AI inference or training infrastructure is a strong plus.
LOCATION
On-site in San Francisco, CA. This role involves a follow-the-sun on-call rotation and requires timezone flexibility. Visa sponsorship is not available.

This job was verified from Ashby (US). Applications are completed on the original source.

[Apply on the original listing ↗](https://jobradar.live/ilan/4f623075-4038-4fe4-aa80-95255e1c8a50/git)

## Stop searching one by one for roles like this.

Upload your resume or enter your target roles to see your first 3 matches for free.

[Find jobs for me →](https://jobradar.live/uye/kayit)
Your resume is never shared with employers; it is processed only for matching.

Something wrong with this job?

## Similar jobs

[Clera Jobs](https://jobradar.live/company/clera) · [Jobs in California](https://jobradar.live/jobs/california)
· [DevOps Engineer Jobs in California](https://jobradar.live/jobs/california/devops-engineer) · [Jobs by location](https://jobradar.live/jobs) · [Jobs by company](https://jobradar.live/company)

- [Senior Platform Engineer (9976) – Department of Technology](https://jobradar.live/ilan/dfb4ede4-ed45-438c-b3e8-97b1e871d7e9) City and County of San Francisco · San Francisco, CA, United States

- [Senior Platform Engineer, Cloud Infrastructure & Live Services - Mattel Digital Studio](https://jobradar.live/ilan/140a73e5-0365-495d-8f9d-e2722bfd96ee) Mattel · El Segundo, CALIFORNIA, United States

- [Senior Staff Data Platform Engineer - Data Access Team](https://jobradar.live/ilan/fd9dfc6f-d8e2-430b-af43-82c2ccc106ce) ServiceNow · Santa Clara, CALIFORNIA, United States

- [Senior Staff Agentic Search Infrastructure Engineer - Moveworks](https://jobradar.live/ilan/d4944a8a-b834-41eb-80e0-6da9b7b8b472) ServiceNow · Mountain View, CALIFORNIA, United States

- [Senior (DevOps) Machine Learning Engineer](https://jobradar.live/ilan/96447fe1-9876-46c8-8980-b17d1e1c70a2) ServiceNow · Santa Clara, CALIFORNIA, United States

- [Senior Software Platform Engineer](https://jobradar.live/ilan/85ecbc1a-f650-4e3a-9118-999525795572) Gradient Robotics · San Francisco, California, United States

- [Software Platform Engineer](https://jobradar.live/ilan/3f37e3a9-1c5b-466e-87d7-59caecdea853) Gradient Robotics · San Francisco, California, United States

- [Senior Site Reliability Engineer](https://jobradar.live/ilan/d6c043a8-f298-41d1-ada4-b8c63fc5d08f) Airbyte · San Francisco, California, United States
