[İş Radarı](https://jobradar.live/) / [Jobs](https://jobradar.live/ilanlar) / Inference

Active
On-site
Bay Area, California, United States
Posted · 30.04.2026
Ashby (US)

# Inference

Genesis

WHAT YOU’LL DO
• Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics
• Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization
• Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks
• Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)
• Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks
WHAT YOU’LL BRING
• Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)
• Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)
• Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling
• Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments
• System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

This job was verified from Ashby (US). Applications are completed on the original source.

[Apply on the original listing ↗](https://jobradar.live/ilan/afe62a46-d48e-40e8-9aa1-65c1823c5c1e/git)

## Stop searching one by one for roles like this.

Upload your resume or enter your target roles to see your first 3 matches for free.

[Find jobs for me →](https://jobradar.live/uye/kayit)
Your resume is never shared with employers; it is processed only for matching.

Something wrong with this job?
