Job RadarAll jobs
Active On-site San Francisco, California, United States Posted · 30.09.2026 Ashby (US)

Software Engineer, Infrastructure & Performance

Dedalus Labs, Inc.

ABOUT THE ROLE Dedalus Labs builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state, and work over time. We're looking for an infrastructure engineer who's bothered by a slow build, an unexplained latency spike, or a workflow that makes engineers do the same work twice. You want to understand where the time went. Then you want to fix it, measure the improvement, and make sure everyone who uses that system benefits from it. At a small startup, those details affect how quickly the whole company can move. A build that wastes ten minutes wastes them repeatedly. A confusing API slows every product built on it. A deployment that needs someone to supervise it interrupts work elsewhere. You'll seek out these problems across our systems and codebase, including the ones people have learned to tolerate. Your work can bring product releases and customer commitments forward by weeks, and over time, months. You'll work directly with our CTO and engineers, with substantial freedom to investigate, experiment, and choose an approach. We want someone who takes pride in the craft of software engineering and cares about the systems they'll leave for the next person. That means useful abstractions, clear failure behavior, tests that establish the right guarantees, and documentation that explains the decisions. WHAT YOU'LL WORK ON Make the development loop faster. Profile builds, tests, CI queues, artifact transfers, and local workflows. Find repeated work, unnecessary dependencies, poor scheduling, and cache misses. Follow a bottleneck into a compiler, linker, filesystem, network, or upstream dependency when the evidence takes you there. Establish a baseline, make the change, and verify the result on a representative workload. You'll use and improve our open-source tooling: Bessemer (BSMR) https://github.com/dedalus-labs/bsmr, our build system, and Hollywood https://github.com/dedalus-labs/hollywood, which generates GitHub Actions from typed TypeScript definitions and runs actions locally. This includes working on the tools themselves when their implementation limits what we can do. Experience with Buck2, Bazel, Blaze, remote execution, or build cache design is particularly relevant. Build platforms other engineers can build on. Design APIs, controllers, CLIs, and services that support multiple internal products and consumers. Think through resource lifecycles, concurrency, cancellation, permissions, and failure reporting. Work with the engineers using those interfaces, understand their constraints, and make common operations straightforward without hiding important system behavior. Make delivery reliable and understandable. Improve CI/CD, GitOps, environment provisioning, staging, production rollout, and recovery. Connect the source change, build inputs, artifact, and running software so engineers can trace a release and diagnose a failure. Instrument the path with useful metrics, logs, and traces. Use incidents and recurring manual work to identify what needs a software fix. Work with hardware we control. We have an in-office homelab where we test and debug on real machines and networks we control. Depending on your strengths, you can help assemble, provision, and maintain machines, configure networking, and build test infrastructure so engineers can reproduce failures and measure changes under conditions we understand. YOU'D BE A GOOD FIT IF • A problem keeps your attention when the first few explanations turn out to be wrong. You read source, build a smaller reproduction, ask a better question, or collect another trace. You ask for help when it will move the investigation forward and keep responsibility for the outcome. • You notice slowness and repeated effort even outside your immediate assignment. You investigate proactively and can explain why fixing it matters to the people using the system. • You enjoy the last 20% of performance work, including the part that takes 80% of the effort. You can also judge when that effort is worth spending and when a deadline calls for a smaller, complete improvement. • You're detail oriented and a little perfectionist. You care about API names, error messages, correctness, performance, and the next engineer's ability to understand your work. You can make a decision, finish it, and ship. • You respect the craft and stay open to better ways of practicing it. You use modern AI coding tools to accelerate exploration, implementation, testing, and review. You understand the resulting code and take responsibility for its behavior. • You work well with ambiguity. You can turn an incomplete goal into a concrete problem, agree on the constraints, and choose a useful next step. You enjoy the freedom to tinker and can turn an experiment into something the team depends on. EXPERIENCE THAT HELPS You should have strong software engineering skills and experience building or operating infrastructure that other people rely on. We're particularly interested in depth in Go, Rust, C++, or another systems language, practical Linux debugging, and thoughtful API design. You'll also work with TypeScript in our tooling. Build systems, CI/CD, GitOps, Kubernetes, observability, and infrastructure as code are relevant areas of experience. Familiarity with Terraform is useful. Tell us where you've gone deep and what you learned from operating the system after you built it. Hands-on hardware and networking experience is a plus. You may have assembled and maintained machines, installed storage or network cards, provisioned Linux, or configured switches, routing, and VLANs. We're interested in people who understand how the parts fit together and can trace a problem across application code, the operating system, storage, and the network. Deeper experience with Ethernet, SFP+/QSFP optics and direct-attach cables, RDMA, or RoCE is useful additional depth. Show us depth in at least one of these areas and the ability to learn the layers around it. We care about the work you can explain and support with evidence. Degrees, employer names, GitHub activity, and owning a homelab are not prerequisites. THIS ROLE MAY BE A POOR FIT IF • You need a complete specification and a fixed twelve-month roadmap before you can make progress. • You prefer an assignment limited to one tool or layer and find it frustrating to follow a problem across software, infrastructure, and hardware. • You routinely accept slow or manual workflows because they've always worked that way. • You keep polishing after the agreed deadline or declare a performance improvement before measuring it. LOCATION AND COMPENSATION This is a full-time, in-person role in San Francisco. The annual base salary range is $120,000-$250,000 USD, depending on relevant experience and role scope. Relocation assistance and visa sponsorship are available. SHOW US YOUR WORK Use your application to show us one thing you built and one difficult problem you investigated. They can come from the same project. Choose work you already have. We are not asking you to build a new project for this application. • Work: Show the code, demo, design, or result. Explain what you personally owned, what you reused, and one decision that made the result better or simpler. • Investigation: Describe the symptom, your first explanation, the evidence that changed your thinking, and how you established the result. Include the workload or users involved and any important limitations. • Writing: Share one piece of your best writing as a separate, required sample. Simple and succinct is welcome. Choose work that shows your thinking and your voice in a professional or similarly thoughtful context: a strong README, design note, blog post, postmortem, proposal, or essay. A serious, self-contained Reddit post is fine. Tweets, Twitter/X threads, and short social posts do not count. Link it or paste a redacted excerpt. Tell us the intended audience, what you wrote, why you chose it, and any coauthors or AI assistance. Work from employment, open source, research, coursework, or independent projects is welcome. Public source code is optional. A redacted excerpt or a concrete account of private work is fine. Keep confidential information out of your application. Tell us which problem in this role you want to own and why. If your strongest evidence is easy to miss on a resume, point us to it. An existing demo or short screen recording is welcome, and video is optional. For this role, a useful example might be a build you made faster, an unreliable deployment you repaired, or an internal tool other engineers adopted. Explain how you separated the cause from the symptoms, what you chose to leave alone, and what happened after people started using the change. You can explore Bessemer (BSMR) https://github.com/dedalus-labs/bsmr and Hollywood https://github.com/dedalus-labs/hollywood to understand some of our work. Contributions are optional. If you want to contribute, discuss the problem and scope with the maintainers first. Existing work is equally welcome.
This job was verified from Ashby (US). Applications are completed on the original source.
Apply on the original listing ↗
Something wrong with this job?