Open roles

AI / ML

Founding Engineer — ML Platforms Engineer

Cumulus Labs

On-siteListed 3d ago
Apply now

Backed by

Y Combinator

HQ

🇺🇸 San Francisco

Open roles

2

Experience Not stated in the posting - worth applying if the stack matches

About the role

structured by ORI

Cumulus LabsThe Fastest Multimodal Inference OSFounding Engineer — ML Platforms Engineer$150K - $300K•1.00% - 5.00%•San Francisco, CA, USJob typeFull-timeRoleEngineering, BackendExperience3+ yearsVisaUS citizen/visa onlySkillsDistributed Systems, Software ArchitectureConnect directly with founders of the best…

What you will do

  • Build and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime
  • Design and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant-facing inference gateway
  • Own observability for the fleet: metrics, logs, and traces at scale
  • Debug hard, systems-level problems across the stack, from scheduling logic down to GPU memory and networking
  • Make real architectural decisions, not just implement someone else's design

What they are looking for

  • 3+ years experience
  • US citizen/visa only
  • Excellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem
  • Real production experience, ideally with systems that had to stay up and scale under load
  • Strong design instincts: you can reason about tradeoffs, not just follow a framework's conventions
  • Fast learner who can go deep in unfamiliar territory
  • Comfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor
  • You want to work in person, in a small team, solving problems nobody has solved before

Nice to have

  • Specific experience with Go or Kubernetes is a plus, not a requirement
Distributed SystemsSoftware ArchitectureData StructuresAlgorithmsGoKubernetesTerraformClaude CodeCUDAC++
Full posting text

Cumulus LabsThe Fastest Multimodal Inference OSFounding Engineer — ML Platforms Engineer$150K - $300K•1.00% - 5.00%•San Francisco, CA, USJob typeFull-timeRoleEngineering, BackendExperience3+ yearsVisaUS citizen/visa onlySkillsDistributed Systems, Software ArchitectureConnect directly with founders of the best YC-funded startups.Apply to role ›Veer ShahCEOVeer ShahCEOAbout the roleAbout the role Cumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're looking for an ML Platforms Engineer to help build and run the orchestration layer underneath our inference and agent products, the system that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization. We care more about how you think than which languages are on your resume. Our stack today includes Go, Kubernetes, and Terraform, but we're looking for someone who can walk into any part of a production system, understand it, and make it better, not someone who only knows one toolchain. What you'll do Build and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime Design and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant-facing inference gateway Own observability for the fleet: metrics, logs, and traces at scale Debug hard, systems-level problems across the stack, from scheduling logic down to GPU memory and networking Make real architectural decisions, not just implement someone else's design Ship fast, own your systems end to end, and work directly with the founder What we're looking for Excellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem Real production experience, ideally with systems that had to stay up and scale under load Strong design instincts: you can reason about tradeoffs, not just follow a framework's conventions Fast learner who can go deep in unfamiliar territory; specific experience with Go or Kubernetes is a plus, not a requirement Comfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor You want to work in person, in a small team, solving problems nobody has solved before Why Cumulus We're small, early, and building the systems layer for the next generation of AI infrastructure. You'll own real infrastructure from day one, not tickets in a backlog. About the interview1. Intro call with a founder (30 min) — background, mutual fit 2. Technical deep dive (30-45 min) — a real problem from our orchestrator, discussed live 3. Take-home or paired session on a scoped systems problem 4. References We move fast: most candidates hear back within a few days at each stage, and we aim to get from first contact to offer in under two weeks. About Cumulus LabsCumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're a YC company (W26) taking on some of the hardest systems problems in inference. We ship three products: Ion: a proprietary C++ inference engine with hand-written CUDA kernels, built for unified-memory NVIDIA Grace hardware (GH200, GB200, GB300). Sub-second model hot-swapping, custom attention scheduling, and split inference for physical AI. It outperforms leading inference providers on vision benchmarks, with numbers on our blog. Paladin: a GPU orchestrator that schedules training and inference across heterogeneous, multi-cloud fleets, allocates GPUs fractionally, and live-migrates running jobs between GPUs with no downtime. Talos: an agent platform that builds, runs, monitors, and continuously optimizes production agents on top of the stack. Companies already run inference and agents on Cumulus in production, and we're backed by YC and many other investors. We're small and early, so you own whole systems end to end rather than tickets. If you want to write kernels, build distributed schedulers, and squeeze every token per second out of the newest NVIDIA silicon alongside people who care about doing it right, this is the place. We're based in San Francisco and work in person. Cumulus LabsFounded:2025Batch:W26Team Size:2Status:ActiveLocation:San Francisco FoundersVeer Shah CEOVeer Shah CEO

Apply