Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Data Center Business

Staff Software Engineer - Compute

Lambda

Senior · 10+ yrsRemote · USFullTime$314K – $465K • Multiple RangesListed 1mo ago
Apply now

ROUND · 17h ago

$4.0B

HQ

🇺🇸 San Francisco, CA, United States

Open roles

85

How to stand out for Staff Software Engineer - Compute at Lambda

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Senior · 6+ yrs (10+ years)

About the role

structured by ORI

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you will do

  • Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
  • Guide design of compute platform multi-tenant security model
  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.

What they are looking for

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Nice to have

  • Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
  • Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
  • Experience with Cloud Service Provider Kubernetes offerings.
  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Benefits

  • Generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use
compute control plane distributed systemsdurable execution modelssoftware defined networkingsemi-conductor hardware enablementC/C++RustPythonGoDOCADOCA SNAPCUDAKVM
Full posting text

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

About the Role

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale.

You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.

The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.

What You'll Do

We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:

  • Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
  • Guide design of compute platform multi-tenant security model
  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
  • Work with customers on translating vague customer technical requirements into concrete engineering deliverables.
  • Set engineering standards and lead design reviews for mission-critical cloud software at scale.

Who You are

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Nice to Have

  • Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
  • Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
  • Experience with Cloud Service Provider Kubernetes offerings.
  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Our values are publicly available: https://lambda.ai/careers
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Lambda

AI computing infrastructure

Provides AI computing infrastructure and cloud services powered by GPU clusters.

Company pageWebsite

More at Lambda

LambdaData Center BusinessFunded

Remote · US

Senior · 6+ yrs · $230K – $340K • Multiple Ranges · 8d ago

Senior Cloud Infrastructure Engineer, Cloud Foundations
View role
LambdaData Center BusinessFunded

Remote · US

Senior · 10+ yrs · $278K – $325K · 8d ago

Senior Workday Payroll Architect
View role
LambdaData Center BusinessFunded

Remote · US

Senior · 10+ yrs · $314K – $465K • Multiple Ranges · 12d ago

Staff Software Engineer - Managed Kubernetes
View role
LambdaData Center BusinessFunded

Remote · US

Senior · 6+ yrs · $230K – $346K • Multiple Ranges · 12d ago

Senior Software Engineer - Managed Kubernetes
View role
Apply