Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Engineering

GPU Cluster Infrastructure Engineer

Standout

3+ yrsRemote · WorldwideContractListed 10h ago
Apply now

Backed by

Y Combinator

HQ

🇺🇸 San Francisco

Open roles

47

How to stand out for GPU Cluster Infrastructure Engineer at Standout

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience 3–6 yrs (3+ years)

About the role

structured by ORI

Back Beam ### GPU Cluster Infrastructure Engineer Remote Contract Visa Sponsorship ### About the role Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs.

What you will do

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.

What they are looking for

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.

Nice to have

  • Recent-generation NVIDIA platforms
  • Bare-metal cloud operations
  • Ansible or similar automation
  • Prometheus/Grafana
  • NVIDIA certifications

Before you apply

  • Visa Sponsorship

Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more
GPU cluster operationsInfiniBandGPU node bring-up and burn-inParallel storageData center operationsHardware troubleshootingFabric troubleshootingTechnical communicationNVIDIA HGXNVIDIA DGXUFMBMC/RedfishDCGMNCCLPXE
Full posting text

Back

Beam

GPU Cluster Infrastructure Engineer

Remote

Contract

Visa Sponsorship

About the role

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

About the Role

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

Skills & Experience

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.

Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

About Beam

AI-Native Cloud Platform

Other roles at Beam

•

Chief of Staff

New York, NYFull-time

•

Applied AI Research Engineer

RemoteFull-time

•

Site Reliability Engineer

RemoteFull-time

•

Distributed Systems Engineer

RemoteFull-time

•

Network Engineer

RemoteFull-time

Interested?

Let me introduce you to the founders.

Skip the process

Job details

Salary

$10,500 - $18,000

Location

Remote

Experience

3+ years

Company

NameBeam

IndustryInfrastructure

Team Size5

View profile

Funding

Total raised

$4M

Last stage

Seed

Investors

Alumni Ventures

Founders

Luke Lombardi

CTO & Co-Founder

LinkedIn

Eli Mernit

Co-Founder & CEO

LinkedIn

What happens next.

No applications, no recruiter spam. Just the intro.

01

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

02

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

03

A meeting lands on your calendar

If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Standout

B2B

Agentic hiring marketplace

Backed by Y Combinator

Company pageWebsite

More at Standout

StandoutEngineeringFunded

On-site · Istanbul, İstanbul

3+ yrs · 10h ago

Founding Engineer
GoKubernetes+12 more
View role
StandoutFunded

On-site · Paris, Paris

Senior · 2+ yrs · 10h ago

Operations Manager - Data & AI
CuriosityAttention to detail+8 more
View role
StandoutEngineeringFunded

On-site · Seattle, WA

3–6 yrs · 10h ago

Founding Engineer
Backend engineering+11 more
View role
StandoutEngineeringFunded

On-site · San Francisco, CA

3–6 yrs · 1d ago

Founding Engineer, Robotics and Flight Software
C++RustPython+8 more
View role
Apply