Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Research & Engineering

Member of Technical Staff, Cluster Administration

Inferact

Senior · 5+ yrsOn-site · SingaporeFullTimeSGD 200K – SGD 400K • Offers EquityListed 7d ago
Apply now

Backed by

Lightspeed India

HQ

🇮🇳 India

Open roles

32

How to stand out for Member of Technical Staff, Cluster Administration at Inferact

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Senior · 6+ yrs (5+ years)

About the role

structured by ORI

- - Back to Inferact’s Job Listings # Member of Technical Staff, Cluster Administration Location Singapore Employment Type Full time Location Type On-site Department Research & Engineering Compensation - SGD 200K – SGD 400K • Offers Equity Inferact's mission is to grow vLLM as the world's AI inference engine and…

What you will do

  • Own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive.
  • Take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across systems.
  • Work closely with engineering leadership and infrastructure owners to standardize how to provision, operate, debug, and scale compute across providers.

What they are looking for

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Nice to have

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.
  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.
  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.

Benefits

  • Generous benefits package, including medical, dental, and vision coverage
  • Offers Equity
LinuxSystems AdministrationShell ScriptingIncident ResponseCluster SchedulingResource AllocationHardware DiagnosticsMonitoringSLURMKubernetesBashPythonAnsibleTerraformHelm
Full posting text

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build. About the Role We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock. You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM. Skills and Qualifications Minimum qualifications: - Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar. - Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters. - Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging. - Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics. - Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling. - Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams. - Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling. Preferred qualifications: - Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments. - Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues. - Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems. - Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems. - Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene. - Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams. Bonus points if you have: - Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team. - Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time. - Operated Kubernetes clusters for ML or GPU workloads at meaningful scale. - Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers. - Carried real operational responsibility for infrastructure used by many engineers or researchers. Logistics - Location: This role is based in Singapore. - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity. - Visa sponsorship: We sponsor visas on a case-by-case basis. - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Inferact

India

Backed by Lightspeed India

Company pageWebsite

More at Inferact

InferactResearch & EngineeringFunded

On-site · San Francisco

Senior · 5+ yrs · $200K – $400K • Offers Equity · 8d ago

Member of Technical Staff, Production Site Reliability Engineer
PythonGoBash+9 more
View role
InferactResearch & EngineeringFunded

On-site · San Francisco

Fresher · 16d ago

Inference Engineering, Co-op
PythonC++Rust+13 more
View role
InferactResearch & EngineeringFunded

Remote · Worldwide

Senior · 5+ yrs · 27d ago

Member of Technical Staff, Inference
PythonPyTorchLLM inference+11 more
View role
InferactResearch & EngineeringFunded

Remote · Worldwide

Senior · 5+ yrs · $200K – $400K • Offers Equity · 1mo ago

Member of Technical Staff, Cloud Orchestration (Remote)
KubernetesPythonRust+8 more
View role
Apply