About the role
structured by ORI← CareersInfrastructure Engineer (Argentina)Remote, anywhere in Argentina · Full time · GPU compute and inference platform (YC S26)Before you apply: a demo video is required, and everything you write and record must be in English. You will need to rate your English 1 to 6, and we are looking for 4 or higher.Enter…
What you will do
- The Terraform that defines both AWS environments, and the discipline of applying it: plan, review, beta, then prod
- The provider fleet: enrollment, node identity (mTLS certificates), health reconciliation, and the lifecycle of VMs scheduled onto bare-metal GPU hosts through Nomad
- The ugly, real parts of GPU hosts: VFIO passthrough, IOMMU groups, kernel flags, why this exact box will not release its GPU
- The network paths: customer traffic through our gateways, the control overlay between nodes and the control plane, and the debugging when a path silently degrades
- Monitoring and the alerts a human actually acts on: dashboards, runbooks, and the discipline of making the next incident boring
What they are looking for
- 2+ years running production infrastructure, professional or a serious homelab/fleet you can show
- Terraform or equivalent infra-as-code on a real cloud account
- Advanced Linux: you have debugged a boot problem, a driver problem, and a network problem, and can tell the stories
- Comfortable reading and writing Go, or strong in one systems language and willing
- Docker and at least one orchestrator (Nomad, Kubernetes, or equivalent)
- Able to own a system end to end: design, rollout, monitoring, incident response
- Working English, written and spoken
Nice to have
- GPU or bare-metal experience
- QEMU/KVM or VFIO
- PKI/mTLS
- Nebula/WireGuard/Tailscale
- DynamoDB
- on-call experience at a small company
Benefits
- High autonomy, high expectations, fair and human culture
- Direct influence on product decisions
- Fully remote, anywhere in Argentina
- Competitive pay, strong mentorship and growth
Full posting text
← CareersInfrastructure Engineer (Argentina)Remote, anywhere in Argentina · Full time · GPU compute and inference platform (YC S26)Before you apply: a demo video is required, and everything you write and record must be in English. You will need to rate your English 1 to 6, and we are looking for 4 or higher.Enter applicationWhy OpenRelayOpenRelay is GPU compute and LLM inference for people who ship. Customers rent GPU machines and call inference APIs; providers plug their hardware into our network and get paid for it. We run a distributed fleet on real hardware, not a reseller skin over someone else's cloud. Small team, high bar, and room for A-players to own things that are live in production.No degree required, show us what you have runWe care about what you can do, not diplomas. The profile we want most is simple: you have built infrastructure and kept it alive. Not a tutorial cluster, real machines with real workloads, where a mistake takes something down and you are the one who brings it back. Show us the systems you have run, what broke, and what you did about it.Core techTerraform, AWS (ECS Fargate, DynamoDB, S3, VPC networking)Nomad on bare-metal GPU hosts (QEMU/KVM, VFIO GPU passthrough)Go (control plane, gateways, node agents)Linux deep enough to argue with: kernel parameters, systemd, networking, driversOverlay networking (Nebula, WireGuard, Tailscale), mTLS PKIVictoriaMetrics + Grafana, GitHub ActionsClaude Code and agentic tooling, dailyWhat you will ownThe Terraform that defines both AWS environments, and the discipline of applying it: plan, review, beta, then prodThe provider fleet: enrollment, node identity (mTLS certificates), health reconciliation, and the lifecycle of VMs scheduled onto bare-metal GPU hosts through NomadThe ugly, real parts of GPU hosts: VFIO passthrough, IOMMU groups, kernel flags, why this exact box will not release its GPUThe network paths: customer traffic through our gateways, the control overlay between nodes and the control plane, and the debugging when a path silently degradesMonitoring and the alerts a human actually acts on: dashboards, runbooks, and the discipline of making the next incident boringOps tooling for a fleet you cannot always SSH into: safe remote execution, self-updating agents, automation that fails loudlySkills and qualifications2+ years running production infrastructure, professional or a serious homelab/fleet you can showTerraform or equivalent infra-as-code on a real cloud accountAdvanced Linux: you have debugged a boot problem, a driver problem, and a network problem, and can tell the storiesComfortable reading and writing Go, or strong in one systems language and willingDocker and at least one orchestrator (Nomad, Kubernetes, or equivalent)Able to own a system end to end: design, rollout, monitoring, incident responseWorking English, written and spokenNice to have: GPU or bare-metal experience, QEMU/KVM or VFIO, PKI/mTLS, Nebula/WireGuard/Tailscale, DynamoDB, on-call experience at a small company.The videoRequired, five minutes or less, in English. Two parts: your face on camera introducing yourself, and a screen recording walking through the best infrastructure you have built or run and the hardest problem it gave you. Casual is fine, content matters more than polish.What to expectFast-moving U.S. startup building GPU compute and inference infrastructureHigh autonomy, high expectations, fair and human cultureDirect influence on product decisionsFully remote, anywhere in ArgentinaCompetitive pay, strong mentorship and growthEnter application
