Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

AI / ML · ML Engineering

Staff / Principal Machine Learning Engineer, Serving - UK

Inworld AI

SeniorOn-site · UKFullTimeListed 5mo ago
Apply now

Backed by

Lightspeed India

HQ

🇮🇳 India

Open roles

19

How to stand out for Staff / Principal Machine Learning Engineer, Serving - UK at Inworld AI

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Senior · 6+ yrs

About the role

structured by ORI

About Inworld Inworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. We build first-party speech models, serve LLMs, and run the inference behind modular APIs designed for high-volume, realtime workloads.

What you will do

  • Take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.

What they are looking for

  • Candidates must already have the legal right to work in the United Kingdom, as visa sponsorship is not available for this role.

Benefits

  • Equity
Inference OptimizationModel AccelerationHigh-Performance SystemsDistributed Systems & ScalingQuantizationDistillationCaching StrategiesContinuous BatchingvLLMTRT-LLMC++CUDARustPythonKubernetesRay
Full posting text

About Inworld Inworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. We build first-party speech models, serve LLMs, and run the inference behind modular APIs designed for high-volume, realtime workloads. Hundreds of millions of users interact with Inworld powered apps every day and we serve over 10 trillion LLM tokens per month. Our models and infrastructure support consumer applications across companions, healthcare, fitness, education, media, and more. Our work spans model research, realtime inference, large-scale serving infrastructure, and the APIs developers use to bring these capabilities into production. We’ve raised more than $125M from Lightspeed Venture Partners, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, Stanford, and others. Our technology has powered experiences from companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Inworld has also been recognized by CB Insights as one of the 100 most promising AI companies globally and named one of LinkedIn’s Top 10 Startups in the USA. Who We're Looking For A year ago, reliably working agentic systems and sub-second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template — we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood. Experience We Find Useful You don't need all of this. But you need enough to make a case. - Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM. - Model Acceleration. Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding. - High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs. - Distributed Systems & Scaling. Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections. - Public work. Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups. - Full-cycle ownership. You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production. - Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems. Who Thrives Here - You don’t need a roadmap to start walking; you’re comfortable picking a direction and building the map as you go. - You believe engineering isn't finished until it’s shipped and stable. You have a bias for impact over purely theoretical optimizations. - You don't just ship code; you obsess over the why. You’re the first to question an architecture if you think there’s a better way to solve the core latency or throughput problem. - You aren't satisfied with "the PM said so." You thrive on deep context and want to understand the fundamental logic behind every decision we make. What Working Here Is Like We hand you unclear problems and expect you to make them clear. We value engineers who say "I don't know yet" and then design the benchmark or prototype that finds out. We treat performance, latency, and reliability as first-class product features, not a box to check before launch. Impact comes before everything else, though we support sharing work and open-source contributions that move the field forward. Your work should be visible. Flat structure, fast iterations, minimal process theater. The base salary range for this full-time position is £140,000 – £200,000. In addition to base pay, total compensation includes equity and benefits. Within the range, individual pay is determined by work location, level, and additional factors, including competencies, experience, and business needs. The base pay range is subject to change and may be modified in the future. Candidates must already have the legal right to work in the United Kingdom, as visa sponsorship is not available for this role. For candidates interested in relocating to the San Francisco Bay Area in the future, full U.S. visa and relocation support may be available, subject to business needs and applicable legal and work authorization requirements. Inworld Jobs Privacy https://inworld.ai/jobs-privacy

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Inworld AI

India

Backed by Lightspeed India

Company pageWebsite

More at Inworld AI

Inworld AIGTMFunded

On-site · Mountain View, California, USA

5+ yrs · 4mo ago

Founding AI Solutions Engineer - USA
PythonJavaScriptTypeScript+9 more
View role
Inworld AIML EngineeringFunded

On-site · Switzerland

Senior · 5mo ago

Staff / Principal Machine Learning Engineer, Serving - Switzerland
Model AccelerationQuantization+14 more
View role
Inworld AIML EngineeringFunded

On-site · Serbia

Senior · 5mo ago

Senior / Lead Machine Learning Engineer, Serving - Serbia
Model AccelerationQuantization+14 more
View role
Inworld AIML EngineeringFunded

On-site · Germany

Senior · 5mo ago

Senior / Lead Machine Learning Engineer, Serving - Germany
Model AccelerationQuantization+14 more
View role
Apply