Back to Job Simulations
Medium → Hard 2–2.5 hoursClient: Kestrel Commerce · OneRoadmap

DevOps & Cloud Engineer Job Simulation

Kestrel's checkout broke ninety seconds after a deploy, and the first thing anyone tried made it worse. Nothing is on fire in a way CPU can see. You get the metrics, the logs, the rollout config and the pipeline - and two hours to work out what is actually wrong, say so in an incident channel, and tell a VP what to change.

Avatar 1Avatar 2Avatar 3Avatar 4Avatar 5Avatar 6Avatar 7Avatar 8Avatar 9Avatar 10

Join 100,000+ One Roadmap-certified candidates

NEW MESSAGE · ONEROADMAP
B

Ben HarlowPlatform Engineering Lead, Kestrel Commerce

Welcome to Kestrel. I'm Ben, I run platform engineering, and you have picked an interesting week to start. We're an online ordering platform - about 40,000 checkouts a day - and right now a good number of those are failing. I'm going to give you the same evidence I have and ask you what you think, because that is genuinely how this works here. Two tickets: the incident happening now, and the plan I have to defend to my VP on Friday. Read the numbers properly. Most of what looks urgent in an outage turns out to be a symptom, and the thing that fixes it is usually boring.

Your manager for this simulation - a fictional role at a fictional company.

Why finish this simulation

A production incident, with the evidence a real on-call engineer would actually have - and none of the hindsight. This is the harder end of our catalogue: several answers look defensible and only one manages the risk properly.

Self-paced startNo cluster to provisionNothing to installWritten by platform engineersMedium → Hard

Read evidence, not vibes. CPU is 22% and the database CPU is 28%, so nothing looks broken - and checkout is 11% broken. The answer is in the numbers you were given, and it is not the number everyone looks at first.

The reflex that makes it worse. Someone scales the deployment to handle the load. It is the most natural thing to do and it deepens the outage. Understanding why is worth more than any command you could memorise.

Rollback as a decision, not a button. You decide when to roll back, what to roll back to, and what else has to move with it. A rollback that restores the image and forgets the replica count looks like a rollback that failed.

A Terraform plan that would end your week. A colleague's branch proposes destroying the production database, and the plan says so in a line most people skim past. Spotting it is a large part of this job.

Alerts nobody was woken by. Two alerts existed. Neither fired. You work out what should have paged instead, which is the difference between monitoring and observability.

Write for the person reading it. One note for an incident channel that Support reads, one plan for a VP who will ask what it costs. Graded on what your writing demonstrates, never on length.

How it works

01Pick up a ticket

Real tasks land on your board like a Jira queue. Download the dataset and work in your tool of choice.

02Submit for review

Your work enters review - results within 24 hours (usually much sooner). The clock pauses while you wait, like a real take-home.

03Pass, then progress

Pass to unlock the next ticket. Clear every ticket to earn a certificate and a spot on the weekly leaderboard.

Skills you will learn and practice

  • Reading metrics for the constraint rather than the loudest number
  • Recognising connection-pool exhaustion from the database's own signals
  • Knowing when adding instances deepens an outage instead of easing it
  • Deciding what to roll back, and what has to move with it
  • Spotting a destroy-and-recreate in a Terraform plan before applying it
  • Designing a rollout that halts itself on a regression
  • Alerting on user-visible symptoms instead of causes
  • Treating a credential in a log as disclosed, and responding accordingly

The tickets

What you'll learn

  • Why a database at 28% CPU can still be the thing that is broken
  • How per-pod connection pools turn adding capacity into removing it
  • What "staging passed" proves when the pipeline builds the image twice

What you'll do

  • Read metrics, an application log, a rollout config and a delivery pipeline
  • Find the real constraint while two plausible explanations sit in front of it
  • Write the incident note Support and two engineers will act on
Prerequisites & Resources →

Weekly leaderboard

Highest ranking points top the board - your score counts 80%, your speed 20%.

Loading…
View full leaderboard →

Your Future Certificate

OneRoadmap job simulation certificate preview

Preview only. Clear both tickets to earn your personalized certificate - free.

Verified Certificate

Unique verification URL & QR code

LinkedIn Integration

Add to your profile with one click

Performance Analysis

Detailed insights into your results

Issued by DPIIT-recognized & MSME-registered company

DPIIT
MSME

Clear it, prove it

Pass both tickets and your verifiable certificate is free - add it to LinkedIn right away. Want to go deeper? Generate a personalized AI review of your performance for a one-time ₹99.