Skip to content

Apply

Situational judgement test questions: how the marking actually thinks

Updated 12 August 2026

A situational judgement test is the only stage of a hiring process where you are marked against a framework you can read in advance and most candidates never do. The scenarios feel like common sense — a slipping deadline, an angry caller, a colleague cutting a corner — and so people answer from instinct. But the test is not asking what you would do. It is asking whether you can recognise what the organisation's published framework says an effective person does, and the two are not always the same thing.

That is not an invitation to fake anything. It is an invitation to prepare properly: the civil service publishes its behaviours, the NHS publishes its values, and most graduate employers describe the strengths their test is built on. Reading those documents before the test is not gaming it — it is doing the reading the test assumes you have done. The scenarios below are worked the way the scoring thinks: not 'what feels right', but 'what does each option cost, and what does the framework weigh'.

One structural point before the questions. Most SJTs ask you either to rank every option from best to worst, or to pick the most and least effective. In both formats the marks live at the extremes: identifying the genuinely strong option and the genuinely damaging one matters more than agonising over the middle. The worked reasoning below spends its effort accordingly.

Your career — see where the written stage fits, and what comes next.

What the screening reviewer needs to see

  • Whether your judgement matches the organisation's published framework — the civil service behaviours, the NHS values, a graduate employer's strengths model. The test is calibrated against that document, not against general good sense, which is why reading it beforehand moves your score.

  • Consistency across scenarios. One collaborative answer proves nothing; the same instinct showing up across fifteen scenarios reads as a disposition. Wild swings between hero-mode and hand-it-upwards read as guessing.

  • Whether you can spot the real weight in a scenario — a safety issue outranks a deadline, a service user outranks internal convenience, an error already in the world outranks the embarrassment of admitting it. Most scored options separate on exactly this.

  • Test discipline itself: reading the instructions, answering the self-report sections honestly enough that your interview does not contradict them, and finishing inside the time. An unfinished test scores the unanswered scenarios at zero.

A colleague on your team is visibly behind, and a piece of joint work is due with your manager tomorrow. You have your own deadline today. Rank the responses.

Why it's asked: The classic collaboration-versus-own-work scenario, and the most common shape in civil service and graduate tests. It is testing whether you can support a colleague without either abandoning your own commitments or quietly absorbing theirs — and whether escalation reads to you as good communication or as telling tales.

Model answerRank-order format, civil service style

The strongest option is almost always the one that talks to the colleague first and looks for a small, concrete assist — not the one that takes the whole task over, and not the one that goes straight to the manager.

Here is the reasoning the scoring rewards. Taking over their work entirely rescues tomorrow at the cost of today: your own deadline slips, and the framework reads that as poor judgement, not generosity. Going to the manager first, before speaking to the colleague, solves the problem while damaging the relationship the test cares about — most frameworks phrase this as working together or treating people with respect. Doing nothing protects your deadline and fails the team outright; it usually belongs at the bottom, just above any option that hides the problem.

So the ranking logic runs: speak to the colleague and agree what small piece you can realistically lift; then, jointly, flag the risk to the manager today rather than tomorrow — early warning is a virtue in every public-sector framework. The trap option is the heroic one. It feels kind and marks badly, because it converts one missed deadline into two.

You notice a factual error in a report your team sent to a senior stakeholder last week. Nobody else has spotted it. What is the most effective and least effective response?

Why it's asked: The error-already-in-the-world scenario. It separates candidates who treat mistakes as reputational events from candidates who treat them as service events. Every serious framework — civil service honesty and integrity clauses, NHS candour language — wants the error surfaced fast, through the right person, with a fix attached.

Model answerMost/least format, EO–HEO level

Most effective: tell the report's owner or your manager today, with the correction drafted. The three components matter together — speed, the right channel, and arriving with the fix rather than just the alarm. An option that says 'flag it to your manager and suggest the corrected figure is reissued' is doing all three and will sit at the top.

Least effective is not, usually, the option that quietly corrects the master copy for next time — poor, but at least motion in the right direction. The floor is the option that decides it probably doesn't matter because nobody has noticed. That option fails on honesty grounds and on service grounds at once, and in values-based tests it is placed there deliberately as the integrity tripwire.

The tempting middle mistake is emailing the senior stakeholder directly yourself. It looks brave and marks modestly: it is fast, but it routes around the report's owner, who both needs to know and may know context you do not — perhaps the figure was corrected in a later version. Surface errors through the person who owns the work. That instinct, applied consistently, is worth several marks across a full test.

A member of the public is angry about a decision you have no power to change, and the queue behind them is growing. Rank the responses.

Why it's asked: The frontline service scenario — standard in civil service, local government, and NHS tests. It is testing whether you can hold a rule and a person at the same time: neither bending the decision to end the discomfort, nor hiding behind the rule and letting the interaction curdle.

Model answerRank-order format, frontline service role

Top of the ranking is the option that acknowledges the frustration, explains the decision's basis in plain terms once, and offers the legitimate next step — a review process, a complaints route, a written explanation. It scores because it treats the person as owed an explanation and the decision as not yours to trade away. Both halves matter; options that do only one half sit in the middle.

The two floor candidates are usually promising to see what you can do — which manufactures false hope about a decision you cannot move, and fails on honesty — and summoning a manager at the first sign of anger, which abandons the interaction while the queue grows. A manager option ranks higher only when the scenario adds threat or abuse; the moment safety enters, escalation stops being avoidance and becomes the correct weight, and well-built tests include exactly that variant to see whether you can tell the two apart.

Note what the strong option does not do: it does not apologise for the decision itself. It apologises for the situation's difficulty. Frameworks with service language notice the difference, and so do the real conversations this test is rehearsing.

Your manager asks you to set aside a data-quality check to hit this week's numbers. You think the check matters. Most and least effective?

Why it's asked: The conflicting-instruction scenario, and the one candidates most often get backwards. It is testing whether you can disagree upwards professionally — the frameworks call it courage, integrity, or making effective decisions — without either silent compliance or self-righteous refusal.

Model answerMost/least format, graduate scheme

Most effective: say what you think, once, with the reason attached — 'my concern is that skipping the check has caused rework before; can we cut scope somewhere else?' — and then, if the manager holds the call and nothing unsafe or improper is involved, do as asked. That compound option scores at the top because it demonstrates both halves of the behaviour: the courage to raise it and the professionalism to accept a legitimate decision that goes against you.

Least effective splits by flavour. Silent compliance — saying nothing and dropping the check — marks badly but not at the floor; the floor is usually deceptive middle ground, such as pretending to run the check while skipping it, or complaining to colleagues instead of the manager. Tests place those options to find candidates whose response to disagreement is indirect.

The calibration point: notice the scenario said data quality, not safety or legality. If the scenario had made the check a safeguarding step or a legal requirement, the weighting inverts and the persistent option — escalating beyond your manager if necessary — becomes the strong answer. Read what kind of check it is before you rank anything; that single habit resolves most of the hard middle orderings.

Two managers each hand you urgent work for the same afternoon. Rank the responses.

Why it's asked: The competing-priorities scenario. The scoring wants the conflict made visible to the people who own it — the strong option asks the two managers to agree the order, or proposes an order and confirms it, rather than silently choosing a favourite or working late to hide the collision. Absorbing the overload unseen reads as poor organisation, not heroism.

You are asked to cover a task you have never been trained on, with nobody senior around until tomorrow.

Why it's asked: The competence-boundary scenario, sharpest in NHS and regulated-sector tests. The framework logic: attempting work beyond your competence is a risk decision, not initiative. Strong options name the limit, do the safe subset, and park the rest with a clear note for tomorrow; weak options either refuse everything or attempt everything.

A teammate presents your analysis in a meeting as their own. What do the options separate on?

Why it's asked: The credit scenario — testing proportion as much as principle. The scored gap is between the private, direct conversation with the teammate (strong: addresses it, preserves the relationship, keeps the issue its actual size) and the public correction mid-meeting or the immediate escalation (weak: right on the facts, wrong on the weight). Doing nothing at all usually sits just above the public options, not above the private one.

Rank-every-option versus most-and-least: how should your strategy differ between formats?

Why it's asked: Because marks are format-shaped. In most-and-least, the middle options are unscored noise — spend your seconds nailing the top and the floor. In full ranking, adjacent middle pairs typically carry partial credit, so a defensible middle order matters and dithering over it still costs time. Knowing this before the clock starts is worth more than one extra practice paper.

The self-report section asks how you typically behave at work. How honest is honest?

Why it's asked: The civil service judgement test opens with exactly this, and the guidance says plainly that your answers may be raised at interview. Answer as the person your referees would describe, not the person the framework sketches: a self-report at odds with your interview reads worse than a modest one. This section is calibration, not a place to spend your ambition.

What is the practice test actually for, and what does a practice score tell you?

Why it's asked: Less than candidates hope, more than cynics say. Official practice tests exist to teach you the format — the timing, the response mechanics, the scenario register — so the real attempt spends nothing on orientation. Most are deliberately unscored or easier than the live test, so treat a practice result as familiarity gained, not a forecast, and treat the format lesson as the point.

FAQ

Can you fail a situational judgement test?
Yes — most are sift instruments with a cut-off, and in high-volume campaigns the cut-off does real work. The civil service judgement test compares your score to a reference group of applicants and presents it as a banding, so 'failing' means falling below the mark the vacancy set rather than below some universal line.
How do I prepare for a situational judgement test in a week?
Three things, in order: read the framework the test is built on — the civil service behaviours, the NHS values, or the employer's strengths pages; take the official practice test to learn the format and timing; then rehearse the reasoning pattern above on a handful of scenarios, always asking what each option costs and what the framework weighs. Bought question banks add little once those three are done.
Are situational judgement tests timed?
Most are, though generously compared with numerical tests — the pressure is decision fatigue across many scenarios rather than seconds per question. Check the invitation email for the format, complete it somewhere quiet, and do not leave scenarios unanswered: an unfinished test scores its blank scenarios at zero.
Should I answer as myself or as the ideal employee?
In the scenario sections, answer as a competent professional applying the organisation's published framework — that is what the scoring is calibrated against, and preparing that way is legitimate. In self-report sections, answer as yourself on a good day: those answers can be revisited at interview, and a gap between your test self and your interview self costs more than a modest self-report ever would.
Do employers see my individual answers?
Usually they see scores and bandings rather than your option-by-option choices, but the civil service explicitly warns that self-report answers may be discussed at interview, and some employers review responses when candidates are close to a borderline. Assume anything you submitted can come up, and answer accordingly.

Read next

Sources

Take the next step with your own application.

Use the worked examples to test the evidence you have, then see how Apply connects to the conversation that follows.