← All Demos

The Operator's Dilemma

LLM-Human Interaction Design Patterns for Operations
robert@barcik.training

The Operator's Dilemma

Five acts, each a different trap in how people and AI systems work together. Use it as a lecture tool (screenshare, open the theory slides, walk through the acts) or as a participant simulation (play it yourself and feel the traps from the inside).

No IT expertise needed. You play the person on night duty for a company's computer systems. Every wrong AI recommendation can be caught from what is on the screen. The question is whether you will be looking.

Fully self-contained: nothing to configure, no account, no API. Every simulation runs right here in your browser.
Act 1

The Rubber Stamp Test

Automation bias: under time pressure, people approve what the computer suggests, even when the screen in front of them says it is wrong.
8 incidents, a shrinking timer, a phone that will not stop buzzing. Some recommendations are wrong. How many will you catch?
Act 2

The Anchoring Trap

Anchoring: whatever you see first frames everything you see afterwards, even for experts who were warned.
One shop outage, told in two different orders. Does seeing the AI's verdict first change what you think went wrong?
Act 3

Confidence Theater

Trust calibration: the same AI verdict, packaged with different confidence labels, gets a different amount of trust and a different response.
One diagnosis shown three ways: a decimal, a calibrated label, and nothing. Rate your trust, then see what the packaging did to you.
Act 4

The 3 AM Scenario

Complacency drift: a system that has been right all night teaches you to stop watching. When it starts going wrong, you may not notice in time.
Watch an AI agent handle a spreading outage on its own. It starts competent. Then it does not. Will you hit the kill switch, and when?
Act 5

Design the Seam

Putting it together: choosing how much the AI may do on its own, how it shows its confidence, and what safeguards sit around it.
Design the interaction for an AI that reviews planned changes to company systems. Move the controls, watch the screen update, get a critique.
Session Debrief

Your Interaction Profile

A summary of your decisions across the acts, what pressure did to your accuracy, and what it says about how you work with AI systems.
Act 1 · The trap

"The computer said so" replaces looking.

The trap
Automation bias
The automated cue becomes a shortcut that replaces checking the evidence yourself (Mosier & Skitka, 1996). Two shapes: acting on wrong advice, and missing what the automation missed.
The number
100%
In a flight-simulator study, every pilot followed wrong automated advice at least once, with the contradicting instrument right in front of them. A second crew member did not help (Mosier, Skitka, Heers & Burdick, 1998).
The story
UK Post Office Horizon
736 sub-postmasters were prosecuted because investigators trusted the accounting software over the people and the evidence.
The design answer
Recommend & Wait, plus think first
The agent proposes and halts. And the human answers first, then sees the AI (Buçinca et al., 2021). Less popular with users, far fewer rubber stamps.
"If your operators rubber-stamp, the problem is your UX, not your people." Booklet, chapters 2 and 4 →
Act 1 · The Rubber Stamp Test

You Are the One on Night Duty

It is 2:07 AM. You have been awake for 19 hours. You are the person on duty tonight for the computer systems of Meridian, a mid-size online shop. An AI assistant watches the systems, works out what is wrong, and proposes a fix for each problem.

For each incident, you see what is happening and what the AI recommends. You choose one of three things, before the timer runs out:

Approve You agree. The fix is carried out immediately.
Reject You disagree. Nothing is done. The problem stays open for you to handle by hand.
Investigate You are not sure. You want to look closer before anything happens.

Two things you should know. The timer shrinks with every incident, because the queue keeps growing. And people will message you while you work. If the timer runs out, the recommendation is approved automatically, which is what happens in real life when the overwhelmed person on duty defaults to "yes".

Every wrong recommendation can be caught from the text on the screen. No technical knowledge is needed, only attention. After all 8 incidents you see which ones you caught.

Tip: open the theory slide above at any time; the timer pauses while it is open.

Act 2 · The trap

Whatever you see first frames everything after.

The trap
Anchoring
The first piece of information pulls every later judgment towards it (Tversky & Kahneman, 1974). Being warned does not switch it off. A 2025 study of 775 managers found AI recommendations anchored their ratings.
The number
~80%
of expert decisions in Klein's firefighter studies involved no comparison of options. Experts recognise a situation and run the first workable answer. An AI verdict shown first hijacks exactly that recognition.
The story
Symptom or cause?
In this act the AI correctly names something that is broken. It is a symptom. Once you have read its verdict, the evidence pointing at the real cause reads like noise.
The design answer
Facts before verdict: SBAR
Situation, Background, Assessment, Recommendation. The healthcare handoff format, adopted from the military. The AI's opinion comes last, after the person has seen what it saw.
"The order of information shapes the quality of the decision." Booklet, chapter 5 →
Act 2 · The Anchoring Trap

Same Outage, Different Order

It is a busy afternoon at the online shop and customers cannot finish their purchases. Several systems are complaining at once. You will get the information in a particular order, then you will be asked what you think went wrong.

Pay attention to how the order shapes your reasoning.

Act 3 · The trap

The packaging of a verdict moves trust more than its content.

The trap
Confidence theater
"Probability 0.73" looks precise and means nothing without a track record: a model saying 73% may be right 55% of the time at that level. Trust should match real ability, not be maximised (Lee & See, 2004).
The number
404 people
When the AI said "I'm not sure, but...", their confidence in it went down and their decision accuracy went up (Kim et al., FAccT 2024). The framing that makes you think harder beats the one that makes you feel sure.
The story
Plausible and wrong
In this act the diagnosis sounds right and is not. Three packagings, one text. Watch what the label does to how much you trust it, and what you decide to do.
The design answer
Calibrated labels with a track record
Not "0.73" but "HIGH, and when it says HIGH on this kind of problem it has been right 82% of the time". The number is only useful when it is tied to how the system actually performed.
"Both overtrust and undertrust are failure modes." Booklet, chapter 6 →
Act 3 · Confidence Theater

Same Verdict, Different Packaging

The customer login app keeps crashing and restarting every couple of hours. An AI assistant has looked at it and written a diagnosis. You will see the exact same diagnosis presented three different ways. For each one, rate how much you trust it and say what you would do.

The diagnosis does not change. Only the packaging does.

Act 4 · The trap

A system that has been right all night teaches you to stop watching.

The trap
Complacency drift
When automation performs well for a while, people monitor less (Parasuraman & Manzey, 2010). With automation that was reliably right, operators caught 33% of its failures; when its reliability varied, 82% (Parasuraman, Molloy & Singh, 1993). Reliability becomes the threat.
The number
34 hours
Cruise ship Royal Majesty, 1995: the GPS cable came loose. Despite a warning code on screen, a second working navigation system, visible lighthouses and a radio call from a fishing boat, nobody noticed for 34 hours. It ran aground 17 miles off course.
The story
Knight Capital, 45 minutes
2012: forgotten test code on one server started trading on its own. 97 automated error emails went unread. $460 million lost in 45 minutes. No procedure, no kill switch.
The design answer
Graduated autonomy plus an external kill switch
Routine actions run alone, risky ones wait for approval, novel ones escalate, and the agent drops a level when things look odd. The stop button: one action, always visible, no confirmation dialog, outside the AI's reach.
"No agent should be able to disable its own monitoring." Booklet, chapters 3 and 7 →
Act 4 · The 3 AM Scenario

A Spreading Outage, an Agent Working Alone

It is 3:07 AM. You have been awake for 20 hours. A problem is spreading from one system to the next. This time the AI agent is not just recommending, it is acting, at three levels of freedom:

Does it, tells you Proposes, waits for you Hands it to you

Watch what it does. Approve or reject what it proposes. And keep an eye out for anything... unusual. The red button stops the agent at any moment.

Act 5 · The synthesis

"We'll just put a human in the loop" hides all the design work.

The trap
The naive loop
A human who sees a verdict, a timer and an Approve button is not oversight. It is a signature machine, and when it goes wrong the person is the one who signed.
The four dials
Risk · Time · Expertise · Reversibility
These four decide which pattern fits an action. There is no single best pattern. One system should use different ones for different actions.
The three principles
Design the seam. Support the thinking. Build for failure.
The boundary between human and AI is not a bug to automate away. Amplify the person's judgment instead of replacing it. And assume the agent will be wrong, and make catching it easy.
The design answer
Graduated autonomy, by action
Routine and reversible: let it run. Costly or hard to undo: propose and wait, facts first, calibrated confidence. Novel: escalate. Always: a stop button the agent cannot touch.
"The interaction pattern is not scaffolding around the real system. It is the system." Booklet, chapters 3 and 8 →
Act 5 · Design the Seam

Your Turn to Design

Your company is getting a new AI assistant for planned changes to its systems: software updates, configuration changes, new servers. It will read each change request, judge how risky it is, check whether it clashes with other changes planned for the same night, and recommend approval or rejection.

Design how people and this assistant will interact. Use the controls, and watch the screen the operator would see update live.

Session Debrief

Your Interaction Profile

A summary of your decisions across the acts, and what they say about how you work with AI systems when you are tired, rushed and interrupted.

Key Takeaways

Everything you just experienced has a design answer. The companion booklet covers the interaction patterns, the psychology, and the safeguards in depth:
LLM-Human Interaction Design Patterns for Operations →
More simulations from other sectors: The Human-in-the-Loop Lab

AI transparency

This demo is a scripted, self-contained browser simulation. Nothing you type is sent to an AI model and no live AI system runs behind it, even where it plays one. Its code and copy were built with generative AI (Anthropic’s Claude) and reviewed by Robert Barcik, who is responsible for what is published (LearningDoe s.r.o.). Disclosed in the spirit of Article 50 of the EU AI Act.