How to Set the Right Difficulty for Video Questions
Calibrate video questions for the learner, outcome, and evidence you need without relying on trivia, time pressure, or confusing wording.
A question is not well calibrated because it feels difficult to the person who wrote it. It is well calibrated when a named group of learners must use the intended knowledge, at the intended stage of learning, to produce evidence you can interpret.
This guide gives you a practical way to control that challenge. Instead of adding obscure details or a short timer, you will adjust four parts of the task: evidence distance, reasoning, response production, and available support.

Difficulty is a relationship, not a label
Difficulty exists between a task and a learner. A question about selecting a microscope objective may be demanding for a beginner, routine for a laboratory technician, and irrelevant for someone whose job never involves microscopy. The words “easy,” “medium,” and “hard” are incomplete until you name the audience, prerequisite knowledge, and purpose.
Begin with three statements: who is answering, what they have already learned, and what decision the response will support. A low-stakes practice check can stretch memory and invite another attempt. A certification decision needs stronger evidence, documented scoring, and tighter control over what the question actually measures.
The Standards for Educational and Psychological Testing make the same starting point explicit: score interpretations must be tied to a proposed use and the construct being measured. In plain language, decide what the answer is supposed to tell you before you decide how hard the prompt should feel.
Write the question from observable evidence
Define the outcome, answer, timestamp, distractors, and feedback before fine-tuning the level of challenge.
Separate meaningful challenge from avoidable difficulty
Meaningful challenge comes from the knowledge and reasoning you intend to develop. Avoidable difficulty comes from something else: ambiguous wording, irrelevant detail, hidden evidence, unfamiliar jargon, poor contrast, or a response control that is hard to operate.
The distinction matters because both can lower the correct rate. If a viewer misses a diagnosis because they overlooked a relevant symptom, the question may have revealed a learning gap. If they miss it because the prompt used an undefined abbreviation, the result says less about the diagnosis skill.
The answer depends on the intended knowledge
A stronger mental model produces a better response for a reason you can name.
The required evidence is available
The learner has seen the necessary information, even when they must connect or transfer it.
Difficulty comes from ambiguity or access
Dense prose, hidden evidence, tiny targets, or unclear conditions raise errors for the wrong reason.
Only speed, stakes, or obscurity changed
A shorter timer, more points, or an incidental fact can increase failure without deepening the task.
The NBME Item-Writing Guide describes flaws that add difficulty unrelated to the testing objective as construct-irrelevant. It also warns that overly long options can shift a task from content knowledge toward reading speed. The lesson applies well beyond medical education: difficulty should come from the intended performance, not from preventable friction.
Avoidable difficulty
Considering all aforementioned procedural contingencies, which option is not inconsistent with safe practice?
Meaningful challenge
The motor has stopped, but the pressure gauge remains above zero. What must happen before restart?
Use four levers to shape the task
Difficulty is easier to control when you stop treating it as one slider. The four levers below describe what the learner must reach for, do, produce, and complete without support. Each can be adjusted without making the language less clear.
The difficulty profile
Change the thinking, not the clarity
Move one lever at a time so you can explain what became more demanding and why.
LEVER 01
Evidence distance
How far must the learner reach for the evidence?
Nearby
Use one visible cue from the current moment.
Integrated
Connect evidence from more than one moment.
Transferred
Apply the principle to a new but comparable case.
LEVER 02
Reasoning move
What must the learner do with the evidence?
Identify
Recognize a fact, feature, step, or relationship.
Explain
Compare, connect, classify, or state why.
Decide
Diagnose, predict, prioritize, or choose an action.
LEVER 03
Response production
How much of the response must the learner construct?
Select
Choose among well-designed alternatives.
Produce
Enter a word, number, sequence, or short response.
Justify
Explain reasoning or identify supporting evidence.
LEVER 04
Available support
How much guidance remains while the learner responds?
Guided
Keep a cue, example, or relevant frame available.
Reduced
Remove one scaffold while preserving fair evidence.
Independent
Ask for retrieval or application without a prompt cue.
The levels are not a universal staircase. A short response may be easier than multiple choice when the alternatives are extremely close. A transfer case may be routine for an expert. Use the map to describe the task, then judge that profile against the audience.
Change one lever at a time during revision. If you remove a cue, add a new case, and switch to an open response simultaneously, you will not know which change caused the new response pattern.
Start from the decision the learner must make
Difficulty calibration begins with performance, not a question format. Ask what the person will need to notice, remember, explain, calculate, or decide outside the lesson. Then write the simplest prompt that produces evidence of that performance.
Suppose a video teaches food storage safety. If the outcome is recalling the approved temperature, a numeric or selected response may be enough. If the outcome is deciding whether a delivery is safe to accept, give the temperature, elapsed time, and packaging condition, then ask for the decision. Adding a rare regulation number would not deepen the same skill.
Topic-shaped question
Which statement about food storage is correct?
Performance-shaped question
This delivery has been at 9°C for three hours. Which action follows the procedure shown in the video?
Adjust evidence distance
Evidence distance is the gap between the prompt and the information needed to answer. A nearby question may ask about the frame still visible behind the interaction. An integrated question may require connecting a diagram from one section with a demonstration from another. A transfer question presents a new case governed by the same principle.
Increase distance only after the learner has encoded the source material. If the required evidence never appeared, or appeared after the question, the prompt is not demanding retrieval. It is demanding a guess. If the evidence is visual, check that the interaction does not cover the exact feature the learner needs.
Nearby evidence
Which valve is highlighted in the current frame?
Integrated evidence
How does the valve position shown now change the pressure pattern demonstrated two minutes earlier?
Research on retrieval offers a useful caution. In two experiments, Pyc and Rawson found that more effortful successful retrieval during practice was associated with better later memory. Successful matters here. Making retrieval impossible, withholding the necessary instruction, or hiding the evidence does not create the same learning condition.
Adjust the reasoning move
The reasoning move describes what happens between evidence and response. Identifying asks the learner to locate or recognize something. Explaining asks them to connect, distinguish, or state why. Deciding asks them to diagnose, predict, prioritize, or choose an action under stated conditions.
Do not discard recall. Some performances depend on prompt access to a fact, term, or step. The better question is whether recall is the intended performance. When the real job requires deciding, a sequence of recall-only prompts cannot establish that the learner can make the decision.
Identify
Which number is the target pressure?
Decide
The pressure drops below target, while flow remains stable. Which cause should be checked first?
Retrieval itself can support learning. In a study using educational prose, Roediger and Karpicke reported better delayed retention after repeated testing than after repeated study. That supports making viewers retrieve important ideas, but it does not mean every question should use the same reasoning move. Match the retrieval to what people will need later.
Adjust response production
Selecting an answer and producing an answer create different evidence. Multiple choice can reveal a misconception when its alternatives represent distinct reasoning paths. Numeric input can test a result without offering candidate values. A short or open response can show whether the learner can formulate the explanation, but it also introduces typing, language, and review demands.
Choose the lightest response burden that still captures the outcome. Do not require a paragraph when one selected action is sufficient. Do not use multiple choice when recognition lets viewers succeed without producing the calculation or phrase they must recall independently.
Use distractors to expose distinct reasoning paths
Make selected responses more diagnostic by building wrong options from real misconceptions, incomplete rules, and procedural errors.
In Interakly, automatically scored types can return an immediate result, while free text and other constructed media may require educator review depending on their configuration. That changes the learner experience and the evidence available at completion. If a response awaits review, describe the criteria in advance and avoid presenting a provisional result as a final judgment.
Adjust the support that remains available
Support includes visible frames, labels, examples, answer choices, formulas, and hints. Removing one support can show whether the learner can perform more independently. Removing every support at once can turn a focused check into a memory contest.
Plan a deliberate fade. First ask with the relevant diagram visible. Later provide the same kind of case without the highlighted cue. Finally ask for application in a new case. Keep the response conditions otherwise stable so the change in independence is interpretable.
Keep realistic context without adding noise
A realistic case can make a question more authentic because the learner must identify which details matter. That does not justify adding random names, decorative numbers, or a long backstory. Every detail should affect the decision, represent a cue that must be ruled out, or make the case resemble the real environment.
Write the minimum case first. Add one contextual detail at a time and ask what job it performs. If the answer remains identical and no trained person would need to evaluate the detail, remove it. This keeps the challenge in filtering relevant evidence rather than enduring prose.
Decorative context
At 10:43 a.m. on a rainy Tuesday, a technician named Morgan approaches a blue machine installed in 2018...
Decision context
The machine has stopped, the isolation indicator is off, and the pressure gauge reads 18 psi. What is the next safe action?
Do not confuse pressure or points with rigor
Interakly supports point values and optional time limits in relevant question editors. Video Player settings can pause on a question, allow retakes, prevent skipping, and control whether viewers can change individual answers. These controls shape the conditions and consequences of a response. They do not rewrite the intellectual task.
Use a time limit only when speed belongs to the outcome or when a carefully tested practice constraint serves a clear purpose. A rapid visual hazard check may need seconds. A calculation, translation, or evidence-based explanation may need enough time for accurate work. Tightening the timer can introduce reading speed, motor speed, and anxiety into the score.
Points communicate weight in a score. They should reflect importance and the evidence produced, not act as a difficulty knob. Retakes alter the consequence of an error. They can make formative practice safer without lowering the reasoning demand of the question itself.
Build a progression across the video
One question cannot carry an entire learning progression. Use several moments to move from supported recognition toward independent transfer. Early prompts can confirm prerequisites. Mid-video prompts can connect ideas. Later prompts can present a new case or ask for a decision.
Progression does not mean every question must be harder than the last. After a demanding transfer task, a short consolidation check can help the viewer name the principle. The sequence should follow the learning logic of the video, not a mechanically rising difficulty line.
Establish the prerequisite
Use a clear, low-burden check to confirm the fact, feature, or rule required by the next section.
Connect evidence
Ask the learner to relate two cues, compare examples, or explain the relationship demonstrated in the video.
Apply the principle
Present a changed case that preserves the governing idea while removing a familiar surface cue.
Consolidate the rule
Use feedback or a short follow-up to help the learner state what should transfer to the next situation.
Make feedback part of the progression
Connect the response to evidence, explain the reasoning, and give the learner a clear next move before the next challenge.
Calibrate with response evidence
Before publishing, record the result you expect and why. “Most trained staff should answer correctly, while new starters may choose the option that skips the pressure check” is more actionable than “medium difficulty.” After a pilot, compare the observed pattern with that prediction.
The classical item difficulty statistic is simply the proportion or percentage of people who answer correctly. The NBME guide calls this the p-value. It is descriptive, not explanatory. A low percentage may reflect a demanding task, a missing lesson, an ambiguous key, an attractive misconception, or an inaccessible presentation.
Interakly Analytics can show correct rates and answer distributions for supported graded questions, while review surfaces preserve individual and subjective responses. Use the distribution to locate the pattern. One dominant wrong option may reveal a misconception. An even spread across every option can indicate guessing or unclear wording. Strong performance paired with an instant response may indicate a cue or an obvious distractor.
You may encounter claims that learners should answer 85% of practice questions correctly. The often-cited Eighty Five Percent Rule study derived an optimum for a class of gradient-based learning models and demonstrated it with artificial neural networks and a model of perceptual learning. It does not establish one mandatory percentage for every classroom, compliance check, open response, or complex video lesson. Use percentages as evidence within your context, not as a universal design target.
Three worked difficulty profiles
The same equipment lesson can support several levels of evidence without adding tricks. Each profile below changes the task in a named way and keeps the language direct.
Foundation check
- Audience
- A new employee after the first equipment demonstration
- Evidence target
- Recognize the condition that prevents a safe restart
- Question
- Which visible reading shows the system is not ready to restart?
- Response design
- Multiple choice with three parallel readings from the same panel
Integrated check
- Audience
- A learner who has completed the full isolation sequence
- Evidence target
- Connect two signs of stored energy before choosing the next action
- Question
- Name the two observations that require another isolation step.
- Response design
- Short response, followed by evidence-specific feedback
Transfer check
- Audience
- An experienced technician preparing for independent work
- Evidence target
- Apply the isolation principle to unfamiliar equipment
- Question
- Which source must be isolated first in this new setup, and what evidence supports that choice?
- Response design
- Open response reviewed against a stated decision and evidence criterion
Notice that the final profile is not harder because it has more words. It is harder because the learner must transfer the principle, choose an action, and justify the decision with less support. The response also needs a review criterion before it can support a score.
A practical calibration workflow
Name the audience and use
State who will answer, what they have learned, and whether the result supports practice, diagnosis, progression, or a consequential decision.
Define the performance evidence
Describe the decision, explanation, calculation, or action that should count as understanding outside the video.
Write a clear baseline question
Create the simplest fair prompt that measures the outcome before attempting to raise or lower its challenge.
Profile the four levers
Record the evidence distance, reasoning move, response production, and support level in plain language.
Adjust one lever
Change one source of cognitive demand while keeping the wording, access, and required evidence fair.
Set response conditions separately
Choose points, timing, pausing, and retake rules for their own purposes instead of using them as substitutes for rigor.
Preview every path
Answer correctly and incorrectly on desktop and phone, inspect feedback, and confirm the relevant video evidence remains available.
Pilot and compare with the prediction
Review correct rates, option distributions, response explanations, and learner interpretation before deciding what to revise.
Build and test the complete video quiz
Follow the verified workflow for adding questions, configuring response behavior, previewing each path, publishing, and reviewing results.
Strengthen the stem and answer choices
Keep difficulty focused on the intended evidence by writing clear stems, defensible keys, and appropriate response formats.
The final difficulty audit
Aligned
The task asks for the same kind of thinking required by the learning outcome.
Fair
Every fact needed for a defensible answer has been taught or deliberately established as prior knowledge.
Interpretable
A wrong response points to a plausible knowledge gap rather than confusing wording or missing context.
Proportionate
The reading, typing, and navigation burden match the skill being measured.
Calibrated
The expected result is recorded, then compared with actual response evidence from the intended audience.
Teachable
Feedback or review can explain the gap and prepare the learner for the next attempt or case.
Ask one final diagnostic question: if a learner gets this wrong, what specific explanation will you be able to defend? If the only answer is “the question was hard,” the design still needs work. A calibrated question produces an interpretable gap, not just a lower score.
Research sources
- American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, Standards for Educational and Psychological Testing. The guide's alignment and interpretation advice follows its validity framework.
- National Board of Medical Examiners, NBME Item-Writing Guide. Its chapters on irrelevant difficulty, item flaws, and percent-correct statistics inform the calibration and diagnostic sections.
- Pyc and Rawson, Testing the retrieval effort hypothesis. The article supports the discussion of effortful successful retrieval and later memory.
- Roediger and Karpicke, Test-Enhanced Learning. Their experiments support the role of retrieval practice in delayed retention.
- Wilson and colleagues, The Eighty Five Percent Rule for optimal learning. The article is cited with its task and model boundaries rather than treated as a universal classroom target.
FAQ
What is a good difficulty level for a video question?
A good level asks the intended learner to perform the thinking required by the outcome with enough effort to expose understanding, but without ambiguity, missing evidence, or irrelevant reading burden. The same question can be appropriate for an experienced employee and too difficult for a new starter, so calibrate for a named audience and purpose.
What percentage of learners should answer a video question correctly?
There is no universal target. The expected percentage depends on whether the question introduces practice, checks prerequisite knowledge, certifies mastery, or diagnoses a misconception. Compare the observed result with your expectation, then inspect answer choices, timing, feedback, and individual responses before changing the difficulty.
Is a higher-order question always harder?
No. Difficulty depends on the learner's knowledge, the familiarity of the case, the cues provided, and the response required. An expert may find a realistic diagnosis easier than recalling an obscure label. Use reasoning verbs to align the task with the outcome, not as automatic difficulty ratings.
Do time limits make video questions more rigorous?
A time limit adds speed pressure. That is appropriate only when speed is part of the real performance, such as rapid hazard recognition. Otherwise it can measure reading speed, motor speed, language fluency, or anxiety instead of the knowledge you intended to assess.
How can I make an easy video question more challenging?
Change one meaningful lever at a time. Ask the learner to combine evidence from two moments, apply the rule to a new case, produce an answer rather than recognize one, or respond with fewer instructional cues. Keep the wording clear and retain every fact needed for a defensible answer.
How can I tell whether a video question is too difficult?
Look beyond a low correct rate. Check whether one wrong option captures a genuine misconception, whether responses are scattered across every option, whether the required evidence appeared before the prompt, and whether the wording or interface adds unrelated burden. Pilot with people from the intended audience and ask them to explain how they interpreted the task.
Build a question that is challenging for the right reason
Start with a supported, embeddable public or unlisted YouTube video, or upload your own file. Define the learner and evidence first, adjust one difficulty lever, then preview and pilot the complete learning moment.
Get started free