The exercise is over and the scoreboard has a winner. The programme owner still needs to know whether the team investigated accurately, made sound decisions and can defend the relevant systems with less help next time.
A score may answer part of that question. Its meaning depends on what was rewarded, what evidence was available and how much support participants received. Evaluation needs to explain those conditions as well as the result.
The expert philosophy behind ObsidianCorps' exercise work values learning over winning. Challenge should reveal limitations and create opportunities to improve, with difficulty matched to objectives and participant maturity. That does not mean every participant must fail, or that practical defensive performance is secondary.
The methods below are recommendations for planning an assessment. They are not a statement of an established ObsidianCorps scoring standard, historical participant outcomes or Scenarium product features.
Decide what the exercise is meant to demonstrate
An assessment begins with the objective. “Respond to an incident” can cover many different capabilities. A more precise objective might ask whether participants can investigate suspicious access, select a justified containment action and verify the effect on a required service.
That illustrative objective has several parts. Fast containment would not necessarily demonstrate accurate investigation. A correct explanation would not demonstrate that a configuration change was actually made. A technically effective change would not establish that the team understood its operational consequences.
Choose which parts matter for this exercise and what evidence each needs. Make those decisions before assigning points. Otherwise, the available scoreboard may end up deciding the objective on behalf of the commissioner.
For an exercise provider, this is also a scoping decision. Recording detailed observations, checking service behaviour and reviewing decisions require planned responsibilities. They should not be added as an afterthought when the client asks what the final score means.
Assess practical defence, decisions and learning separately
Use separate observations for different capabilities, then explain how they relate. A proposed assessment plan might look like this:
| Area |
Example question |
Possible evidence |
Limit to acknowledge |
| Practical defence |
Did the participant investigate and implement an appropriate action? |
Relevant logs, changes and independent service checks |
Only the implemented environment was observed |
| Decisions |
Did the team use available evidence and acknowledge uncertainty? |
Decision records and observer notes |
The record may omit reasoning that was not expressed |
| Coordination |
Did relevant information reach the role that needed it? |
Briefings, handoffs and requests for assistance |
Simulated external roles do not prove real external response performance |
| Learning |
Can the participant explain and apply feedback? |
Debrief reasoning and a comparable later task |
Familiarity with the exact scenario may affect repeat performance |
This is a recommended structure, not a mandatory list for every exercise. Select the areas that fit the objective. A focused technical drill may need detailed practical evidence; a crisis decision exercise may place greater emphasis on authority, uncertainty and coordination.
Research on learning assessment can inform that choice. JYVSECTEC's account of measuring learning in a cyber security exercise describes using knowledge categories and participant questionnaires. A questionnaire is one possible perspective. It should not be presented as direct proof that a participant can perform an unobserved defensive task.
Make the criteria descriptive before making them numerical
A useful rubric tells an observer what to look for. Start with descriptions that distinguish meaningful performance, then decide whether numerical weights add value.
For a hypothetical containment task, the criteria might require participants to identify the evidence behind the decision, involve the appropriate authority, perform the permitted action and verify the effect on the agreed service. These criteria express the objective more clearly than a single completion flag.
Consider three descriptive outcomes: demonstrated independently, demonstrated with support, and not yet demonstrated. Record the observation that justifies the category. “Demonstrated with support” should say what help was given; “not yet demonstrated” should distinguish an unsuccessful attempt from no opportunity to attempt the task.
If a score is required, agree weights and thresholds for this exercise rather than presenting them as universal measures of competence. Document how partial completion, alternative valid approaches and assistance affect the score. Participants should not be penalised simply for using a defensible route the scenario author did not anticipate.
Check the effect of defensive actions
Practical capability includes understanding what an action changed. A team that applies a control should have an opportunity to verify its effect and identify any consequence relevant to the objective.
For example, an illustrative exercise might require both restricting suspicious access and preserving a particular authorised workflow. Blocking all access may satisfy one condition while violating the other. An evaluation that checks only whether the suspicious connection stopped would miss that trade-off.
Recommend an independent observation of the required service after the action. Record the evidence the participant used and the checks they performed. If the service is represented by facilitator information instead of implemented behaviour, say so in the evaluation.
The environment-design brief should make these observations possible. Evaluation cannot recover evidence that the range was never built to expose.
Interpret timing in context
Detection and response times can be useful when their start and end points are defined. They become misleading when different participants encounter different information or when a system fault delays one team's progress.
Decide what starts the clock: the technical event, the first available indication or an explicit exercise prompt. Record the event consistently. If the facilitator pauses the exercise or supplies a clue, preserve that context.
Avoid treating speed as a substitute for sound reasoning. A fast but unsupported conclusion and a slower investigation that correctly separates competing explanations demonstrate different things. Which is appropriate depends on the agreed objective and operational context.
For a combined tabletop and technical exercise, also distinguish simulated time from elapsed exercise time. A scenario can represent a longer disruption within a shorter session, but evaluation should not confuse a compressed story deadline with measured real-world performance.
Separate participant limitations from exercise faults
A participant may be unable to complete a task because a required log is missing, a permission is wrong or an instruction contradicts the environment. These are exercise conditions that need their own record.
Recommend allowing an observer to classify a task as not assessable when the environment did not provide a fair opportunity. Explain why, and identify whether the problem affected other observations. Do not silently convert an environment failure into a zero score.
Participant guidance also needs context. A prompt that clarifies the exercise rules is different from a prompt that reveals the investigative answer. Both may be useful for learning, but they support different conclusions about independent performance.
The article on coordinating tabletop and cyber range activity discusses how facilitators can keep these interactions coherent during execution.
Use the debrief to test explanations
Ask participants to reconstruct an important decision using the evidence available at the time. What did they know, what remained uncertain and why did they choose that action? Compare the explanation with the observations without turning the discussion into a search for blame.
A useful debrief can identify a misunderstood dependency, a missing procedure or a skill that needs practice. Preserve effective actions as well as limitations. The aim is to understand what should be repeated, changed or explored further.
For each priority learning point, recommend an owner and an appropriate next step. That might be a focused practical task, a revised handoff procedure or a correction to the exercise itself. Do not assume all findings require more training.
Look for learning in a comparable later task
Repeating the same exercise can show familiarity with its details. A comparable task can provide a more useful opportunity to see whether a participant applies the reasoning in a changed situation.
Keep the objective comparable and document relevant differences, including difficulty, tools and assistance. Improvement is a conclusion that needs evidence; it should not be inferred from attendance or a positive reaction to the session.
Even a well-designed follow-up remains an exercise observation. It does not guarantee how the participant or organisation will perform in a real incident. Reporting the limits makes the assessment more useful to the people deciding what to do next.
ObsidianCorps provides cyber exercise scoring and evaluation expertise alongside design, engineering and delivery. If you need to connect an exercise objective with meaningful evidence of defence, decisions and learning, Contact us.