S5.5 and S5.7 Difficulty distribution and blueprint bank

S5.5 and S5.7 are not adjacent sub-stage numbers. S5.6, quiz and exam assembly, sits between them in the project notes’ own numbering. Its own content is different enough, and rich enough, that this guide gives it two pages of its own instead: S5.6 Quiz and exam assembly and the delivery-formats page. This page covers S5.5 and S5.7 together because they are two tightly linked halves of one idea: fix the bank’s own difficulty mix first, then weight its domain coverage against the blueprint.

Outcome

At the end of these two sub-stages, a whole item bank’s own difficulty mix is fixed up front and mapped to Bloom’s levels. The bank’s own domain coverage also mirrors the blueprint’s own weights: the same arithmetic S1.3 Draft the blueprint already defines for its own bank of test items, applied here at Stage 5’s own larger scale.

Where it fits

These two sub-stages take in finished stems from S5.3 and S5.4 Distractors and format rules, and the blueprint S1.3 approved. They hand a checked bank on to S5.6 Quiz and exam assembly, where the bank’s own stems are assembled into a deliverable quiz.

Why this way

Fixing the difficulty mix and the domain weighting before assembly, rather than checking only after a bank is finished, catches a gap early. A gap in one domain or one difficulty tier is caught while there is still time to write the stems that would close it.

Steps

Step Who Basis
Fix the bank’s own difficulty split at 30% easy, 50% medium, 20% hard, mapped to Bloom’s levels Agent documented
Weight each domain’s own target item count from the blueprint Agent documented
Run the bank composition check Script suggested
Approve the bank’s own composition before assembly Person suggested

The difficulty split and its Bloom crosswalk are documented: the project notes fix them exactly, and this guide has already stated the split publicly elsewhere (see below). The blueprint arithmetic is also documented, not this guide’s own invention. The project notes describe matching a bank’s own item counts to a blueprint’s weights, the same way S1.3’s own Parameters table already defines it, under its “Bank-size arithmetic” row. The section below says how this differs from a concept-count check. The bank composition check itself, and the judgment call of approving a finished composition, have no described precedent, so both are marked suggested.

The difficulty split and the Bloom crosswalk

Pipeline overview already states: “The 30/50/20 split is one of the reference implementation’s parameters, that project’s choice and not a universal rule.” This page reuses that exact split and its own crosswalk onto Bloom’s six-level scale (documented, the reference implementation’s own mapping of that already-public split onto Bloom’s levels):

  • Remember and Understand map to easy (30 percent).
  • Apply and Analyze map to medium (50 percent).
  • Evaluate and Create map to hard (20 percent).

A separate, industry-generic range

A separate, industry-generic exam-development guide gives a wider range for the same three-tier idea, not tied to any one certification:

  • Roughly 20 to 30 percent foundational (Remember and Understand).
  • Roughly 50 to 60 percent intermediate (Apply and Analyze).
  • Roughly 15 to 20 percent advanced (Evaluate and Create).

Keep the two figures distinct. The industry-generic range above is a general reference point, useful when no project-specific split exists yet. The reference implementation’s own fixed 30/50/20 is one project’s chosen instance inside that kind of range, not the same figure restated twice.

Blueprint arithmetic applied to a bank of test items

S1.3’s own Parameters table defines a “Bank-size arithmetic” row (suggested, at that sub-stage’s own small scale): “A domain’s weight, times the bank size, divided by 100, should be a whole number.” The running example’s own blueprint already applies this: four domains weighted 25, 30, 25 and 20, against a bank of 20 stems, gives target counts of 5, 6, 5 and 4.

This is not a mechanism first used for counting concepts and then carried over to items. The blueprint’s own bank-size field already counts stems, the finished test items, at both scales. Stage 5 reuses the same mechanism, not a new one, at its own larger scale.

Without naming the real target certification anywhere: a real project’s own blueprint can weight a much larger number of domains unevenly across many more skills than the running example’s four domains and twelve objectives. This guide demonstrates the same arithmetic on the running example’s own small, real blueprint instead, because the mechanism is the same regardless of scale. Multiply a domain’s weight by the bank size, then divide by 100: a domain’s own real item count should land close to that number, no matter how many domains a real program’s blueprint weights.

Worked illustration

The running example’s own real, already-published blueprint (scripts/sample_data/git_basics/blueprint.json, read-only, not copied here) has four domains, weighted 25, 30, 25 and 20. Applied to a bank of 20 items, the arithmetic above gives:

Domain Weight Target (weight * 20 / 100)
D1 Snapshots and history 25 5.0
D2 Branching and merging 30 6.0
D3 Working with remotes 25 5.0
D4 Recovering and collaborating safely 20 4.0

A small, invented bank composition (bank/composition.json, this guide’s own suggested illustration) splits each domain’s own target across the three difficulties, so the bank-wide split also lands on 30/50/20:

Domain Easy Medium Hard Total
D1 2 2 1 5
D2 2 3 1 6
D3 1 3 1 5
D4 1 2 1 4
Bank-wide 6 (30%) 10 (50%) 4 (20%) 20

Every domain’s own real count matches its target exactly, and the bank-wide difficulty split matches 30/50/20 exactly. So bank_composition_check.py reports no warnings on this file, shown below.

Artifacts and formats

This sub-stage’s own artifact is a bank-composition record: item counts by domain and by difficulty. Its JSON shape (bank/composition.json) is a single object with one key, domains, mapping each domain id to an object holding an easy, a medium and a hard count. This is this guide’s own suggested shape. The project notes give no machine-checkable format for this record.

Prompts

P-S5-04 Propose a bank composition proposes a per-domain, per-difficulty item-count table from a blueprint excerpt and a target bank size, using the arithmetic above. It is written for this guide and has not been run against any model in this build.

Scripts

X-S5-04 Bank composition check reads a bank composition file and the blueprint it is meant to match. For each blueprint domain, it checks whether the composition’s own real item count is within 1 of weight * bank_size / 100. It also checks whether the bank-wide easy, medium and hard split is within 10 percentage points of 30/50/20. A composition gap is a person’s judgment call, not a hard failure. This is the same limit Stage 2’s readability report already sets for its own two scores: this script never exits 1 on its own, and exits 0 whenever it can read both files.

Run it from the repository root, with the composition file and the blueprint file as its two arguments:

python3 -B scripts/s5/bank_composition_check.py \
    scripts/sample_data/git_basics_stage5/bank/composition.json \
    scripts/sample_data/git_basics/blueprint.json
domain by target and actual item count (bank_size=20):
D1  weight=25 target=5.0 actual=5
D2  weight=30 target=6.0 actual=6
D3  weight=25 target=5.0 actual=5
D4  weight=20 target=4.0 actual=4
difficulty split (bank_size=20): easy=6 (30.0%) medium=10 (50.0%) hard=4 (20.0%) target=30/50/20
domains=4 warnings=0

A clean run only shows that this one small, invented composition happens to match its blueprint and its difficulty split exactly. It is not proof that every stem behind those counts is itself sound, only that the counts add up.

Break it on purpose: copy the file, then change domain D2’s own hard count from 1 to 5. It now drifts past its own target, and the bank-wide hard share balloons past its own 10-point tolerance. Run the check again on the copy:

warning domain-count: domain D2 weight 30 gives bank_size * weight / 100 = 7.2 target, actual count is 10 (+2.8 away)
warning difficulty-split: bank-wide split is easy=25.0% medium=41.7% hard=33.3%, more than 10 points off the reference implementation's 30/50/20
domain by target and actual item count (bank_size=24):
D1  weight=25 target=6.0 actual=5
D2  weight=30 target=7.2 actual=10
D3  weight=25 target=6.0 actual=5
D4  weight=20 target=4.8 actual=4
difficulty split (bank_size=24): easy=6 (25.0%) medium=10 (41.7%) hard=8 (33.3%) target=30/50/20
domains=4 warnings=2

The one-domain change trips two warnings at once, the same way a single weight edit trips two checks on S1.3’s own blueprint check. Adding four items to one domain moves that domain away from its own target, and pulls the whole bank’s hard share away from 30/50/20, since the added items were all hard. See Reading exit codes on the Stage 1 index for what the exit code means; here it stays 0, since a composition gap is always a warning on this script, never a failure.

Definition of done

  • Every domain’s own real item count is within 1 of its blueprint target.
  • The bank-wide difficulty split is within 10 percentage points of 30/50/20.
  • The bank composition check has run, and a person has looked at any warning it printed.
  • A person has approved the bank’s own composition before assembly.

Common failures

  • A bank whose difficulty mix drifts from 30/50/20 as stems are added piecemeal, one chapter at a time, with no one checking the running total.
  • A domain’s own item count that does not divide evenly against its blueprint weight, the same arithmetic warning S1.3 already gives for its own bank of test items.
  • Conflating the generic 20-30/50-60/15-20 range above with the reference implementation’s own fixed 30/50/20, as though they were the same figure restated twice.

Adapting to your platform

  • llm: proposes the per-domain, per-difficulty table from a blueprint excerpt and a target bank size, using P-S5-04.
  • file-read: reads the blueprint file the proposal, and the check, are both measured against.
  • shell: runs the bank composition check; without it, compute each domain’s target by hand and compare it to the real count.
  • human-approval: covers the approval step below; a platform with no built-in approval step still needs a person to read and record it somewhere, such as a shared document.

Where humans decide

  • Approving a finished bank’s own composition before assembly.
  • Whether a domain’s own item count that misses its target by more than 1 needs new stems, or is close enough to accept as is.
  • Whether a bank-wide difficulty split more than 10 points off 30/50/20 needs rebalancing before assembly, or reflects a deliberate choice for this particular program.

Next: S5.6 Quiz and exam assembly.


To the extent possible under law, copyright and related rights in this work are waived under CC0 1.0 Universal.

This site uses Just the Docs, a documentation theme for Jekyll.