From ba873d45ae04984340cec79bc80b76ba6f5394d9 Mon Sep 17 00:00:00 2001 From: DineshGirbide Date: Mon, 10 Aug 2026 16:59:39 +0530 Subject: [PATCH 1/4] =?UTF-8?q?feat(ccaf):=20align=20to=20exam=20guide=20v?= =?UTF-8?q?1.0=20=E2=80=94=20multiple-response=20items,=2030=20task=20stat?= =?UTF-8?q?ements,=20prep=20guide?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The revised CCAF exam guide (v1.0, effective July 2026, exam code CCAR-F) changes the item format and publishes a full objective index. Aligns the plugin to both. Multiple-response items. The guide specifies multiple-choice *and* multiple-response items, each stating how many responses to select, so a mock is now 45 `select: 1` + 11 `select: 2` + 4 `select: 3` (25% multiple-response, enforced exactly; choose-three capped at 5). Items carry a `select:` line that survives into the key-free questions file, because the administering skill needs the count without ever seeing the key. `init` rejects any item whose `select:` disagrees with its `answer_key`, repeats a letter, or lists letters out of A-D order — and holds a pre-filled `user_answer` to the same shape, since scoring is exact string comparison and an unsorted "DB" would otherwise score a correct "BD" as wrong. `record` accepts a letter set in any case or order and normalizes it, so scoring is exact set equality: all-or-nothing, no partial credit. Multiple-response items render with `multiSelect: true`; a wrong response count is re-asked once. The key-spread guard now covers the single-answer items only, with a band proportional to their count. Known divergence, disclosed in the README rather than hidden: every item has four options (A-D) because AskUserQuestion renders at most four. That makes a choose-three item easier than a five-option one would be, which is why choose-three stays a minority. The 30 task statements are now real. The blueprint formalizes D1.1-D5.6 with self-authored descriptions naming the exact identifiers items use (Task / allowedTools, AgentDefinition, fork_session, PostToolUse, context: fork, allowed-tools, argument-hint, Explore subagent, @import, ~/.claude.json, --json-schema, Pydantic, state manifests, the interview pattern, the Edit->Read+Write fallback). This closes a dangling reference: ccaf-tutor and ccaf-check-author already cited these codes against a blueprint that never defined them. Scenario facts corrected. Case briefs re-authored (still our own prose) around the guide's actual identifiers get_customer / lookup_order / process_refund / escalate_to_human, the 80%+ first-contact-resolution target, and four research subagents including a separate report generator. The previously invented names are gone. Bank grown 12 -> 24. Twelve new self-authored anchors covering D1.2, D1.4, D1.5, D2.1, D2.4, D3.3, D3.4, D4.5, D4.6, D5.1, D5.2, D5.3 — five of them multiple-response so that format has anchors too. All 24 tagged with `task:` and `select:`. Six task statements still lack an anchor (D1.7, D2.5, D3.5, D4.2, D4.4, D5.4); flagged in the bank, not hidden. Reporting. `score` emits `pct=` per domain, matching the real score report's percent-correct-by-domain, presented as diagnostic-only since pass/fail is the total scaled score. New data/ccaf-prep-guide.md. Study routes, four hands-on exercises mapped to task statements, multiple-response answering strategy, and certification logistics (fee, Pearson VUE delivery, 14/30/90-day retake ladder, 12-month validity, free renewal assessment). The tutor can assign exercises at domain boundaries. Everything shipped stays self-authored: the guide's published facts are encoded, none of its prose is, and its own sample questions were rewritten as original items rather than quoted. The guide PDF is gitignored so it can never ship with the plugin. Fork-cost discipline: process creation is slow on some machines (~640ms/fork measured on Windows + AV), so normalize_answer is pure bash and each validation check is a single awk pass. Net effect is that `score` got faster than before this change despite the added checks. Co-Authored-By: Claude Opus 5 (1M context) --- .claude-plugin/marketplace.json | 2 +- .gitignore | 6 + ccaf/.claude-plugin/plugin.json | 4 +- ccaf/CLAUDE.md | 29 +- ccaf/README.md | 82 ++- ccaf/agents/ccaf-check-author.md | 79 ++- ccaf/commands/mock-exam.md | 22 +- ccaf/commands/practice.md | 14 +- ccaf/commands/prepare.md | 14 +- ccaf/data/ccaf-blueprint.md | 805 +++++++++++++++++---------- ccaf/data/ccaf-prep-guide.md | 210 +++++++ ccaf/data/ccaf-question-bank.md | 403 +++++++++++++- ccaf/scripts/ccaf-exam.sh | 240 ++++++-- ccaf/scripts/tests/ccaf-exam.test.sh | 306 ++++++++-- ccaf/skills/ccaf-exam/SKILL.md | 213 ++++--- ccaf/skills/ccaf-practice/SKILL.md | 197 ++++--- ccaf/skills/ccaf-tutor/SKILL.md | 98 ++-- docs/specs/ccaf-mock-exam.md | 4 +- 18 files changed, 2052 insertions(+), 676 deletions(-) create mode 100644 ccaf/data/ccaf-prep-guide.md diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 1d39aed..2340c87 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -17,7 +17,7 @@ { "name": "ccaf", "source": "./ccaf", - "description": "CCAF (Claude Certified Architect – Foundations) mock-exam readiness gate. /ccaf:mock-exam assembles a 60-question case-study-framed mock with machine-enforced domain weighting — every question generated fresh per attempt and independently verified — administers it 4 per screen with resumable progress, and scores it on the real 100–1000 band with a 720 pass line plus a per-domain breakdown." + "description": "CCAF (Claude Certified Architect – Foundations, exam code CCAR-F) readiness gate, aligned to exam guide v1.0. /ccaf:mock-exam assembles a 60-item case-study-framed mock mixing multiple-choice and multiple-response items, with machine-enforced domain weighting and item mix — every item generated fresh per attempt and independently verified — administers it 4 per screen with resumable progress, and scores it on the real 100–1000 band with a 720 pass line plus a per-domain percent breakdown. /ccaf:prepare teaches the 30 task statements turn by turn; /ccaf:practice drills chosen domains." }, { "name": "discovery", diff --git a/.gitignore b/.gitignore index 63d77d8..61f23e6 100644 --- a/.gitignore +++ b/.gitignore @@ -3,3 +3,9 @@ .claude/bee-insights/ .claude/ccaf-exam.local.md .claude/ccaf-exam.local.answers.md +.claude/ccaf-practice.local.md +.claude/ccaf-practice.local.answers.md + +# Reference PDFs (e.g. the CCAF exam guide) are read while authoring the plugin's +# self-authored content but must never be redistributed with it. +ccaf/*.pdf diff --git a/ccaf/.claude-plugin/plugin.json b/ccaf/.claude-plugin/plugin.json index c2c0347..cc56c67 100644 --- a/ccaf/.claude-plugin/plugin.json +++ b/ccaf/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "ccaf", - "version": "0.1.1", - "description": "CCAF (Claude Certified Architect – Foundations) mock-exam readiness gate. Assembles a 60-question case-study-framed mock with machine-enforced domain weighting — every question generated fresh per attempt and independently verified — administers it 4 per screen with resumable progress, and scores it on the real 100–1000 band with a 720 pass line plus a per-domain breakdown. Run /ccaf:mock-exam.", + "version": "0.2.0", + "description": "CCAF (Claude Certified Architect – Foundations, exam code CCAR-F) readiness gate, aligned to exam guide v1.0. Assembles a 60-item case-study-framed mock mixing multiple-choice and multiple-response items with machine-enforced domain weighting and item mix — every item generated fresh per attempt and independently verified — administers it 4 per screen with resumable progress, and scores it on the real 100–1000 band with a 720 pass line plus a per-domain percent breakdown. Learn with /ccaf:prepare, drill with /ccaf:practice, gate with /ccaf:mock-exam.", "author": { "name": "Incubyte" }, diff --git a/ccaf/CLAUDE.md b/ccaf/CLAUDE.md index 6b8c094..8ebad08 100644 --- a/ccaf/CLAUDE.md +++ b/ccaf/CLAUDE.md @@ -1,8 +1,13 @@ # CCAF: Claude Certified Architect – Foundations mock exam -A self-serve readiness gate for the CCAF certification. `/ccaf:mock-exam` administers a faithful -mock exam and reports a scaled /1000 score with the 720 pass line, so a candidate can check -readiness before booking the real (paid) exam. +A self-serve readiness gate for the CCAF certification (exam code `CCAR-F`). `/ccaf:mock-exam` +administers a faithful mock exam and reports a scaled /1000 score with the 720 pass line, so a +candidate can check readiness before booking the real (paid) exam. + +Aligned to exam guide **v1.0 (effective July 2026)**. The guide's published *facts* are encoded +here; none of its prose is. Every stem, option, explanation, scenario brief, task-statement +description, and exercise is self-authored for this plugin — keep it that way when editing, and do +not commit the guide PDF (it is gitignored). ## Layout @@ -13,9 +18,10 @@ readiness before booking the real (paid) exam. - `skills/ccaf-exam/SKILL.md` — the assemble → administer → score engine. - `skills/ccaf-practice/SKILL.md` — the domain-selection → assemble → administer → score engine for focused domain practice; uses a separate state file so it never conflicts with `/ccaf:mock-exam`. - `agents/ccaf-check-author.md` — mini-agent the tutor spawns to author one scenario check at a time. -- `data/ccaf-blueprint.md` — domains, weights, scenarios, the syllabus, scope lists, scoring. Shared curriculum for both commands. -- `data/ccaf-question-bank.md` — 12 self-authored reference questions; style/difficulty anchors only, never served in an exam (helper-enforced). -- `scripts/ccaf-exam.sh` — silent state helper (init / get / record / blanks / audit / score / clear); never use Write/Edit on the attempt files. `init` takes one payload (with keys) and splits it: questions file (write-once) + answers file (hot, ~60 lines) — so `record` rewrites only the tiny answers file and `get` output is key-free. Guards: `init` validates the payload — and, for 60-question exams, enforces the blueprint composition (domain quotas 16/11/12/12/9, 4 scenarios in contiguous sections each headed by its own `[[CASE:]]` brief — so a screen's brief always matches its questions — non-degenerate key spread) — and refuses to overwrite an in-progress attempt (unless `--force`); `record` takes one or more `--q/--answer` pairs atomically (one call per screen) and requires an in-progress attempt; `score` cross-validates the pair and requires `--partial` to score with unanswered questions. All writes serialize through a directory lock (stale locks are stolen), so mid-exam `record` calls run **in the background** while the next screen shows; the final screen records in the foreground and completion is verified before scoring. +- `data/ccaf-blueprint.md` — domains, weights, item composition, the 30 task statements (D1.1–D5.6), scenarios, case-study briefs, scope lists, scoring. Shared curriculum for all three commands, and the only authority for item content. +- `data/ccaf-question-bank.md` — 24 self-authored reference questions, each tagged with its `task:` and `select:`; style/difficulty/format anchors only, never served in an exam (helper-enforced). Five are multiple-response so that format has anchors too. +- `data/ccaf-prep-guide.md` — study routes, four hands-on exercises, multiple-response answering strategy, and certification logistics (fee, retakes, recertification). Read by the tutor; read by the exam skills only for post-result guidance, never for item content. +- `scripts/ccaf-exam.sh` — silent state helper (init / get / record / blanks / audit / score / clear); never use Write/Edit on the attempt files. `init` takes one payload (with keys) and splits it: questions file (write-once) + answers file (hot, ~60 lines) — so `record` rewrites only the tiny answers file and `get` output is key-free. Guards: `init` validates the payload — every item's `select:` count must equal its `answer_key` length, with distinct letters in A–D order — and, for 60-item exams, enforces the blueprint composition (domain quotas 16/11/12/12/9; exactly 15 multiple-response items with at most 5 choose-three; 4 scenarios in contiguous sections each headed by its own `[[CASE:]]` brief — so a screen's brief always matches its items; non-degenerate key spread across the single-answer items) — and refuses to overwrite an in-progress attempt (unless `--force`); `record` takes one or more `--q/--answer` pairs atomically (one call per screen), normalizes each answer set (uppercased, sorted, de-duplicated) so selection order never matters, and requires an in-progress attempt; `score` cross-validates the pair and requires `--partial` to score with unanswered items. All writes serialize through a directory lock (stale locks are stolen), so mid-exam `record` calls run **in the background** while the next screen shows; the final screen records in the foreground and completion is verified before scoring. - `scripts/tests/ccaf-exam.test.sh` — shell test harness for the data files + helper logic. ## Conventions @@ -27,5 +33,16 @@ readiness before booking the real (paid) exam. so the two modes never interfere. - Untimed, honor-system, fully offline. Self-serve: nothing is reported or persisted as history. The real exam's 120-minute budget is stated once up front for self-pacing; no time is ever captured. +- **Two item formats**, as the real exam has. `select: 1` is multiple-choice; `select: 2`/`3` is + multiple-response, rendered with `multiSelect: true` and stating its count in bold in the stem. + A full mock is 45 / 11 / 4. Every item has exactly four options A–D — `AskUserQuestion` renders no + more than four, which the README's fidelity table discloses rather than hides. +- Multiple-response scoring is **all-or-nothing**: the recorded set must equal the key exactly. - Scoring: `scaled = 100 + 15 × correct` (linear over the real 100–1000 band); pass = 720 (≥ 42/60). + Results always show per-domain correct/total **and percent**, labelled diagnostic-only — pass/fail + is the total scaled score, as on the real criterion-referenced exam. - The scaled score is an honest estimate, never presented as Anthropic's proprietary equating curve. +- Process creation can be slow on some machines (Windows + AV in particular), so hot paths in the + helper avoid gratuitous subprocesses — `normalize_answer` is pure bash, and each validation check + is a single `awk` pass rather than a loop of `grep`s. Keep it that way; `record` runs once per + exam screen. diff --git a/ccaf/README.md b/ccaf/README.md index d60d65a..48f21ac 100644 --- a/ccaf/README.md +++ b/ccaf/README.md @@ -2,7 +2,26 @@ CCAF is a Claude Code plugin that helps you **prepare for** and **mock-test** against the **Claude Certified Architect – Foundations** exam — right in your terminal — so you get an honest readiness verdict before you book the real (paid) exam. -**Why this exists.** Incubyte is having everyone get CCAF-certified, with a simple rule: *practice first, and only sit the real exam once you can reliably score 720+.* Instead of every engineer hand-rolling a quiz from the exam-guide PDF, this plugin makes that learn-and-gate flow a couple of commands, consistent for the whole team. All content is self-authored or publicly corroborated — no Anthropic exam material is reproduced. +**Why this exists.** Incubyte is having everyone get CCAF-certified, with a simple rule: *practice first, and only sit the real exam once you can reliably score 720+.* Instead of every engineer hand-rolling a quiz from the exam-guide PDF, this plugin makes that learn-and-gate flow a couple of commands, consistent for the whole team. + +Aligned to the published exam guide **v1.0 (effective July 2026, exam code `CCAR-F`)**: the same 60 items, the same five weighted domains, the same 30-objective index, and the same two item formats. **Every question, scenario brief, task-statement description, and exercise in this plugin is self-authored** — the guide's published *facts* (weights, item counts, the objective index, backend tool names, policies) are reflected, but none of its prose and no exam item is reproduced. + +## The real exam at a glance + +| | | +| --- | --- | +| Exam code | `CCAR-F` | +| Items | 60 — multiple-choice **and** multiple-response (each item states how many responses to select) | +| Structure | 4 scenarios drawn from a bank of 6 | +| Time limit | 120 minutes | +| Delivery | Proctored by Pearson VUE — online or at a test centre | +| Passing score | Scaled **720** on a 100–1000 range | +| Reporting | Pass/fail + scaled score + percent correct per domain | +| Fee | $125 USD per attempt | +| Validity | 12 months, renewable with a free non-proctored assessment | +| Retakes | 14 / 30 / 90-day waits after successive failures; max 4 attempts per rolling year | + +That retake ladder is the argument for this plugin: a failed attempt costs $125 *and* two weeks. Full logistics — booking, ID, accommodations, appeals, recertification — are in `data/ccaf-prep-guide.md`. **What you get — three commands, one loop:** @@ -23,17 +42,22 @@ CCAF is a Claude Code plugin that helps you **prepare for** and **mock-test** ag | Real exam rule | This mock | | --- | --- | -| 60 questions | ✅ 60 per attempt | -| 5 domains, weighted 27 / 18 / 20 / 20 / 15 | ✅ same distribution (D1=16, D2=11, D3=12, D4=12, D5=9) | +| 60 items | ✅ 60 per attempt | +| 5 domains, weighted 27 / 18 / 20 / 20 / 15 | ✅ same distribution (D1=16, D2=11, D3=12, D4=12, D5=9), machine-enforced | +| 30 task statements (D1.1–D5.6) | ✅ every item is written against one, and spread across them | | 4 of 6 scenarios, chosen at random | ✅ same | -| Questions organized around case studies | ✅ 4 case-study sections; the case brief stays visible on every screen | -| Single-select, 1 correct + 3 distractors | ✅ same | +| Items organized around case studies | ✅ 4 case-study sections; the case brief stays visible on every screen | +| Multiple-choice **and** multiple-response | ✅ 45 single-answer + 15 multiple-response (11 "Select TWO", 4 "Select THREE"), machine-enforced at 25% | +| Each item states how many to select | ✅ stated in bold in the stem; multi-select rendering; wrong count is re-asked once | +| Multiple-response scored all-or-nothing | ✅ the recorded set must match exactly — no partial credit | | No penalty for guessing | ✅ unanswered = incorrect | | Answers revisable before submit | ✅ ask to change any earlier answer mid-exam | | Scaled 100–1000, pass = 720 | ✅ `scaled = 100 + 15 × correct`; pass at ≥ 42/60 | -| 120-minute time limit | ❌ untimed by design — it shows the 120-min / ~2-min-per-question budget up front so you can self-pace | +| Percent correct per domain on the report | ✅ shown, and labelled diagnostic-only — pass/fail is the total scaled score | +| 5+ options on some multiple-response items | ⚠️ every item here has exactly 4 (A–D) — the terminal's question UI renders at most four options | +| 120-minute time limit | ❌ untimed by design — it shows the 120-min / ~2-min-per-item budget up front so you can self-pace | -Two things it does **not** replicate, on purpose: Anthropic's proprietary scaled-scoring curve (impossible — the score here is a transparent, linear *estimate*, clearly labelled) and the 120-minute clock (deliberate — the mock is resumable and honor-system; time yourself if you want realistic conditions). Treat 720+ as a readiness signal, not a guarantee. +Three things it does **not** replicate, and why. Anthropic's proprietary scaled-scoring curve is unpublished, so the score here is a transparent linear *estimate*, clearly labelled. The 120-minute clock is deliberately absent — the mock is resumable and honor-system; time yourself if you want realistic conditions. And multiple-response items are capped at four options by the terminal UI, which makes a "Select THREE" item easier than a five-option one would be (you need only spot the single option that doesn't belong), so the mock keeps those to a minority. Treat 720+ as a readiness signal, not a guarantee. ## The flow @@ -41,20 +65,22 @@ Two things it does **not** replicate, on purpose: Anthropic's proprietary scaled /ccaf:mock-exam | v - [ ASSEMBLE ] Pick 4 of 6 case studies; generate all 60 questions fresh + [ ASSEMBLE ] Pick 4 of 6 case studies; generate all 60 items fresh | (anchored to the reference bank, each independently verified, | A–D shuffled); group into 4 case-study sections; freeze the - | exam to ~/.claude/ccaf-exam.local.md. The 16/11/12/12/9 - | domain split is machine-enforced at write time (a - | mis-weighted exam is refused), and the composition is shown - | to you up front. + | exam to ~/.claude/ccaf-exam.local.md. Both the 16/11/12/12/9 + | domain split and the 45/11/4 item-format mix are + | machine-enforced at write time (a mis-weighted exam, or one + | whose answer key disagrees with its stated response count, is + | refused), and the composition is shown to you up front. v - [ ADMINISTER ] 4 questions per screen, case brief always visible. Each screen - | saves atomically in the background while the next one shows — - | no save-wait between screens. Quit any time; re-run to resume. + [ ADMINISTER ] 4 items per screen, case brief always visible; "Select TWO" + | items render as multi-select. Each screen saves atomically in + | the background while the next one shows — no save-wait between + | screens. Quit any time; re-run to resume. v - [ SCORE ] Scaled /1000, PASS/FAIL at 720, per-domain breakdown, and an - honest "this is an estimate" disclaimer. + [ SCORE ] Scaled /1000, PASS/FAIL at 720, per-domain percent breakdown, + and an honest "this is an estimate" disclaimer. ``` ## Install @@ -94,7 +120,7 @@ Untimed and conversational. It's **stateless** — Claude Code's native session /ccaf:practice ``` -Select one or more domains to focus on, then choose how many questions you want (10, 20, or 30). Questions are drawn proportionally from the selected domains using the real blueprint weights. At the end you get a per-domain bar chart — no overall score or PASS/FAIL verdict — and a targeted recommendation for any domain that needs work. +Select one or more domains to focus on, then choose how many questions you want (10, 20, or 30). Items are drawn proportionally from the selected domains using the real blueprint weights, with a quarter of them multiple-response so you drill that format too. At the end you get a per-domain bar chart with percentages — no overall score or PASS/FAIL verdict, since a partial session isn't weighted like a real form — and a targeted recommendation for any domain that needs work. ```bash /ccaf:practice fresh @@ -112,7 +138,7 @@ Discard any in-progress or completed practice attempt and start a new domain sel /ccaf:mock-exam ``` -Answer the questions four to a screen. When you finish, you get your scaled score, a PASS/FAIL at 720, and a domain-by-domain breakdown so you know where you're weak. A FAIL points you back to `/ccaf:prepare ` for targeted practice. +Answer the items four to a screen. Most ask for one response; 15 of the 60 ask for two or three and say so in the stem — those render as multi-select and are scored all-or-nothing, so pick every correct option and nothing else. When you finish, you get your scaled score, a PASS/FAIL at 720, and a domain-by-domain percent breakdown so you know where you're weak. A FAIL points you back to `/ccaf:prepare ` for targeted practice. ```bash /ccaf:mock-exam fresh @@ -126,10 +152,12 @@ Discard any in-progress or completed attempt and assemble a brand-new exam. ## How scoring works -- Raw `correct` = questions answered correctly (unanswered count as incorrect). +- Raw `correct` = items answered correctly (unanswered count as incorrect). +- **Multiple-response items are all-or-nothing**: the set you select must match the key exactly. One right and one wrong scores the same as zero right — there is no partial credit, as on the real exam. Selection *order* never matters (the helper normalizes each set). - `scaled = 100 + 15 × correct` — a linear mapping over the real 100–1000 band (equivalently `100 + round(correct ÷ 60 × 900)`; since `900 ÷ 60 = 15`, no rounding is needed). - **Pass** iff `scaled ≥ 720`, i.e. **≥ 42 of 60** correct. -- A per-domain breakdown (correct / total per D1–D5) accompanies every result. +- A per-domain breakdown (correct / total **and percent** per D1–D5) accompanies every result, mirroring the real score report. Like the real exam, those percentages are diagnostic only — pass/fail is decided by the total scaled score. +- The real exam is **criterion-referenced**: you're measured against a fixed standard set by a formal standard-setting study, not graded against other candidates. 720 is a fixed bar. ## What's inside @@ -151,8 +179,9 @@ ccaf/ │ └── ccaf-practice/ │ └── SKILL.md # practice engine: domain-select → assemble → administer → score (internal) ├── data/ -│ ├── ccaf-blueprint.md # public exam mechanics + self-authored syllabus, scenarios, scoring -│ └── ccaf-question-bank.md # 12 self-authored reference questions (anchors only — never served) +│ ├── ccaf-blueprint.md # exam mechanics + item composition + self-authored 30-task-statement syllabus, scenarios, scoring +│ ├── ccaf-question-bank.md # 24 self-authored reference questions (anchors only — never served) +│ └── ccaf-prep-guide.md # study routes, 4 hands-on exercises, multiple-response strategy, certification logistics ├── scripts/ │ ├── ccaf-exam.sh # silent state helper (init / get / record / score / clear) │ └── tests/ @@ -164,9 +193,10 @@ ccaf/ ## Notes -- **Question sourcing.** Every question in every attempt is **generated fresh** from the blueprint syllabus and passes an independent verifier (re-solve cold, plausible distractors, shuffled positions) before being served. The 12 self-authored questions in the bank are style/difficulty anchors only — they never appear in an exam (machine-enforced), because the bank ships in this repo with answers, and re-serving readable questions would inflate your readiness signal. -- **How `prepare` teaches.** The tutor reads the same blueprint as its curriculum, teaches one task statement per turn, and verifies by retrieval. Its apply-to-scenario checks are authored on demand by a small `ccaf-check-author` subagent (built from the syllabus anti-patterns), keeping the main teaching thread lean. Nothing is written to disk. -- **Roadmap.** v1 generates everything per attempt anchored to a small reference bank; growing a larger verified anchor bank and verifying the full lifecycle end-to-end are the next steps (tracked in `docs/specs/ccaf-mock-exam.md`). Per-domain focused practice is now available via `/ccaf:practice`. +- **Question sourcing.** Every item in every attempt is **generated fresh** from the blueprint syllabus and passes an independent verifier (re-solve cold, exactly the required number of defensible options, plausible distractors, shuffled positions) before being served. The 24 self-authored questions in the bank are style/difficulty/format anchors only — they never appear in an exam (machine-enforced), because the bank ships in this repo with answers, and re-serving readable questions would inflate your readiness signal. +- **How `prepare` teaches.** The tutor reads the blueprint as its curriculum, teaches one of the 30 task statements per turn, and verifies by retrieval. Its apply-to-scenario checks — single-answer or "Select TWO" — are authored on demand by a small `ccaf-check-author` subagent (built from the syllabus anti-patterns), keeping the main teaching thread lean. At domain boundaries it can assign one of the prep guide's four hands-on exercises instead of another quiz. Nothing is written to disk. +- **Roadmap.** Everything is generated per attempt against a 24-question anchor bank. Next: anchors for the six task statements that still lack one (D1.7, D2.5, D3.5, D4.2, D4.4, D5.4 — listed in the bank), per-task-statement result diagnostics so a score report can name the exact objectives you missed, and end-to-end lifecycle verification. Tracked in `docs/specs/ccaf-mock-exam.md`. +- **Provenance.** The plugin encodes the exam guide's published *facts* (item count, weights, the objective index, the scenarios' backend tool names and targets, and program policies) and nothing else from it: every stem, option, explanation, scenario brief, task-statement description, and exercise is written for this plugin. The guide PDF itself is deliberately not committed. ## License diff --git a/ccaf/agents/ccaf-check-author.md b/ccaf/agents/ccaf-check-author.md index aa15b6e..be6b24e 100644 --- a/ccaf/agents/ccaf-check-author.md +++ b/ccaf/agents/ccaf-check-author.md @@ -1,11 +1,12 @@ --- name: ccaf-check-author description: >- - Authors ONE fresh single-select knowledge-check question for a given CCAF task - statement and difficulty, then returns it with a compact answer key. Spawned by - the ccaf-tutor skill (the /ccaf:prepare engine) for apply-to-scenario checks so the - main teaching thread stays lean — it never sees the authoring rationale, only the - question. Not user-invokable directly. + Authors ONE fresh knowledge-check question — multiple-choice or multiple-response, + matching the real exam's two item formats — for a given CCAF task statement and + difficulty, then returns it with a compact answer key. Spawned by the ccaf-tutor + skill (the /ccaf:prepare engine) for apply-to-scenario checks so the main teaching + thread stays lean — it never sees the authoring rationale, only the question. Not + user-invokable directly. Context: The tutor just taught D1.4 (programmatic enforcement vs prompt-based ordering) and wants a scenario check. @@ -19,11 +20,20 @@ description: >- Context: The tutor wants to raise difficulty after the learner aced two D2 checks. user: "(tutor delegates) Author a hard check for D2.3 (tool distribution & tool_choice), scenario multi-agent-research. Taught: D2.1, D2.3." - assistant: "Authoring a harder D2.3 scenario that turns on the least-privilege tradeoff (scoped verify_fact tool vs full tool access vs end-of-pass batching vs speculative caching), all four options plausible to a partial-knowledge candidate. Returning question + key=that one option + why the three distractors fail." + assistant: "Authoring a harder D2.3 scenario that turns on the least-privilege tradeoff (scoped verify_fact tool vs full tool access vs end-of-pass batching vs speculative caching), all four options plausible to a partial-knowledge candidate. Returning question + SELECT: 1 + key=that one option + why the three distractors fail." Difficulty is expressed by making distractors closer and the tradeoff finer, not by adding out-of-scope trivia. The agent stays strictly in-scope per the blueprint. + + + Context: The learner has been losing multiple-response items, so the tutor asks for one in that format. + user: "(tutor delegates) Author a medium check for D5.2, scenario customer-support, format multi. Taught: D5.1, D5.2." + assistant: "Authoring a choose-two D5.2 item on escalation triggers: two independently-correct triggers (explicit request for a human; a request the policy is silent on) and two distractors from the unreliable-proxy anti-patterns (sentiment, self-reported confidence). Stem ends in **Select TWO.** Returning SELECT: 2, KEY: two letters in A-D order." + + For `multi` the two correct options must be independently correct and express distinct ideas — not one idea said twice. The count goes in the stem because the real exam always states it. + + model: inherit color: cyan tools: ["Read", "Glob"] @@ -39,55 +49,72 @@ Nothing you reason about leaks to the main thread — only your final output doe statement, in the style and difficulty of the real CCAF exam, then hand back a self-contained question plus a key the tutor can grade against. +You may also receive a **format**: `single` (one correct option) or `multi` (two correct options). +Default to `single` when it is not specified. + **Authority (read it; do not invent content):** - `${CLAUDE_PLUGIN_ROOT}/data/ccaf-blueprint.md` — the syllabus. Find the task statement - (e.g. `D1.4`), read its Knowledge-of / Skills-in bullets, and read the **anti-patterns** - flagged there. The blueprint's in-scope / out-of-scope lists are hard boundaries. -- `${CLAUDE_PLUGIN_ROOT}/data/ccaf-question-bank.md` — the 12 self-authored seed questions. - Use them as **style and difficulty anchors only**. Never reproduce one verbatim; author fresh. + (e.g. `D1.4`) in the 30-task-statement section, read what it says a candidate must be able to do + and the exact identifiers it names, and read its domain's **common mistakes** list — that is your + distractor source. The blueprint's in-scope / out-of-scope lists are hard boundaries. +- `${CLAUDE_PLUGIN_ROOT}/data/ccaf-question-bank.md` — the 24 self-authored reference questions, + each tagged with the `task:` it covers and its `select:` count. Use them as **style, difficulty, + and format anchors only**. Prefer the anchor tagged with your task statement if one exists (six + task statements have none — D1.7, D2.5, D3.5, D4.2, D4.4, D5.4 — so fall back to the nearest + in-domain anchor). Never reproduce one verbatim; author fresh. **Authoring process:** 1. Locate the requested task statement in the blueprint and extract the concept it tests plus - its flagged anti-patterns. + its domain's flagged anti-patterns. 2. Frame ONE realistic production scenario in the requested scenario context. Keep the stem to a few sentences — a concrete situation a working architect would hit. -3. Write exactly four single-select options (A–D): one clearly-correct answer and three - plausible distractors **built from the syllabus anti-patterns** for that task statement - (the wrong answers a partial-knowledge candidate would actually pick). -4. Stay strictly in-scope. Never test an out-of-scope topic. Test judgment, not trivia or +3. Write exactly four options (A–D). For `single`, one is clearly correct. For `multi`, exactly two + are correct, each **independently** correct and expressing a **distinct** idea — two options + restating the same point is the mark of a badly-built multiple-response item. Every remaining + option is a plausible distractor **built from the syllabus anti-patterns** for that task + statement (the wrong answers a partial-knowledge candidate would actually pick). +4. For `multi`, end the stem with `**Select TWO.**` on the same line. The real exam always states + the count, and a candidate who does not see it will answer the wrong question. +5. Stay strictly in-scope. Never test an out-of-scope topic. Test judgment, not trivia or API-parameter memorization. -5. **Calibrate difficulty** by how close the distractors sit to the correct answer and how fine +6. **Calibrate difficulty** by how close the distractors sit to the correct answer and how fine the tradeoff is — NOT by adding obscure facts: - - `easy` — one obviously-right option; distractors are clearly weaker. + - `easy` — the right option is obvious; distractors are clearly weaker. - `medium` — distractors are reasonable-sounding; the learner must apply the concept. - - `hard` — all four are defensible on a quick read; only the correct one survives the + - `hard` — all four are defensible on a quick read; only the correct one(s) survive the tradeoff the task statement turns on. -6. Shuffle the correct option to a varied A–D position (don't default to A). +7. Shuffle correct options to varied A–D positions (don't default to A, or to AB for `multi`). **Output format — return EXACTLY this, nothing before or after:** ``` QUESTION - + A)