Assessment rubric · v1.0

CloudCaive Assessment Rubric v1.0

This is the rubric CloudCaive assessment reports cite. It states what an assessed session is, what is sampled, how the founder reviews evidence, the three possible outcomes for each sampled objective, and the limits of what any report may claim.

Version 1.0 (beta-calibrated)Effective 4 August 2026This version is immutable

1. What an assessed session is

An assessed session is a scheduled, founder-led session of browser-based Kubernetes incident simulations. During the founding beta there are two assessed sessions in each participant’s cycle:

  • The provisional gap diagnostic (Week 1 baseline): three compressed incident probes of 15–25 minutes each, or four of 15–20 minutes, within a hard total ceiling of 90 minutes.
  • The reassessment (end of cycle): new unseen incidents mapped to the same objectives, testing them under different conditions — within the same ceiling and graded against the same rubric version.

Incidents sit in a junior-to-mid difficulty band. Difficulty comes from ambiguity — several plausible explanations that evidence must separate — never from length. The simulation is not a real or production cluster.

2. What is sampled

Each session samples troubleshooting objectives: named, plain-language statements of what is being tested and why it matters operationally. Every assessed item maps to its objectives through a construct map fixed before the session; an outcome is only reported for an objective the session actually gave you an opportunity to demonstrate.

The objectives are drawn from realistic Kubernetes failure classes, including: workload failures (CrashLoopBackOff, out-of-memory), Service/DNS/ingress faults, configuration and Helm-release failures, RBAC and admission denials, and node or scheduling pressure.

The specific incidents, their benchmark answers and the reassessment’s form-comparison material are deliberately held back. Publishing them would let a session be rehearsed rather than sampled and would break the unseen reassessment.

3. What evidence is reviewed

Review considers the structured evidence of your investigation during the session:

  • which evidence you choose to inspect, and the order you inspect it in;
  • the hypotheses your choices support or rule out;
  • whether your conclusion follows from the evidence you gathered;
  • whether you verify the result after making a change;
  • your choices at the incident’s decision points: alert interpretation, evidence prioritisation, rollback-versus-repair, escalate-or-continue, and the written handoff.

These five decision points are the review dimensions of Rubric v1.0. For each sampled objective, the assessment plan pre-registers — before the session — the essential observable evidence required to judge that objective, expressed in terms of these decision points. What counts as sufficient evidence is defined per objective by that pre-registration; there is no universal observation count.

State computed in your browser is never itself the evidence of record; grading is performed offline by the founder against the versioned, held-back benchmark for the scenario.

4. How judgement is made

Judgement is criterion-referenced and qualitative: each sampled objective is judged against its own pre-registered essential evidence, under the session’s stated conditions. You are not ranked against other participants, and Rubric v1.0 defines no aggregate score and no numerical cut score.

  • Observable claims only. A report says what the session showed — for example, “could not isolate the failing Service selector within the scenario” — never a character judgement such as “not ready for on-call”.
  • Alternative solutions count. A technically valid alternative solution is assessed on the rubric objective, not on whether it matches the benchmark keystrokes.
  • What is never judged: accent, confidence, communication style, employer prestige, or memorised command syntax — unless an objective explicitly assesses it (the written handoff, for instance, assesses written communication).
  • Contamination is downgraded, never hidden. If founder assistance during a session contaminates an objective’s result, that objective is reported as insufficient evidence — never a hidden penalty and never a silent pass.

5. The role of founder review

During the founding beta the founder personally schedules every session, delivers each probe, grades the assessed work offline against this rubric’s versioned benchmark material, and writes every report. Assistance during sessions is standardised, and assessed work is reviewed, and its outcomes assigned, before any satisfaction feedback is read.

This process is designed to be bias-resistant, consistent and auditable — it is neither independent nor bias-free, and is not claimed to be. A founder-run process cannot be. Predefined criteria and recorded evidence reduce discretion; the controls that reduce and expose bias include standardised assistance, blind re-review spot-checks of a sample of reports, and a recorded assessment-outcome dispute log. During the founding beta there is no external validation and no independent reliability evidence, and none is implied.

6. The three outcomes

Each sampled objective receives exactly one of three outcomes. Demonstrated and not demonstrated may be assigned only when the session supplied a fair opportunity to observe all of that objective’s pre-registered essential evidence:

OutcomeWhat it means
DemonstratedSufficient uncontaminated evidence supports the sampled objective under the session’s stated conditions, including appropriate verification of the result. A technically valid alternative approach counts.
Not demonstratedA fair opportunity and sufficient evidence existed, but one or more essential requirements were absent, contradicted, unsafe or left unverified.
Insufficient evidenceThe evidence cannot honestly support either a positive or a negative inference.

A missing opportunity, contamination, excessive assistance, interruption or technically invalid evidence produces insufficient evidence for the affected objective whenever it prevents a supported positive or negative inference. CloudCaive does not turn silence, a missing opportunity or an ambiguous action into a confident outcome, and never reports an unsupported number.

“Not demonstrated” describes this sample only. It means the essential evidence was not met in this session — it does not mean you are generally incapable of the objective.

7. Sampling limits — what a report may not claim

A report is an honest account of sampled performance under stated conditions. It describes only what the sampled session demonstrated. It is not:

  • a certification, job guarantee or employer-recognised credential;
  • a claim that you are ready for every incident, role or environment;
  • a verdict on your competence, your job or your overall readiness;
  • a claim that any identified gap is permanently “closed” or a skill “mastered”.

One session samples a small number of objectives under specific conditions. It does not certify total professional capability, and a single demonstrated outcome does not establish broad competence.

8. How reassessment works

  • A new incident mapped to the same objective tests it under different conditions. The reassessment form is unseen and pre-committed — assigned before results are known, so it cannot be chosen to flatter the outcome.
  • CloudCaive does not claim the two forms are psychometrically equivalent: a documented form-comparison protocol has not yet been established. This is one reason every report is labelled beta-calibrated.
  • Both sessions in a cycle are judged against the same rubric version. Outcomes are compared objective by objective. If either session produces insufficient evidence, no directional change is claimed for that objective.
  • If comparability is lost — for example, a session had to be interrupted — the objectives are reported honestly and the comparison is marked as lost. A “before/after” is never manufactured.
  • The comparison shows how observed performance changed between two sampled sessions. It does not show general competence, workplace readiness, or that any skill is mastered.
  • Targeted practice is not itself evidence that a gap has resolved. Any reported change must be supported by evidence gathered during the new unseen reassessment incident.
  • An assessment-outcome dispute may be raised on any objective; disputes are re-checked against this rubric and recorded in the dispute log.

9. Version, effective date and change history

  • Version: 1.0 (beta-calibrated) — the initial published rubric for the founding beta.
  • Effective date: 4 August 2026 — the date this version first became publicly accessible in production. Reports may cite this version from that date.
  • Change history: v1.0 is the first version; there are no prior versions.
  • Immutability: after publication, no semantic change is made at this address. Any substantive change produces a new version at a new address, listed in the rubric version history. Purely presentational or accessibility repairs that leave the meaning unchanged may be applied, and each such repair is recorded in the change history. Superseded versions remain published at their original addresses.
  • Citation: every report records the rubric version and scenario version it was graded against — for example, “Assessed using CloudCaive Rubric v1.0”.

Reports produced during the founding beta are labelled beta-calibrated: the item bank and grading process are new, and their long-run reliability is still being established. This rubric makes no claim of predictive validity for job performance.