Saluca Labs
Saluca Labs

The Measurement

How the score is derived

Rubric v1.0 · August 2026


A governance score is worth exactly what its rules prevent. These are the rules, published before an assessment rather than explained after one, so that nothing about how a number was reached is hidden from the organisation being measured.

The number is not the finding. What it refuses to hide is.

The shape

Ninety-three criteria, eight dimensions, four levels

Each dimension is scored independently against the same four-level ladder. A criterion is an atomic, adjudicable statement with a verdict, an evidence grade and a citation. There is no dimension weighting, because a weighting is a judgement smuggled into arithmetic.

DimensionMeasuresThe question underneath
D1 Identity and provenanceHow agents are identified and how their actions are attributedCan you tell which agent did a thing, and prove it later?
D2 Memory governanceWhat agents retain, for how long, and under what controlWhat does it know, and who decided it should?
D3 Capability and authorisationWhat agents may do, and what enforces the boundaryIs the limit enforced, or merely described?
D4 Auditability and receiptsWhether actions and authorisation decisions are recorded and retainedCould you reconstruct an incident afterwards?
D5 Substrate and model strategyModel dependency, portability and change managementWhat happens when the provider changes something quietly?
D6 Safety and failure modesStop mechanisms, blast radius, degradation behaviourHow does it fail, and who can stop it?
D7 Organisation and processOwnership, approval, and the gate new agents pass throughWho said yes to the last one?
D8 Compliance and riskInventory, vendor terms, regulatory mappingDo you know what you are running and on whose terms?
D4 is the one that is usually weakest and the one that decides most scores. It is also the least visible before an incident, because an estate with no receipts behaves identically to an estate with good ones right up until somebody asks a question about last Tuesday.

The ladder

Four levels, and a level is reached only in full

The scale runs 0.00 to 4.00. The anchors are offset by one because absent is not the same as level zero: an estate nobody has looked at scores below one whose agents are at least dispatched in a known way.

AnchorLevelWhat it means
0.00Absent or unknownNothing in place, or nothing anyone can describe
1.00L0 stateless dispatchAgents run. Identity, memory and authority are incidental to how they were built
2.00L1 governed identityAgents are identified, inventoried and owned. Someone approved them
3.00L2 sovereign and enforcedBoundaries are enforced by architecture rather than convention, and the enforcement is observable
4.00L3 measured governed intellectThe governance is itself instrumented, and drift is detected rather than discovered

Levels are conjunctive. A level is achieved only when every one of its criteria is met. There is no promotion on a majority, because a control set with one hole is not most of a control set. It is a control set with a hole, and the hole is the part that gets used.

The decimal is partial credit toward the next level, derived as criteria met divided by criteria total at that level. A dimension anchored at L0 with three of five L1 criteria met scores 1.60.

Evidence

Claims are not evidence, and the rule is enforced in code

Every criterion carries a grade describing how its verdict was established. The grade gates what the criterion is allowed to prove, and the gate is applied by the scoring tool rather than by the assessor, because assessor discipline is exactly what erodes in week four of an engagement the client is paying for.

A. Observed directlyconfiguration, logs, a test run in front of you
B. Produced from the systeman export, a report the system generated itself
C. Documenteda policy, a diagram, a runbook
D. Assertedsomeone credible told you it works

The gate. A criterion claimed at L2 or above and supported only by grade C or D is recorded NOT MET, however senior the person asserting it and however well written the policy is. A document describing an enforcement is evidence that the document exists.

The gate does not apply below L2. An organisation can reach L1 on documentation, because L1 is largely about whether things are known, owned and written down. L2 is a claim about what the system actually does under adversarial conditions, and that claim needs to have been watched.

The floor

The overall score is the lowest dimension, not the average

This is the single rule that most distinguishes the instrument, and it is the one clients most often ask to have changed.

An estate with seven well-governed dimensions and one unguarded one does not have a good average. It has an unguarded dimension, and that is where an incident will go. Averaging converts the most important fact about an estate into a rounding contribution.

Worked. Dimensions scoring 1.80, 1.60, 1.50, 1.25, 2.50, 1.50, 1.67 and 1.75 produce an average of 1.70 and a floor of 1.25. The average describes an estate making reasonable progress. The floor describes one whose receipts would not survive an incident review. Only one of those two is useful.

A consequence worth stating plainly: remediation that does not touch the floor dimension does not move the score at all. That is not a defect in the arithmetic. It is the arithmetic telling you where the work is.

Two rules that surprise people

Both were learned by getting them wrong

Sequencing

Only criteria immediately above the anchor change the score

Partial credit counts at the level directly above where a dimension sits. Closing an L2 criterion on a dimension anchored at L0 scores nothing at all, however good the work is. Any remediation plan that ignores this is sequencing by intuition, and it will produce a quarter of genuine improvement with no movement to show for it.

Exclusion

Not applicable removes a criterion from both sides of the fraction

An NA is legitimate and sometimes necessary, and it is recorded with its justification in the report. An unjustified NA reads as a score inflated by exclusion, and a hostile reviewer will read it exactly that way. The safest assumption is that every NA will be challenged by someone who did not attend the engagement.

Versioning

Criterion identifiers are permanent, and the rubric is frozen on purpose

Criterion IDs are the join key across the scorecard, the gap register, the roadmap, the quarterly re-score and the framework crosswalk. They are never renumbered. A retired criterion is marked deprecated and its identifier is burned rather than reused.

The rubric is frozen at v1.0 through the first engagements. Adding a criterion changes a denominator and therefore every historical decimal in that dimension, which would silently rewrite the past. Defects are logged and a version ships deliberately, rather than the instrument drifting underneath the scores it has already produced.

A crosswalk to ISO/IEC 42001 or the NIST AI RMF may appear in a report. It is a mapping, not an attestation. It does not evidence conformity with either standard and does not substitute for certification by anybody entitled to grant it.

What the number is not

The limits are part of the measurement

The last one is disclosed here rather than in an appendix because a measurement practice that hides its own missing control has already failed the thing it sells. It is restored the day a second assessor exists and not before.

The criteria

Provided in full to anyone measured under them

This page describes the machinery. The complete criterion set, with the adjudication guidance for each, is supplied to any organisation being assessed, because a score the measured party cannot audit is an opinion with a number attached to it.

The role →

What an accountable AI security officer is, what the role produces, and what breaks in an estate that does not have one.

The engagement →

What an assessment involves, what it produces, what it costs, and the limits stated in writing before signature.