August 2026
Organisations have started giving AI systems credentials and the ability to act on them. Almost none can name the person who is answerable when one of them acts wrongly. That accountability is a role, and in most estates it is currently held by nobody.
One role decides what the system may do. This one is answerable when it does it.
This is not a staffing oversight. Each role below is doing its own job correctly, and none of them can hold this one without a conflict, which is why the gap stays open even in well-run organisations.
None of these is a document about a decision. Each one is the decision itself, in a form somebody outside the organisation can check.
The value of the role is easiest to see as a list of things that happen when it is vacant. Each of these gets attributed to something else at the time.
Any competent consultancy can produce a maturity matrix in which an organisation scores respectably. The instrument behind this role is built to make that specific outcome difficult, and three rules do the work.
The first is that the score is the floor. Eight dimensions are measured independently and the reported number is the lowest, because the weak dimension is the blast radius and averaging is how it disappears.
Identity and provenance.
Memory governance.
Capability and authorisation.
Auditability and receipts. This is the score. Not the average of the eight, which would read comfortably. The one that is weakest, because that is where an incident goes.
Substrate and model strategy.
Safety and failure modes.
Organisation and process.
Compliance and risk.
The second rule is that claims are not evidence. Every criterion carries an evidence grade, and the grade gates what it is allowed to prove. This is enforced in the scoring code rather than left to the assessor's discipline, because assessor discipline is exactly what erodes in week four of an engagement the client is paying for.
The gate. A criterion claimed at the enforced levels and supported only by C or D is recorded as not met, however credible the person saying it. This is the single most unpopular rule in the instrument and the one that makes the rest of it worth anything. A policy describing an enforcement is evidence that the policy exists.
The third rule is that levels are conjunctive. A level is reached only when every criterion at that level is met. There is no partial credit that promotes, because a control set with one hole is not eighty per cent of a control set; it is a control set with a hole, and the hole is what gets used.
The same eight dimensions, the same criteria and the same arithmetic apply to a bank and to a ten-person company. Scope, evidence burden, target level and price scale. The measurement does not. The moment the standard softens for a smaller client, the score stops meaning anything to the enterprise reviewer it was built for, and that reviewer is the entire point.
| Estate | The agent typically | Reaches | Who asks for the evidence |
|---|---|---|---|
| Financial services | Reads positions, drafts client communications | the customer record | The regulator, then the auditor |
| Healthcare | Summarises notes, routes referrals | the patient record | The privacy office, and eventually a patient |
| Software and SaaS | Writes code, opens pull requests, deploys | production | Every enterprise customer's security review |
| Legal practice | Drafts, researches, summarises discovery | privileged material | The client, the insurer, the bar |
| Industrial | Schedules, tunes, orders parts | the physical process | Safety, and the incident investigator |
| Small business | Runs the inbox, the calendar, the billing | the whole company | Nobody, until it matters |
The last row is the honest one. A small estate is not a lower standard, it is a smaller surface, and the whole assessment fits in a fortnight rather than a month because there is genuinely less of it. That is a real reason for a lower fee. Scoring generously is not.
Most organisations need the accountability, not a full-time hire to carry it. But a role sold as continuous accountability and backed by one person with a sleep schedule is a breach the provider commits on its own, without the client doing anything wrong, simply by that person being asleep or on another engagement when something happens.
So the conditions below are not fine print. They are what makes the arrangement real, and any provider offering this role without them is selling something they are not staffed to deliver.
"Accountable officer" invites a buyer to assume a pager. The gap between that assumption and the truth is where the unforced breach lives, so the hours are written down before signature and the absence of monitoring is stated in the same paragraph.
Every engagement requires a client-side stop mechanism, held by someone independent of the team that built the systems, tested during the engagement with the date recorded. An officer who becomes the single point of failure is the exact finding the instrument exists to raise.
Where no substitute exists, that is stated to the client before signature rather than discovered afterwards, and the client holds a right to terminate immediately if the named individual becomes unavailable beyond a stated period.
Days per month, divided by days per engagement. A provider who cannot tell you their ceiling has not done the division, and the number they eventually give you under pressure will be the one that fits the deal in front of them.
The category is filling with respectable-looking offerings. These five questions separate them, ordered by how hard they are to answer well without meaning it.
The rubric is at v1.0 and is deliberately frozen through the first engagements. Adding a criterion changes the denominator and therefore every historical decimal in that dimension, so defects are logged and a version ships on purpose rather than continuously and invisibly.
The full criterion set is provided to anyone being assessed under it, because a measurement the measured party cannot audit is not a measurement. Every criterion carries a verdict, an evidence grade and a citation, and the arithmetic over them is published rather than described.
The eight dimensions, the four levels, the evidence grades and the gate, and the floor rule with its arithmetic. What a level means and what it takes to reach one.
What an assessment involves, what the retainer covers and does not, the artifacts produced, and the terms that make an accountable officer arrangement honest rather than notional.