Evidence
Evidence, not content.
Nothing here is demonstrated yet. This page says how it will be, and publishes the empty tables it will fill.
Research and validation programme
Status: draft v0.1, GOVBRM original, not yet validated.
Purpose
This document records how GOVBRM was developed, what it rests on, what it claims, and how those claims will be tested. It is the method's own account of its evidence base. It is written so that a reader can tell, line by line, which parts are borrowed from established practice, which are adapted, which are original, and which are hypotheses that have not yet been tested.
The short version: GOVBRM was built from practice, published in the open, and has not been independently validated. Every outcome claim in this document is marked "Not yet demonstrated." until a study described in this programme has been run and reported.
What GOVBRM is
GOVBRM is the operating layer between AI demand and AI delivery. It shapes requests for AI before they become projects, routes them to the governance that already exists, and forces a value realisation conversation after delivery. It does not replace business cases, appraisal, procurement, security assurance, privacy and equality assessments, transparency records, service assessments or AI management standards. It orchestrates them.
The method has six stages (Discover, Assess, Prioritise, Design, Adopt, Realise), six gates (Request, Shape, Rank, Commit, Build, Review), twenty canvases in the AI Demand Toolkit, three lanes (fast, standard, strategic), an autonomy ladder and a Value Realisation Review that defaults to month nine.
How it was developed
Literature and frameworks reviewed
The method was developed by reading and working alongside the following categories of material. Nothing in these categories is reproduced in GOVBRM beyond names and published numbering; the ideas are used and the expression is GOVBRM's own.
| Category | What was taken from it | Label |
|---|---|---|
| Business Relationship Management body of knowledge | Relationship as a strategic capability, demand shaping, value planning, value realisation and value optimisation | Established (ideas), Adapted (expression) |
| Public sector appraisal and business case guidance (for example the Green Book and its equivalents) | Value as a range with conditions, optimism bias, the separation of options from the preferred option | Adapted |
| Digital service standards (for example the Service Standard and the Technology Code of Practice and their equivalents) | User need before technology; the request is not the need; assess at points, not once | Established |
| AI governance guidance and standards (for example the AI Playbook, the Data Ethics Framework, the Algorithmic Transparency Recording Standard, NIST AI RMF, ISO/IEC 42001) | The list of assessments an AI request must be routed to; the vocabulary of risk, transparency and accountability | Established (as routing targets) |
| IT governance and service management standards (ISO/IEC 38500, ISO/IEC 27001, ITIL) | Separation of governance from management; request intake as a managed process | Established |
| Benefits management and change management practice | Deployment is not use; adoption measured from the work; benefits owned by a named person | Established |
| Public guidance on automated decision-making and human oversight | The idea that autonomy is a design choice, not a given | Adapted |
The BRM relationship
GOVBRM's lineage is Business Relationship Management. BRM's contribution is the idea that the person who receives demand should shape it rather than take orders, that value is planned before it is delivered and realised after, and that a relationship with the business is a capability in its own right. GOVBRM takes those ideas and applies them to one problem: requests for AI in public and regulated organisations. It is not a BRM certification, does not reproduce BRM assessment material, and is not affiliated with any BRM body.
Practitioner experience
The method was written by one practitioner who has sat on the receiving end of AI requests in public and regulated organisations. That experience is the source of the pattern GOVBRM addresses: requests arrive as solutions, ownership is unclear, value is asserted rather than measured, governance is either avoided or applied in full regardless of scale, and nobody returns to ask whether the thing was worth it.
That experience is described here generically and deliberately. No organisation, department, supplier or product is named, no case from that experience is presented as evidence, and nothing in this document should be read as describing any employer's position or practice.
Written from practice is a claim about provenance, not about validation. The two are kept apart throughout.
Design rationale
Each of the main mechanisms was designed to answer a specific failure seen in practice.
| Mechanism | Failure it answers | Label |
|---|---|---|
| Six gates as a model of demand, not delivery | Delivery methods start once a project exists; the damage is done before that | GOVBRM original |
| Request gate separates the ask from the need | Requests arrive as solutions | Established idea, GOVBRM gate design |
| Explicit business owner at Shape | Nobody owns the outcome after go-live | Established idea, GOVBRM placement |
| Value as a range with named conditions | Single-point benefit claims are wrong and unowned | Adapted |
| Three lanes, burden in proportion to blast radius | Governance either avoided or applied in full | GOVBRM original; the proportionality principle is Established |
| Autonomy ladder, lowest rung that solves the task | Autonomy chosen by default, not by need | GOVBRM original as a shaping device; Hypothesis as a proxy for risk |
| Four agent questions | Agent requests arrive without a statement of what the agent may do alone | GOVBRM original |
| Value Realisation Review, default month nine | Nobody returns to check | Hypothesis (the timing); Established (that a review should exist) |
| Twenty canvases, ninety-minute sessions | Governance asks for documents, not conversations | GOVBRM original |
Current evidence
Every outcome claim GOVBRM might make is listed here with its current evidence status. This table is the honest baseline against which future evidence will be added.
| Claim | Evidence | Status |
|---|---|---|
| GOVBRM reduces the share of AI requests that reach build without a stated need | Not yet demonstrated. | Hypothesis |
| GOVBRM increases the share of AI requests with an explicit business owner | Not yet demonstrated. | Hypothesis |
| GOVBRM increases the share of AI requests with a measurable value hypothesis | Not yet demonstrated. | Hypothesis |
| GOVBRM routes requests to the correct existing governance more often than the organisation's current practice | Not yet demonstrated. | Hypothesis |
| Trained practitioners score and route the same request consistently | Not yet demonstrated. | Hypothesis |
| GOVBRM reduces time to triage and time to decision | Not yet demonstrated. | Hypothesis |
| GOVBRM reduces governance burden for low blast radius requests without increasing missed escalations | Not yet demonstrated. | Hypothesis |
| Interventions shaped through GOVBRM are adopted and realise benefits at a higher rate | Not yet demonstrated. | Hypothesis |
| The autonomy ladder is a useful shaping question | Not yet demonstrated. | Hypothesis |
| The month-nine default is the right default review point | Not yet demonstrated. | Hypothesis |
| The scoring weights and thresholds in the toolkit are calibrated | Not yet demonstrated. They are version 1.0 initial parameters. | Hypothesis |
Limitations
The following limitations are known now and are not resolved by this programme; they bound what any future evidence can show.
- The method was designed by one person and has no external advisory group yet.
- The risk canvas is a triage device, not a risk methodology. It routes to the assessments that weigh likelihood, impact, velocity and detectability; it does not do that work itself.
- The value range is not an appraisal and the business case canvas does not replace a full business case.
- Public value, who bears the benefit and the cost and who has no seat at the table, is asked but not modelled.
- Sustainability and enterprise architecture are routed to, not covered.
- The method was developed in one jurisdiction's practice and framed for several. Mappings to other jurisdictions' guidance are a starting point, not a certified equivalence.
- Any study run by the method's author carries an obvious conflict of interest. The programme below says how that is handled.
Validation status
As of 13 September 2026:
- No pilot organisation has run the method.
- No inter-rater reliability study has been run.
- No controlled comparison has been made.
- No outcome statistics exist.
- No customers, testimonials, endorsements or accreditations exist.
- The verification registry is empty.
- There is no Method Council or advisory group.
- Accessibility is built to WCAG 2.2 AA and has not been audited.
- No legal review has taken place.
Planned research
The programme runs in the following order. Each stage is described in its own document in this folder.
| Stage | Document | What it establishes | Depends on |
|---|---|---|---|
| 1. Inter-rater reliability study | Inter-rater reliability study | Whether trained practitioners apply the method consistently | A trained cohort of practitioners and a case set |
| 2. Pilot study | Pilot study methodology | Whether the method changes what happens at the front door in real organisations | Stage 1 reaching acceptable agreement; 3 to 10 consenting organisations |
| 3. Evidence framework | Evidence framework | How anonymised evidence from stages 1 and 2 is graded and published | Stages 1 and 2 producing data |
| 4. Research backlog | Research backlog and hypotheses | The questions and hypotheses that stages 1 to 3 test, with their status | Nothing; maintained throughout |
| Supporting | Case study template | The structure every published case follows and the labelling rules | Nothing; used from stage 2 |
Independence
Because the method's author has an interest in the result, the programme adopts the following rules.
- Study designs are published before data is collected, so the analysis cannot be changed to fit the result.
- Raw anonymised data behind any published metric is held and can be shared with a reviewer under agreement.
- Where an external reviewer, academic partner or advisory group exists, they sign off the analysis before publication. Until one exists, that absence is stated on every published result.
- Negative and null results are published in the same place and format as positive ones.
How the method changes in response
Findings feed method revision through the version history on the provenance page. A revision states which finding prompted it, what changed, and which hypothesis it affects. Calibration parameters (weights, thresholds, review timing) carry a version number and a note of the evidence that set them.
Evidence framework
Status: draft v0.1, GOVBRM original, not yet validated.. No evidence has been collected.
Purpose
This framework sets the structure for publishing evidence about GOVBRM. It exists so that the first published number, whenever it arrives, sits in a table whose columns were designed before the number existed, is graded on a scale published before the number existed, and cannot be confused with an invented one.
Principles
- Every published claim carries an evidence grade.
- Every metric is published with its baseline, its denominator and the study it came from.
- Blank cells are published as blank. "Not yet demonstrated." is a legitimate entry and is the default.
- No case study, quotation, figure or testimonial is published unless it is real, consented and labelled under the scheme in Case study template.
- Negative and null results are published in the same table as positive ones.
- The author's conflict of interest is stated wherever evidence is published.
Metric table
This is the table that will carry GOVBRM's outcome evidence. It is published now, empty, so that readers can see what will be measured and what has not been.
| Metric | Definition | Baseline | GOVBRM result | Evidence | Grade |
|---|---|---|---|---|---|
| Time to triage | Working days from request logged to Request gate decision | Not yet demonstrated. | |||
| Time to decision | Working days from request logged to Commit gate decision | Not yet demonstrated. | |||
| Share stopped before build | Requests stopped or redirected before Build, as a share of requests logged | Not yet demonstrated. | |||
| Share with explicit business owner | Requests with a named outcome owner at or before Shape | Not yet demonstrated. | |||
| Share with measurable value | Requests reaching Commit with a measure, baseline and range | Not yet demonstrated. | |||
| Share routed to correct governance | Requests whose routed assessments match an independent reviewer's judgement | Not yet demonstrated. | |||
| Assessor agreement | Agreement among trained practitioners on the five judgements | Not yet demonstrated. | |||
| Governance burden | Hours from Request to Commit, by lane | Not yet demonstrated. | |||
| Stakeholder satisfaction | Requester and owner rating on the fixed scale | Not yet demonstrated. | |||
| Adoption | Evidence of use from the work at Review | Not yet demonstrated. | |||
| Realised benefits | Benefits observed at Review against the value hypothesis | Not yet demonstrated. | |||
| False escalation | Requests routed to a heavier lane than an independent reviewer judges needed | Not yet demonstrated. | |||
| Missed escalation | Requests routed to a lighter lane than an independent reviewer judges needed | Not yet demonstrated. |
Column rules:
- Baseline: the organisation's measure before GOVBRM, or "not recorded" if the organisation could not reconstruct it. Never estimated.
- GOVBRM result: the measure during or after GOVBRM, with the number of requests it rests on.
- Evidence: the study identifier and wave, linking to the published design and anonymised data description.
- Grade: from the scale below.
Evidence grading scale
| Grade | Meaning | What it takes |
|---|---|---|
| E0 | Not yet demonstrated | No data. The default for every claim until a study reports. |
| E1 | Practitioner observation | The author's or a practitioner's account from practice, described generically, with no data. This is provenance, not evidence, and is never used to support an outcome claim. |
| E2 | Illustrative | A constructed case or worked example showing how the method is intended to work. Never presented as a result. |
| E3 | Single organisation, uncontrolled | One consenting organisation's before-and-after measures from a published pilot design. Indicative only. |
| E4 | Multiple organisations, uncontrolled | Three or more organisations' measures from the same published design, with the range across organisations shown. |
| E5 | Independently reviewed | E3 or E4 evidence where the analysis was checked and signed off by a reviewer with no interest in the result, named or described. |
| E6 | Reproduced | A finding at E4 or E5 reproduced by a separate study with different organisations or raters. |
| E7 | Controlled comparison | A study with a comparison condition and a published design, independently reviewed. Not planned in the current programme. |
The scale is ordinal. A higher grade does not make a finding larger, only better supported. A finding may move down the scale if a later study fails to reproduce it, and that movement is recorded.
Claims of the form "GOVBRM reduces", "GOVBRM improves" or "organisations using GOVBRM see" require grade E4 or above. Below E4 the wording is "in one pilot organisation" or "not yet demonstrated".
Rules against invented evidence
- No case study is published unless it is labelled under Case study template as Illustrative, Anonymised real, or Published with consent. Illustrative cases are never placed in the metric table, never quoted as results and never given figures that could be mistaken for measurements.
- No number is published in the metric table that did not come from an instrument in a published study design.
- No quotation is attributed to a requester, owner, sponsor or organisation unless it was said, consented to and checked for identifying detail.
- No testimonial, endorsement, accreditation or affiliation is stated unless it exists in writing.
- No aggregate figure ("organisations report", "practitioners find") is published unless the organisations and practitioners exist and the figure is in the table with its denominator.
- Worked examples in the course, the toolkit and the workbook are labelled "Illustrative" in the text where they appear.
- If any of the above is found to have been breached, the item is withdrawn, the withdrawal is recorded on the provenance page, and the reason is stated.
Anonymised data model for the annual report
The programme intends, once evidence exists, to publish an annual "State of AI Demand in Government" report drawing on anonymised data from pilot and later implementing organisations. The report does not exist and no year is committed. The data model is designed now so that data collected from the first pilot fits it.
What is held
| Entity | Fields | Identifying? |
|---|---|---|
| Organisation | Code; sector (broad band); jurisdiction (broad band); size band; date joined; date left; consent scope | Code only; mapping held separately |
| Practitioner | Code; training completed (yes or no); credential level | Code only |
| Request | Code; organisation code; date logged; lane; autonomy rung; risk triage outcome; gate decisions and dates; owner named (yes or no); value hypothesis present (yes or no); assessments routed to; hours by role; stopped before Build (yes or no) | No free text; request description held only inside the organisation |
| Intervention | Request code; date adopted; review date; adoption evidence type; realised benefits as range and conditions, in categories not amounts | No free text beyond category |
| Survey response | Request code; role (requester or owner); scale items; free text held separately and paraphrased before any use | Free text separated |
| Review judgement | Request code; reviewer code; routing, lane and assessment judgements | Code only |
What is never held in the model
- Organisation names, department names, service names, supplier names, product names, place names
- Personal names, job titles specific enough to identify a person, contact details
- Request descriptions in the organisation's own words
- Amounts of money; benefits are held as categories and ranges
- Anything that would let two fields be combined to identify an organisation
Consent
- The organisation's participation agreement states that its coded data may be included in the annual report in aggregate, and lists the fields above.
- Consent is per report cycle. An organisation may withdraw from a future cycle without affecting earlier published aggregates.
- Individuals consent to survey and interview use separately, per Pilot study methodology.
Privacy
- Aggregates are published only when the number of organisations behind them is large enough that no organisation can be inferred from them, and that minimum is stated in the report. Until the minimum is met, the report says so rather than publishing thin aggregates.
- Sector and jurisdiction bands are set wide enough that their combination does not single out an organisation.
- Data is held separately from the code mapping, with access limited to the study lead and any reviewer under agreement.
- Retention is fixed in the participation agreement; data is deleted at the end of it.
Governance
- The report's method section is published before its findings section each year, and states the grade of every finding.
- Where an advisory group exists, it reviews the report before publication. Until one exists, the report states that it was not externally reviewed.
- Organisations receive their own coded data back on request, and the aggregate before publication.
- Any organisation may challenge an aggregate it believes identifies it, and publication is held until the challenge is resolved.
- The data model, the fields, and every change to them are versioned on the provenance page.
Current status
The metric table is empty. The grading scale is version 1.0 and has been applied to nothing. The data model exists as a design only. No annual report is scheduled.
Research backlog and hypotheses
Status: draft v0.1, GOVBRM original, not yet validated.. Every hypothesis below is untested.
Purpose
This is the register of what GOVBRM claims and how each claim will be tested. It is maintained throughout the programme. When a study reports, the status column changes and the finding is linked. Until then, every status reads "untested".
Research questions
| ID | Question | Why it matters | Studied in |
|---|---|---|---|
| RQ1 | Premature delivery. How many AI requests in public and regulated organisations reach build without a stated user need, a named owner and a value hypothesis, and does GOVBRM change that share? | The failure GOVBRM was designed to address | Pilot study |
| RQ2 | Need identification. When a request is taken through the Request and Shape gates, how often does the underlying need differ from what was asked for, and what happens to those requests? | Tests whether the request-is-not-the-need mechanism does work | Pilot study |
| RQ3 | Explicit ownership. Does requiring a named business owner at Shape change who owns outcomes after delivery, and does ownership persist to the Review gate? | Ownership is the precondition for value realisation | Pilot study |
| RQ4 | Measurable value. Does the value canvas produce value hypotheses that can be measured at Review, and how often are they? | A value range nobody can measure is a formality | Pilot study |
| RQ5 | Governance routing. Does GOVBRM route requests to the existing assessments an independent reviewer judges required, more often than the organisation's prior practice? | GOVBRM's claim to orchestrate, not replace, governance | Pilot study with independent review |
| RQ6 | Inter-rater agreement. Do trained practitioners score, route and classify the same case consistently? | If not, no other result is interpretable | Inter-rater reliability study |
| RQ7 | Adoption. Are interventions shaped through GOVBRM used in the work after deployment, as measured at Review? | Deployment is not use | Pilot study, Review follow-up |
| RQ8 | Realised benefits. Do interventions shaped through GOVBRM realise benefits within the value range stated at Commit? | The point of the method | Pilot study, Review follow-up |
| RQ9 | Governance waste. How much governance effort is spent on requests that are stopped, and does GOVBRM reduce effort on requests that should stop early? | Burden is a real cost to organisations | Pilot study, effort sheet |
| RQ10 | Minimum burden per blast radius. What is the least governance burden per lane that does not increase missed escalations? | The proportionality principle needs a calibrated answer | Pilot study, false and missed escalation measures; later calibration studies |
Hypotheses
Each hypothesis is stated so that it can be false. Each names its measure, the study that tests it, and its status. Where the hypothesis is directional, the direction is stated; no size of effect is claimed.
| ID | Hypothesis | Label | Measure | Tested by | Status |
|---|---|---|---|---|---|
| H1 | Requests taken through the Request and Shape gates are more often stopped or redirected before Build than requests under the organisation's prior practice | Hypothesis | Share stopped before build, baseline against pilot | Pilot study (RQ1, RQ2) | Untested |
| H2 | Requests taken through the Shape gate more often have a named business owner who is still the owner at Review than requests under prior practice | Hypothesis | Share with explicit business owner at Shape; owner unchanged at Review | Pilot study (RQ3) | Untested |
| H3 | Requests reaching Commit through GOVBRM more often carry a value hypothesis with a measure, a baseline and a range, and that hypothesis is more often measured at Review | Hypothesis | Share with measurable value at Commit; share measured at Review | Pilot study (RQ4, RQ8) | Untested |
| H4 | Requests routed at Shape match an independent reviewer's judgement of required assessments more often than requests under prior practice, without an increase in missed escalations | Hypothesis | Share routed to correct governance; missed escalation | Pilot study with independent review (RQ5, RQ10) | Untested |
| H5 | Trained practitioners, scoring the same case set independently, agree on lane, autonomy rung and risk triage outcome at a level the method will define after the first study | Hypothesis | Assessor agreement per judgement type | Inter-rater reliability study (RQ6) | Untested |
| H6 | Governance burden in hours from Request to Commit is lower for fast-lane requests than for standard and strategic ones, and the reduction does not come with more missed escalations | Hypothesis | Governance burden by lane; missed escalation by lane | Pilot study (RQ9, RQ10) | Untested |
Hypotheses about the method's own parts
These are hypotheses about design choices rather than outcomes. They are tracked here because a negative result changes the method.
| ID | Hypothesis | Label | Measure | Tested by | Status |
|---|---|---|---|---|---|
| H7 | The autonomy ladder, asked as "the lowest rung that solves the task", leads practitioners and requesters to choose a lower rung than the request arrived with, in a meaningful share of cases | Hypothesis | Rung requested against rung chosen at Shape | Pilot study | Untested |
| H8 | Month nine is a reasonable default for the Value Realisation Review across the interventions seen in pilots, in that most value curves neither peak much earlier nor much later | Hypothesis | Review date set against month nine, with reason for deviation | Pilot study, Review record | Untested |
| H9 | The version 1.0 scoring weights and thresholds produce lane assignments that independent reviewers agree with | Hypothesis | Share routed to correct governance; false and missed escalation | Pilot study; calibration revision | Untested |
| H10 | Ninety-minute canvas sessions are long enough to complete the canvases a lane requires without a second session, for most requests | Hypothesis | Sessions per canvas from the effort sheet | Pilot study | Untested |
Rules for this register
- A status changes only when a study with a published design reports. The permitted statuses are: untested, in progress, supported (with grade from Evidence framework), not supported, inconclusive, withdrawn.
- A hypothesis that is not supported is not deleted. It stays with its result and a note of what changed in the method in response.
- New hypotheses are added with the next number and never reuse a withdrawn one.
- No hypothesis is reworded after its study starts. If the wording was wrong, a new hypothesis is added and the old one is marked withdrawn with the reason.
- The register is versioned with the provenance page.
Current status
All ten research questions are open. All ten hypotheses are untested. No study has started.
Pilot study methodology
Status: draft v0.1, GOVBRM original, not yet validated.. No pilot has run.
Purpose
The pilot study asks one question: when an organisation runs its AI demand through GOVBRM, does what happens at the front door change, and in which direction? It is the first place the method's outcome hypotheses meet real requests, real owners and real governance.
The pilot is an observational study with a before-and-after comparison within each organisation. It is not a controlled trial. Its findings will be indicative, not conclusive, and will be published as such.
Participants
Number of organisations
The pilot runs with 3 to 10 organisations. Fewer than three gives no way to tell an organisation effect from a method effect. More than ten cannot be supported by one person while keeping data quality. The final number depends on who consents.
Selection
Organisations are sought that:
- are public or regulated, in any of the jurisdictions the method is framed for
- receive AI requests through some identifiable route, however informal
- have at least one person willing to act as the GOVBRM practitioner for the pilot period and to complete the practitioner training first
- have a sponsor at a level that can grant access to the request record and to the decision-makers
- will sign a participation agreement covering consent, data handling and publication
Selection will not be random. Organisations that volunteer are likely to be those already inclined to shape demand. This is a known bias and is stated in every published result.
What the pilot asks of an organisation
| Ask | Detail |
|---|---|
| Practitioner time | One or two people trained on the method, then running the gates for the pilot period |
| Sponsor time | A short briefing at start, a review at the midpoint, a debrief at the end |
| Access | The record of AI requests received during the baseline and pilot periods, and the decisions taken on them |
| Data | Completion of the instruments below for each request that enters the front door |
| Consent | A signed participation agreement, and individual consent from anyone interviewed |
What a pilot collects
A pilot collects data on every AI request that enters the organisation's front door during two periods: a baseline period before GOVBRM is applied and a pilot period during which it is.
Baseline period
For the baseline, the organisation reconstructs, as far as its records allow, the AI requests received in a defined period before the pilot. For each request the same measures are recorded where they can be, and marked "not recorded" where they cannot. Missing baseline data is reported as missing, never estimated.
Pilot period
During the pilot, each request is taken through the gates and the canvases the lane requires. The practitioner records the measures below at the gate where each is decided.
Measures
Each measure is defined so that two organisations record it the same way. The definitions below are version 1.0 and will be revised if pilots show they are ambiguous.
| Measure | Definition | Recorded at | Unit |
|---|---|---|---|
| Time to triage | Working days from the request being logged to the Request gate decision (accept for shaping, redirect, decline) | Request | days |
| Time to decision | Working days from the request being logged to the Commit gate decision (proceed, hold, stop) | Commit | days |
| Share stopped before build | Requests stopped or redirected at any gate before Build, as a share of all requests logged | Build | share, with denominator |
| Share with explicit business owner | Requests where a named person accepted ownership of the outcome at or before Shape | Shape | share, with denominator |
| Share with measurable value | Requests that reached Commit with a value hypothesis stating a measure, a baseline and a range | Commit | share, with denominator |
| Share routed to correct governance | Requests where the assessments routed to at Shape match those an independent reviewer judges required | Shape, reviewed later | share, with denominator |
| Assessor agreement | Agreement between the pilot practitioner and an independent second scorer on a sample of requests (see Inter-rater reliability study) | Sample | agreement statistic, to be chosen |
| Governance burden | Practitioner and requester hours spent from Request to Commit, by lane | Commit | hours |
| Stakeholder satisfaction | Requester and owner rating of the process on a short fixed scale, plus free text | Commit and Review | scale point |
| Adoption | For interventions that reach Build, evidence of use from the work (not deployment) at the Review gate | Review | defined per intervention |
| Realised benefits | For interventions that reach Review, benefits observed against the value hypothesis, as a range with conditions | Review | defined per intervention |
| False escalation | Requests routed to the standard or strategic lane that an independent reviewer judges belonged in a lighter lane | Shape, reviewed later | share, with denominator |
| Missed escalation | Requests routed to the fast lane that an independent reviewer judges belonged in a heavier lane | Shape, reviewed later | share, with denominator |
Counts are reported as counts. Shares are computed and reported only when the denominator is large enough to be meaningful, and the denominator is always published alongside them.
Data collection instruments
| Instrument | Purpose | Completed by | When |
|---|---|---|---|
| Request log | One row per request: identifier, date logged, lane, gate decisions and dates, owner named (yes or no), value hypothesis present (yes or no), assessments routed to | Practitioner | Every gate |
| Effort sheet | Hours by role for each request, Request to Commit | Practitioner and requester | Commit |
| Canvas pack | The completed canvases for the request, with identifying detail removed before leaving the organisation | Practitioner | Each gate |
| Requester survey | Fixed-scale satisfaction items and free text | Requester | Commit |
| Owner survey | Fixed-scale satisfaction items and free text | Business owner | Commit and Review |
| Independent review sheet | Second scorer's judgement of routing, lane and assessments for a sample of requests | Independent reviewer | After Shape |
| Review record | Adoption evidence and realised benefits at the Value Realisation Review | Practitioner and owner | Review |
| Sponsor interview | Semi-structured interview on what changed and what did not | Sponsor | Midpoint and end |
The instruments are built from the toolkit's canvases so that the pilot adds recording, not new work. Instrument templates are a deliverable of this study and do not yet exist.
Consent
- Each organisation signs a participation agreement before any data is collected. The agreement states what is collected, who holds it, how it is anonymised, what may be published and the organisation's right to withdraw.
- Each person who completes a survey or interview gives individual consent. They may decline any item and may withdraw their contribution until the point of anonymisation.
- Requesters and owners are told that the pilot is a study of the method, not of them.
- Consent covers publication only in the anonymised forms described in Evidence framework. Publication in any identifying form requires separate written consent.
Anonymisation
- Each organisation receives a code. The mapping from code to organisation is held by the study lead only, separately from the data.
- Each request receives a code. Request descriptions in the canvas pack are rewritten to remove service names, supplier names, product names, place names and any detail that would identify the organisation or a person.
- Free text from surveys and interviews is paraphrased, not quoted, unless the contributor has consented to quotation and the quotation has been checked for identifying detail.
- Sector and jurisdiction are recorded at a level broad enough that no organisation can be identified from them combined.
- Data that cannot be anonymised to this standard is not published and is deleted at the end of the retention period stated in the agreement.
Analysis plan
The analysis plan is fixed before the first pilot starts and published with this document.
- For each organisation, each measure is reported for the baseline period and the pilot period side by side, with the number of requests in each.
- Direction of change is reported per organisation. No pooled figure is reported across organisations unless the definitions were applied identically and the number of organisations supports it, and then only with the range across organisations shown.
- Governance burden is reported by lane, so that a reduction for fast-lane requests cannot hide an increase for strategic ones.
- False and missed escalations are reported together. A reduction in burden that comes with more missed escalations is a negative result and is reported as one.
- Adoption and realised benefits are reported only for interventions that have reached their Review gate within the study period. Interventions that have not are listed as "not yet at Review", not omitted.
- Assessor agreement on the sample is reported using the statistic chosen in the inter-rater study, with its interpretation stated.
- Qualitative findings from interviews are coded by two people where possible and reported as themes with the number of organisations in which each appeared.
- Every finding is graded using the scale in Evidence framework.
- Null and negative results are reported in full.
Duration
- Baseline reconstruction: 4 to 8 weeks before the pilot period starts.
- Pilot period: 6 to 12 months, long enough for a reasonable number of requests to pass Commit and for some to reach Build.
- Review follow-up: to the Value Realisation Review of interventions built during the pilot, which by default falls at month nine after go-live, so up to 12 months after the pilot period ends.
- Total per organisation: 12 to 24 months from agreement to final report.
Organisations may join in waves. Results are published per wave and updated as follow-up completes.
Exit criteria
An organisation's pilot ends when any of the following applies:
- the pilot period has run its planned length and all instruments are complete for requests logged within it
- the organisation withdraws, in which case its data is used only if it consents at withdrawal, and its withdrawal and the reason (if given) are reported
- the study lead ends the pilot because data quality cannot be maintained, and reports why
The pilot study as a whole ends when:
- at least three organisations have completed the pilot period and the follow-up, or
- the study lead judges that no further organisations will complete and publishes what exists with that stated
What the pilot will not show
- It will not show that GOVBRM causes any change. Organisations that volunteer, practitioners who are trained and the attention that comes with a study are all confounders.
- It will not compare GOVBRM with another method.
- It will not produce figures that transfer to organisations unlike the pilot ones.
These limits are printed with every result.
Inter-rater reliability study
Status: draft v0.1, GOVBRM original, not yet validated.. No study has run.
Purpose
A method that two trained people apply differently to the same request is not yet a method. Before GOVBRM is tested in pilot organisations, this study asks whether trained practitioners, given the same case, reach the same scores, the same lane, the same autonomy classification, the same risk classification and the same priority.
This is the first study in the programme because every later result depends on it. If practitioners disagree, pilot results measure the practitioners, not the method.
What is being tested
| Judgement | Where in the method | What agreement means |
|---|---|---|
| Scoring | The scored canvases in the toolkit (readiness, value, exposure and the others that carry a scale) | Two practitioners give the same or adjacent score on each scale |
| Route | The lane assigned at Shape (fast, standard, strategic) and the assessments routed to | Two practitioners assign the same lane and the same set of assessments |
| Autonomy classification | The rung of the autonomy ladder chosen as the lowest that solves the task, and the answers to the four agent questions | Two practitioners choose the same rung |
| Risk classification | The triage outcome of the risk canvas | Two practitioners reach the same triage outcome |
| Prioritisation | The rank order at the Rank gate | Two practitioners produce rank orders that agree on the top and bottom of the set |
Design
Case set
A set of written cases is prepared, each describing an AI request as it would arrive at the Request gate: the ask, the requester, the context, what is known and what is not. Cases are constructed, not drawn from any real organisation, and are labelled "Illustrative" under the scheme in Case study template.
The case set is designed to cover:
- every lane
- every rung of the autonomy ladder
- requests that arrive as solutions and requests that arrive as needs
- requests with and without an obvious owner
- requests where the value is easy to state and where it is not
- requests at the boundary between lanes, where disagreement is most likely and most informative
- requests where the correct outcome is to decline or redirect
Cases are reviewed by at least one person other than the author before use, to remove leading detail and to check that the intended classification is not signalled in the text.
Raters
Raters are practitioners who have completed the GOVBRM practitioner training. Raters must not have seen the case set before the study. Raters work independently, without conferring, and are told not to discuss cases until all scoring is complete.
The method's author does not act as a rater. The author's own classification of each case is recorded before the study, sealed, and used only to see whether raters agree with the author's intent as well as with one another. Agreement with the author is a secondary finding; agreement among raters is the primary one.
Procedure
- Each rater receives the case set, the toolkit, the practitioner guidance and a scoring sheet.
- Each rater scores every case, completing the canvases the lane requires and recording the five judgements above.
- Raters record time taken per case and any point at which they found the guidance ambiguous.
- Scoring sheets are returned before any discussion.
- Agreement is computed.
- Raters then meet to discuss every case on which they disagreed. The discussion is recorded and coded for the source of disagreement.
Sample size
The number of raters and the number of cases are not fixed in this draft, and no number is claimed to be sufficient. The reasoning that will set them is as follows.
- Agreement statistics become unstable with very few raters or very few cases. The study needs enough of each that a single unusual rater or a single ambiguous case does not dominate the result.
- Each judgement has a different number of possible outcomes (three lanes, five rungs, a small number of triage outcomes, a scored scale). The judgement with the most outcomes needs the most cases to estimate agreement on.
- Boundary cases must be numerous enough to say something about the boundaries, since those are where the method will be revised.
- Practical limits apply: each case takes a practitioner real time, and the trained cohort is small at the start.
The study will state, before it runs, how many raters and cases it has, which agreement statistic will be used for each judgement and why, and what precision is expected as a result. If the numbers available are too small to support a stable estimate, the study will say so and report the result as a pilot of the study design rather than as a reliability result.
What acceptable agreement would mean
No threshold is claimed yet. Agreement statistics carry conventional interpretation bands in the literature, and those bands are contested and depend on the statistic and the stakes. GOVBRM does not adopt any band in advance.
What acceptable agreement means for GOVBRM will be argued from consequence:
- For route and risk classification, the cost of a missed escalation is higher than the cost of a false one, so agreement on the boundary between lanes matters more than agreement in the middle.
- For autonomy classification, disagreement by one rung is less serious than disagreement by two or more.
- For scoring, agreement within one scale point may be acceptable where the score feeds a ranking, and not where it sets a threshold.
- For prioritisation, agreement on which requests are at the top and which at the bottom matters more than agreement on the exact order in the middle.
After the first study, GOVBRM will state a threshold per judgement, with the reasoning, and label it "Hypothesis" until a second study with different raters reproduces it.
How disagreements feed method revision
Every disagreement is treated as information about the method, not about the rater.
- Each disagreement is coded by source: ambiguous guidance, ambiguous case, missing rule, conflicting rules, judgement the method leaves open by design, or rater error.
- Disagreements coded as ambiguous guidance, missing rule or conflicting rules go to the method backlog as revision candidates, each tied to the canvas or gate concerned.
- Disagreements coded as judgement left open by design are reviewed to decide whether the method should close them or should state explicitly that they are open.
- Disagreements coded as ambiguous case lead to revision of the case set, not the method.
- Rater error is reported as a training finding.
- Revisions are made, versioned on the provenance page with the finding that prompted them, and the study is repeated with the revised method and a fresh cohort before any agreement claim is made.
Outputs
- A published study design, before data collection.
- An agreement result per judgement type, with the statistic, the number of raters and cases, and the interpretation.
- A coded list of disagreement sources.
- A list of method revisions made in response, with version numbers.
- A statement of whether the method is ready for the pilot study, and if not, what must change first.
Current status
Not yet demonstrated. No case set exists. No trained cohort exists. No agreement statistic has been chosen. Any figure quoted for GOVBRM inter-rater agreement is unfounded.
Case study template
Status: draft v0.1, GOVBRM original, not yet validated.. No case study exists.
Purpose
This template is the single structure every GOVBRM case follows, whether it is a constructed example in a course or a consented account from a real organisation. Using one structure for both makes the label on each case the only thing that distinguishes them, which is why the labelling rules below are strict.
Labelling scheme
Every case carries one of three labels in its first line, and the label is repeated wherever the case is referenced.
| Label | Meaning | May it carry figures? | May it appear in the metric table or as evidence? |
|---|---|---|---|
| Illustrative | Constructed to show how the method is intended to work. No real organisation, request or person behind it. | Only clearly hypothetical ones, marked as such in the text | Never |
| Anonymised real | A real request in a real organisation, taken from a study with a published design, with all identifying detail removed to the standard in Evidence framework, and with the organisation's consent to anonymised publication | Yes, from the study instruments, with the evidence grade | Yes, at its evidence grade |
| Published with consent | A real request in a real organisation, published with the organisation's separate written consent to be identified | Yes, from the study instruments, with the evidence grade | Yes, at its evidence grade |
Rules
- An invented case study is never presented as evidence. It is never given a metric-table row, never quoted in support of an outcome claim, and never described with words that imply it happened ("a department found", "one organisation saw").
- Illustrative cases use obviously generic descriptions ("a licensing service", "a case management team") and never a name that could be mistaken for a real organisation.
- Illustrative cases never carry a testimonial, a quotation from a named or role-described person, or a date that suggests an actual event.
- A case may move from illustrative to anonymised real only if it is replaced entirely by a real case from a study; an illustrative case is never "confirmed" by later data.
- Anonymised real cases are checked against the anonymisation standard by someone other than the author before publication.
- Published with consent cases carry the consent date and the scope of consent, and are withdrawn on request.
- Where the label is missing, the case is treated as illustrative and withdrawn until labelled.
Template
Label: [Illustrative | Anonymised real | Published with consent] Evidence grade: [from Evidence framework; E2 for illustrative] Study reference: [study and wave, or "none" for illustrative] Version: [date and version of the method the case ran under]
Organisation
Sector band, jurisdiction band, size band. For illustrative cases, a generic description only. For anonymised real cases, the organisation code and the bands. For published with consent cases, the organisation as consented.
Context
Why AI was on the table. What pressure, statute or policy, minister, board or executive interest, or citizen or customer need brought the request forward. What the organisation's route for AI requests looked like before this one.
Original demand
The request as it arrived, in the requester's own framing (paraphrased for real cases). Who asked, what for, and by when.
What the requester initially wanted
The solution the requester had in mind, including any product, autonomy level or delivery route already assumed. Named generically ("a document summarisation tool from the model provider").
Underlying need
What the Request and Shape gates found the need to be. Whether it matched the ask. What the need was in the words of the people doing the work and the citizens or customers affected.
Evidence
What was known about the need, the current process and the baseline, and what was not. What the Discover stage collected. What was assumed.
Value hypothesis
The value range stated at Commit: the measure, the baseline, the range, and the named conditions on which the range depends. Who owns it.
Risk
The risk triage outcome, the assessments routed to, and anything the triage flagged for those assessments to weigh.
Autonomy
The rung requested, the rung chosen as the lowest that solves the task, and the answers to the four agent questions where the request involved an agent.
Governance route
The lane, the gates passed, the existing governance the request was routed to (business case, appraisal, procurement, security assurance, privacy and equality assessments, transparency records, service assessments, AI management standards), and how each responded.
Decision
The Commit gate outcome: proceed, hold, stop or redirect. Who decided. On what basis.
Delivery
What was built or bought, by whom, over what period. Where the Build gate found deviation from the Commit decision and what was done.
Adoption
Evidence of use from the work, not deployment. Who used it, for what, how often, measured how. Where adoption fell short of the plan and why.
Benefits
What was realised against the value hypothesis, as a range with conditions, in categories rather than amounts unless consented. Where benefits fell short, exceeded or differed in kind from the hypothesis.
Value Realisation Review
When the review was held relative to adoption and why that date was chosen against the month-nine default. What it found. What went back to the front door.
Lessons
What the organisation and the practitioner would do differently. What the case suggests about the method, tied to a canvas, gate or parameter where possible and logged in the research backlog if it affects a hypothesis.
Metrics
For real cases only: the measures from Pilot study methodology that this request contributed to, with their values and the denominator they sit in. For illustrative cases this section reads "Illustrative case: no metrics."
Current status
No case study of any label has been written under this template. The Academy courses and the toolkit contain worked examples that must be checked against the labelling scheme and marked "Illustrative" where they are not already.
