UNUS London Official Governance Academy ISO/IEC 42001:2023 SRA Legal Track
Course Progress
09 / 10
Module 09 — The Operational Control — ISO 42001 Legal Track

"I reviewed it" is not a procedure.
The Human Oversight Protocol That Makes AI Review Evidenced.

This module takes Halstead & Cole from "fee earners read AI output before using it" — which produces no audit trail — to a three-tier oversight procedure, PROC-AIMS-HITL-001, where every GPT-4o output used in a client matter is reviewed against a structured 10-point checklist and a sign-off record lands in the matter file. Every question of the Pemberton Capital questionnaire is now answerable with documented evidence rather than reassuring generalities.

ISO/IEC 42001:2023 Cl. 9.1 Annex A.6 Human Oversight Three Review Tiers Ten-Item Tier 1 Checklist SRA AI Guidance (Feb 2026) Pemberton Q5 Self-Paced · Interactive
Section 01 — The Scenario

The Review That Happened but Cannot Be Proved

You are looking at Sophie Chen's recent file. She is a junior paralegal in the dispute resolution team at Halstead & Cole. She uses GPT-4o for case chronologies. She will tell you she reviews every output before using it. She is not lying. She does read it. But "reading it" is not a review standard, and a review with no standard produces no audit trail — and that is precisely the failure mode the courts have already addressed.

Firm
Halstead & Cole LLP
SRA Number
HC-SRA-2891
Fee earner in scope
Sophie Chen · Paralegal · DR team
System in scope
AIMS-SYS-001 · GPT-4o · Tier 1 (High)
Review practice
Unstructured read-through · No record
Court precedent
Ayinde v LB Haringey · [2025] EWHC 1383
Pemberton Q5
"What oversight do you apply to AI-generated work product?"
Severity rating
High
What Sophie found when she tried to evidence her review

Sophie was asked by the COLP, Priya Anand, to evidence her review of a GPT-4o-generated chronology that had gone into a matter file the previous month. Sophie said: "I read it."

Priya asked three follow-up questions. Sophie's answers exposed the absence of procedure:

"Did you check the citations?" Sophie could not point to a record of having done so. She said she thought she had. She was not certain.

"Did you check the dates against the source bundle?" Again, no record. She had scanned the chronology in the document viewer and assumed the dates were correct.

"Did you record your sign-off in the matter file?" No. There was no entry, no form, no email, no annotation. The matter file showed the chronology had been produced and used. It showed no evidence that it had been reviewed.

Sophie's review was not dishonest. It was real, in the sense that she did read the document. But it was not supervision. Supervision requires a defined standard, applied consistently, with evidence that it occurred. Without those three things, the review is invisible — to the SRA, to the courts, and to the firm's own defence if the chronology turns out to contain an error.

"An undocumented review that leaves no trace in the matter file is not supervision under SRA Principles. It is supervision in name only."

— UNUS London · Gap 8 Article · Cl. 9.1 Analysis

The gap that turns supervision into a feeling

ISO/IEC 42001:2023 Clause 9.1 requires the firm to retain appropriate documented information as evidence of the results of monitoring and measurement. The SRA's Principle 5 requires proper supervision. The courts have made clear — most recently in Ayinde [2025] EWHC 1383 — that reading AI output is not the same as supervising it. A review with no standard, no checklist, no record, and no escalation path produces no audit trail and no defence. Gap 8 is rated High because the failure mode is invisible to the firm until a regulator, a court, or a client asks the question that demands evidence.

Section 02 — Why This Gap Is High Severity

What Happens When Review Leaves No Trace

A missing human oversight procedure is not an administrative oversight. It is a procedural blind spot — and the consequences fall on three parties simultaneously: the fee earner who cannot evidence their work, the COLP who cannot evidence supervision, and the firm that cannot answer Pemberton Q5 with anything more substantial than a hopeful promise.

Court Sanction Exposure

In Ayinde v The London Borough of Tower Hamlets [2025] EWHC 1383, eighteen fabricated case citations were submitted in a judicial review claim. The submissions were presumably read before filing. The problem was not an absence of eyes on the document — it was the absence of a structured review that would have caught hallucinated citations. A "verify every citation against the primary source" checklist would have caught the failure. An unstructured read-through did not. Halstead & Cole's GPT-4o usage sits in the same failure pattern unless a structured checklist is in place.

SRA Supervision Finding

SRA Principle 5 requires proper supervision of legal work. The SRA's February 2026 AI guidance makes clear that this obligation extends to AI-generated content. Supervision has a specific meaning: it is a process — a defined standard of review, applied consistently, with evidence that it occurred. An undocumented review that leaves no trace in the matter file is not supervision under SRA Principles. It is supervision in name only — and a future SRA thematic review will be able to demonstrate the gap with a single document request.

Repeating Failure Patterns

When a fee earner finds an AI error — a wrong citation, a transposed date — and corrects it without recording the error, the same error type recurs in the next output, and the one after that. The risk register in REG-AIMS-RISK-001 is not updated because the pattern is invisible. The same prompt configuration that produces the error continues to be used. The residual risk score for the relevant factor becomes progressively less accurate over time. The firm's governance loses integrity in slow motion.

Pemberton Q5 Unanswered

Pemberton Capital's fifth questionnaire question asks: "what oversight do you apply to AI-generated work product?" The honest answer today is "fee earners read the output." A response of "we apply professional supervision to AI-generated content" is not an answer either — it is a description of intent. Without a documented three-tier review standard, a Tier 1 checklist, a sign-off form, and an incident escalation pathway, the firm cannot point to a process. The £620,000 panel relationship is exposed.

The ITIL perspective: monitoring without records is not monitoring

In ITIL 4 terms, Gap 8 represents a failure in the Monitoring and Event Management practice — and specifically the absence of a "detection and resolution record." ITIL requires that monitoring produce actionable records: events classified, incidents logged, problems investigated, resolutions documented. A fee earner noticing an AI error and fixing it without recording the finding is the ITIL equivalent of a service desk detecting an incident and resolving it without opening a ticket. The operational data needed to prevent recurrence never accumulates.

Pattern visibility — the argument for recording minor errors too

The four oversight non-conformances identified in Clause 9.1 audits (NF1: no review standard; NF2: no matter-file record; NF3: no escalation path; NF4: no aggregated performance data) are interconnected. NF2 is the most damaging because it disables NF3 and NF4 — without per-matter records, patterns cannot be detected, escalation cannot be triggered, and quarterly performance reports cannot be produced. The whole oversight framework depends on the single act of recording each review.

Section 03 — What the Standard Actually Requires

Reading Clause 9.1 in Plain English

ISO/IEC 42001:2023 Clause 9.1 is short in its written text but operationally demanding. Per ISO/IEC Directives Part 2, the word shall denotes a requirement; should denotes a recommendation; may denotes a permission. The interactive table below lets you filter by obligation type — try it.

Where this clause lives in the standard

Clause 9.1 sits inside Clause 9 (Performance Evaluation). It is the clause that converts "we review AI output" from an aspiration into a documented obligation. The "shall retain appropriate documented information as evidence of the results" requirement is the operative phrase — it admits no exception. Clause 9.1 is mandatory and auditable; firms cannot justify omitting it.

Three-Tier Oversight Model — Risk-Proportionate Review Standard
1
Tier 1 — 100% Review
Every AI output reviewed against the 10-point Tier 1 checklist before use in any client matter. Matter-file sign-off required per output. No AI-generated content reaches a client without a named fee earner's documented approval.
High risk
Applies to: AIMS-SYS-001 (GPT-4o) · AIMS-SYS-003 (Copilot Outlook — pending return to Active)
2
Tier 2 — Supervised Spot-Check
Supervising solicitor reviews 1 in every 5 AI-assisted outputs per fee earner per month. Spot-checks recorded on the Tier 2 log. A pattern of errors triggers Tier 1 escalation for that fee earner's use of the system.
Medium risk
Applies to: AIMS-SYS-002 (Copilot Word) · AIMS-SYS-004 (LEAP AI)
3
Tier 3 — Periodic Audit
AGL or IT Lead conducts a quarterly sample audit of AI outputs from Low-risk systems. No per-output sign-off required. Audit findings reported in the AIMS quarterly performance report.
Low risk
Applies to: Administrative AI tools · Non-client-facing AI features · Grammar and formatting assist
Tier 1 · Every output, structured checklist, matter-file record
Tier 2 · 1 in 5 outputs, supervising partner sample review
Tier 3 · Quarterly audit sample, AIMS performance report
Clause 9.1 obligations — click a filter to narrow the view
Shall The organisation shall determine what needs to be monitored and measured. For law firm AI systems, this means the specific output quality dimensions relevant to each system — citations, dates, scope, legal propositions — not generic "output quality."
Shall The organisation shall determine the methods for monitoring, measurement, analysis and evaluation to ensure valid results. Valid results require a structured checklist calibrated to the system's known failure modes — not a generic read-through.
Shall The organisation shall determine when monitoring and measuring shall be performed. For Tier 1 systems, the answer is "every output, before use." The cadence is not discretionary — it is built into the procedure.
Shall The organisation shall determine when results shall be analysed and evaluated. This means a defined cycle — the COLP's quarterly performance report — for the aggregated view, plus per-incident analysis for material errors.
Shall The organisation shall determine who shall analyse and evaluate these results. The AGL and COLP are named in the RACI matrix (RACI-AIMS-001); the procedure identifies them again here for the oversight-specific role.
Shall The organisation shall evaluate the performance and effectiveness of the AI management system. Performance = are the systems delivering the intended benefits? Effectiveness = is the governance framework operating as designed? Both require data.
Shall The organisation shall retain appropriate documented information as evidence of the results. This is the operative phrase that turns "I read it" into "I can evidence that I read it." Every review leaves a trace.
Shall Per Annex A.6 (human oversight of AI systems), the organisation shall ensure that AI systems are used under appropriate human oversight proportionate to their impact. The three-tier model implements this proportionate oversight.
Should The organisation should define the review standard by reference to the specific failure modes of each AI system — citation hallucination, date transposition, scope drift, confident misstatement of law. Generic review standards fail to catch system-specific errors.
Should The organisation should maintain aggregated performance data at practice-area and firm-wide level, enabling pattern detection across fee earners and over time. Individual reviews are necessary; aggregated analysis is what makes them useful.
May The organisation may apply different review standards to different outputs from the same system — for example, applying Tier 1 to outputs used in pleadings and Tier 2 to outputs used in internal memos. Risk-proportionate review is allowed.

Why Clause 9.1 vs Annex A.6 matters — and why both apply

Annex A.6 (human oversight) is a recommended control. Clause 9.1 (monitoring and measurement) is a core requirement. The distinction matters: Annex A.6 controls can be omitted if a firm documents why they are not applicable; Clause 9.1 cannot. PROC-AIMS-HITL-001 satisfies both simultaneously — the three-tier review model implements Annex A.6, while the per-matter sign-off records and quarterly performance report satisfy Clause 9.1's documentation requirement. A single procedure, two compliance obligations, one audit trail.

The integration point with Clause 6.1 and the lifecycle procedure

PROC-AIMS-HITL-001 draws escalation paths from REG-AIMS-RISK-001 (risk register scores) and feeds into PROC-AIMS-LIFE-001 (lifecycle procedure) when material errors are detected. A Material or Serious error triggers T2 (incident) in the lifecycle, which moves the affected system to Under Review. Without the oversight procedure, the lifecycle procedure has no input from operational reality; without the lifecycle procedure, the oversight procedure has nowhere to escalate its findings.

Section 04 — The Regulator's View

How the SRA's Supervision Standard Maps to Clause 9.1

The SRA's existing supervision framework — never written specifically with AI in mind — already creates the same requirement when applied to AI-generated work. Three instruments operate together. The mapping below shows where each PROC-AIMS-HITL-001 element comes from.

"AI does not reduce the professional obligations of solicitors. Where AI is used to produce work product, that product requires the same standard of supervision as work produced by any other means."

— SRA · Compliance Tips for Solicitors Regarding the Use of AI and Technology · 9 February 2026
SRA · AI Guidance

SRA Compliance Tips for Solicitors (Feb 2026) · "Same standard of supervision"

The SRA's framing — "the same standard of supervision as work produced by any other means" — makes two things clear. First, AI output is work product, and work product requires supervision. Second, the standard cannot be relaxed because the producer is an AI system. For Halstead & Cole, the Tier 1 review checklist implements this "same standard" obligation in operational terms — with the additional point that AI systems have specific, documented failure modes that the checklist is calibrated to detect.

SRA · Competence

SRA Code of Conduct for Solicitors 2019 · Rules 3.2 & 3.3

Rule 3.2 requires solicitors to maintain the level of competence and legal knowledge needed to practise effectively. Rule 3.3 requires supervision of staff and adequate training. The SRA has confirmed both obligations extend to the AI tools used in practice — including understanding their current limitations. A solicitor who reviews AI output without understanding that LLMs hallucinate citations is not exercising the competence Rule 3.2 requires.

SRA · COLP Duty

SRA Code of Conduct for Firms 2019 · Rule 8.1(a)

The COLP must take all reasonable steps to ensure the firm, its managers, employees, and interest holders comply with their obligations under the SRA's regulatory arrangements. Without a documented oversight procedure, the COLP cannot evidence that "reasonable steps" have been taken to govern AI use. The Tier 1 sign-off records and the quarterly performance report are the evidence base for Rule 8.1(a) compliance.

Court · Precedent

Ayinde v The London Borough of Tower Hamlets [2025] EWHC 1383 (Admin)

In Ayinde, eighteen fabricated case citations were submitted in a judicial review claim. The court drew explicit attention to the responsibility of legal representatives to verify AI-generated material before relying on it. The judgment is now routinely cited as the operational baseline for AI output review in legal practice. A Tier 1 checklist that mandates verification of every citation against the primary source is precisely the standard the court expected.

The mapping in one sentence

One procedure. Two regulators. One court-tested standard.

ISO 42001 Clause 9.1, the SRA's February 2026 AI guidance, SRA Rules 3.2 / 3.3 / 8.1(a), and the Ayinde judgment all converge on the same requirement: a documented, evidence-producing human oversight procedure for AI-generated work product. PROC-AIMS-HITL-001 — built in the next two sections — is the single procedure that satisfies the international standard, evidences compliance with the SRA's supervisory duty, meets the standard the court expected in Ayinde, and answers Pemberton Capital's fifth questionnaire question with a structured process rather than a hopeful promise.

Section 05 — The Fix (Part One)

The Three Tiers of PROC-AIMS-HITL-001

The oversight requirement is not uniform across all AI systems. A risk-proportionate model — derived from the risk register scores in REG-AIMS-RISK-001 — applies different review standards to High, Medium, and Low risk systems. Click each tier below to reveal its content. Together they form the firm's complete oversight architecture.

Tiers reviewed
0 / 3

Tier 1 is the most demanding review standard and applies to every AI system classified as High risk in REG-AIMS-RISK-001. At Halstead & Cole, this currently includes AIMS-SYS-001 (GPT-4o) and AIMS-SYS-003 (Copilot in Outlook) once it returns to Active status.

What the review looks like in practice:

  • Every output is reviewed against the 10-point Tier 1 checklist (covered in detail in the next section) before the output is used in any client matter.
  • The reviewer is a qualified solicitor (or a supervised trainee acting under a qualified solicitor's oversight) with sufficient expertise in the relevant area of law to assess the output's accuracy.
  • The HITL Sign-Off Form is completed per output. The form captures: the system used; the matter reference; the reviewer name and date; each of the 10 checklist items confirmed or flagged; any errors detected (F1–F4 classification); the reviewer's overall sign-off or non-sign-off decision.
  • The completed form is filed in the matter file — either as a separate document or as an entry in the matter management system. The form is evidence that the review occurred, what it checked, and what it found.
  • Material or Serious errors trigger the incident escalation pathway (covered in the next section) within 2 working days of detection.

Why Tier 1 cannot be delegated downward: The review must be conducted by a qualified solicitor — or a supervised trainee under qualified oversight — because the assessment involves judgement calls about the accuracy of legal propositions, the sufficiency of verification, and the materiality of any errors detected. A paralegal checking citations against a checklist is operating at the verification layer; the judgement layer requires qualified supervision. This is the SRA's Principle 5 applied to AI output.

Tier 2 is a lighter review standard that applies to AI systems classified as Medium risk in REG-AIMS-RISK-001. At Halstead & Cole, this currently includes AIMS-SYS-002 (Copilot in Word) and AIMS-SYS-004 (LEAP AI). The Tier 2 model acknowledges that reviewing every output is disproportionate for lower-risk tools — but rejects the alternative of no review at all.

What the review looks like in practice:

  • The fee earner uses the AI output in the ordinary course of work. Tier 2 does not require per-output sign-off — that level of scrutiny is reserved for Tier 1 systems where the error cost is highest.
  • The supervising solicitor reviews 1 in every 5 AI-assisted outputs per fee earner per month. The sample is selected at random — not self-selected by the fee earner — to preserve the integrity of the oversight.
  • The Tier 2 spot-check log captures: the output reviewed; the date of review; the supervising solicitor's name; any errors or concerns identified; the F1–F4 classification if applicable; and the action taken (no action / correction / escalation).
  • A pattern of errors triggers Tier 1 escalation for that fee earner's use of the system. Specifically: 3 or more spot-checks in a quarter with identified errors → the AGL reassigns that fee earner's use of the system to Tier 1 for the next quarter.

Why 1-in-5 is the working standard: The cadence is calibrated to two competing requirements: enough sample size to detect patterns, but light enough to be operable in a busy practice. ISO 42001 auditors typically accept 20% sample rates for medium-risk oversight as proportionate. Smaller samples make pattern detection unreliable; larger samples make the procedure unworkable. The 1-in-5 figure is the working consensus.

Tier 3 is the lightest oversight standard and applies to AI systems classified as Low risk in REG-AIMS-RISK-001. At Halstead & Cole, this currently covers administrative AI tools (billing narrative drafting, scheduling assistance), non-client-facing AI features (grammar suggestions, code-completion), and similar utility functions where the consequence of an error is contained.

What the review looks like in practice:

  • No per-output sign-off is required. Tier 3 systems are used in the ordinary course of work without a per-output review record.
  • The AGL or IT Lead conducts a quarterly sample audit. A random sample of 5–10 AI outputs from each Tier 3 system is reviewed against a brief Tier 3 audit checklist (focused on obvious errors, security issues, and scope anomalies).
  • Audit findings are reported in the AIMS quarterly performance report. The report goes to the COLP, the AGL, and senior leadership. Findings include: number of outputs sampled; number of errors found; any patterns detected; and any recommendations for tier reclassification.
  • A pattern of errors triggers tier reclassification. If Tier 3 audit findings show error rates that are disproportionate to the risk tier, the AGL recommends moving the system to Tier 2 for the next quarter, with the prospect of returning to Tier 3 if performance improves.

Why even low-risk systems need oversight: The Tier 3 standard is not "no oversight." It is oversight proportionate to the risk. Administrative AI tools can still produce errors that affect client invoices, internal records, or scheduling — and the absence of any oversight means those errors accumulate silently. The quarterly audit is the minimum viable mechanism to keep Tier 3 systems in check.

Important — tiers are reassessed quarterly

The tier classification of each AI system is not permanent. The AGL reviews the tier classification of every system as part of the quarterly AIMS performance report (RACI-AIMS-001). A Tier 2 system with consistent clean spot-checks may move to Tier 3; a Tier 3 system with persistent audit findings may move to Tier 2; a Tier 1 system with no errors in a quarter remains Tier 1 unless the risk register itself is revised. The tiers are a living governance instrument — not a one-time classification.

Section 06 — The Fix (Part Two)

The Tier 1 Checklist and the Incident Escalation Pathway

The Tier 1 review checklist is the operational core of the oversight procedure — the structured instrument that turns an unstructured read-through into a documented, evidenced review. The checklist below is calibrated to the four specific failure modes (F1–F4) that AI systems produce in legal contexts. Every checkbox matters; every gap leaves a known failure pattern undetected.

The four failure patterns the checklist must catch
F1 · Citation hallucination

Plausible but non-existent case citations, or real citations attributed to the wrong proposition. The Ayinde failure mode.

F2 · Date/figure transposition

Dates, party names, and financial figures silently transposed. Invisible to a fast reader because surrounding text is correct.

F3 · Scope drift

Output extends beyond the requested task — unsolicited analysis, content from outside the submitted documents.

F4 · Confident misstatement of law

Incorrect legal propositions stated in authoritative, fluent language. Passes casual read; fails specialist review.

PROC-AIMS-HITL-001 · Tier 1 Review Checklist · AIMS-SYS-001 (GPT-4o) · All HIGH-risk system outputs
01
Every case citation in the output has been verified against the primary source (Westlaw UK, Lexis+, BAILII, or the official law report). No citation used that has not been individually checked against the primary source.
02
Every date, deadline, and time period in the output has been cross-checked against the source documents provided to the AI. No dates taken from AI output without source verification.
03
Every party name, company name, and identifying reference has been confirmed correct. No AI output used where party names are incorrect or transposed.
04
Every financial figure, sum, or statutory threshold has been verified against the source document or current statutory authority. No figures relied upon from AI output without independent verification.
05
The scope of the AI output matches the scope of the instruction. Content that goes beyond the requested task has been identified and either removed or separately verified before inclusion.
06
Any legal proposition stated as fact by the AI has been independently verified against case law, statute, or authoritative commentary. No AI-stated legal proposition relied upon without independent confirmation.
07
The data minimisation rules for this system (REG-AIMS-DATA-001) were followed when preparing the input — no prohibited data categories were submitted, and any special category data was authorised in advance.
08
The output has been reviewed by a qualified solicitor (or a supervised trainee acting under a qualified solicitor's oversight) with sufficient expertise in the relevant area of law to assess its accuracy.
09
Any errors detected during the review have been corrected in the output, classified using the F1–F4 scheme, and recorded in the HITL Sign-Off Form. Material or Serious errors have been escalated per the incident pathway (below) within 2 working days.
10
The reviewer is satisfied that the output, as amended, accurately reflects the law and facts of the matter and is fit for the purpose for which it will be used. The reviewer signs the HITL Sign-Off Form with their name, signature, and date.
The incident escalation pathway

When the Tier 1 checklist detects a material error, the finding does not simply get corrected and forgotten. The escalation pathway routes the finding through the governance framework — preserving the operational signal that the risk register and lifecycle procedure depend on.

⚑ AI Incident Escalation Path — Material Error Detected During Review
Step 1 Fee earner corrects the error and completes the HITL Sign-Off Form, marking the error type (F1, F2, F3, or F4) and severity (Minor / Material / Serious). Minor errors are recorded in the form only.
Step 2 Material or Serious errors are reported to the supervising solicitor and the COLP within 2 working days of detection. The report includes the F-classification, the severity assessment, and the proposed correction.
Step 3 COLP assesses whether the error constitutes an AI incident under POL-AIMS-001, whether client notification is required (under DOC-AIMS-DISC-001 and SRA transparency obligations), and whether the error triggers T2 (incident) in PROC-AIMS-LIFE-001.
Step 4 AGL updates REG-AIMS-RISK-001 — the relevant risk factor row is re-assessed. If the incident reveals that the residual risk score for that factor is understated, the score is revised and the risk acceptance record updated.
Step 5 Pattern detection — if the same error type recurs across multiple matters or multiple fee earners, the COLP flags this in the quarterly performance report. Three or more instances of the same error type within a quarter triggers a mandatory review of the prompt version (Part B of REG-AIMS-SYS-001) and the Tier 1 checklist.

Why "minor" errors are still recorded

The distinction between Minor, Material, and Serious errors determines the escalation path — it does not determine whether the error is recorded. Every error detected during a Tier 1 review is recorded in the HITL Sign-Off Form, regardless of severity. Minor errors that recur become Material patterns. The pattern is only visible if the individual instances are recorded. A firm that only records Serious errors will systematically miss the early-warning signal that precedes them.

The integration point — HITL feeds the lifecycle procedure

PROC-AIMS-HITL-001 is not a standalone procedure. It is the operational input that makes PROC-AIMS-LIFE-001 (the lifecycle procedure from Module 06) actually function. T2 — the incident trigger — is fired only if there is an incident record to fire it. Without the HITL Sign-Off Form capturing detected errors, T2 never fires, the Under Review stage never opens, and the firm's AI systems drift undetected. Tier 1 reviews are the sensor network; the lifecycle procedure is the response system.

Section 07 — Document Control & Application

How PROC-AIMS-HITL-001 Is Filed, Signed & Applied

A human oversight procedure without document control is a memo. ISO 42001 Clause 7.5 requires that all AIMS documentation carry a defined set of metadata: ID, version, owner, classification, issue date, review cycle, retention, and related documents. The block below shows what Halstead & Cole's control block looks like when PROC-AIMS-HITL-001 is ready to issue.

CONTROLLED DOCUMENT — PROC-AIMS-HITL-001 v1.0 ● READY FOR ISSUE
Document IDPROC-AIMS-HITL-001
Version1.0
TitleHuman-in-the-Loop Oversight Procedure
StatusDRAFT — AWAITING ISSUE
Standard refsISO/IEC 42001:2023 Cl. 9.1; Annex A.6; ISO/IEC Directives Part 2; SRA Code of Conduct 2019 Rules 3.2/3.3/8.1(a); SRA Compliance Tips Feb 2026; SRA Principle 5
OwnerCOLP (review and quarterly reporting); AGL (operational execution and tier assignments)
ClassificationInternal — Controlled
Date of issue[Date of COLP signature]
Next reviewAnnual minimum; immediate trigger on regulatory change, material new AI deployment, or any pattern of errors detected in quarterly report
Retention7 years from date of supersession (UK Companies Act 2006)
Related docsPOL-AIMS-001 (policy); RACI-AIMS-001 (roles); REG-AIMS-SYS-001 (system register); REG-AIMS-RISK-001 (risk register); PROC-AIMS-LIFE-001 (lifecycle); REG-AIMS-DATA-001 (data); DOC-AIMS-DISC-001 (disclosure)
What happens after PROC-AIMS-HITL-001 is issued
01
Day 0 · COLP issues the procedure

PROC-AIMS-HITL-001 becomes ACTIVE

Status changes from DRAFT to ACTIVE. The date of issue is recorded. The seven-year retention clock starts. From this date forward, every Tier 1 system output must be reviewed against the 10-point checklist and a sign-off form completed before use in any client matter.

02
Within 7 days · Fee earner briefing

Every fee earner briefed on the Tier 1 checklist

The COLP, supported by the AGL, briefs every fee earner who uses AIMS-SYS-001 (GPT-4o) or any other Tier 1 system on the 10-point checklist. The briefing covers what each item checks, why it matters, and how to complete the sign-off form. Each fee earner signs an acknowledgement. Briefings are recorded in the AIMS training log.

03
Within 14 days · First Tier 1 sign-off

The first evidenced review is in a matter file

The first Tier 1 review of a GPT-4o output produces the first completed HITL Sign-Off Form. The form is filed in the matter file. From this point forward, every matter using Tier 1 AI output has documentary evidence of the review — meeting the SRA's supervision standard and Clause 9.1's documentation requirement.

04
Day 30 · Pemberton questionnaire answered

Question 5 answered with documented procedure

Halstead & Cole responds to Pemberton Capital's fifth questionnaire question with a copy of PROC-AIMS-HITL-001 attached, the tier classification of each system, the date of the first completed sign-off form, and a sample anonymised form showing the review standard. All five questionnaire questions are now answered with documented procedures.

05
Day 90 · First quarterly AIMS performance report

Aggregated performance data begins to accumulate

The COLP produces the first quarterly AIMS performance report covering Tier 1 reviews (number, error rate, F1–F4 distribution), Tier 2 spot-checks, Tier 3 audit findings, and any pattern signals detected. The report establishes the baseline against which future quarterly performance is measured. Pattern detection — the strategic value of the procedure — becomes operational.

What Gap 8 closure delivers — and what remains

PROC-AIMS-HITL-001 closes the operational oversight loop. With this procedure in place, the firm can evidence that every High-risk AI output is reviewed against a structured standard, that errors are recorded and escalated, and that aggregated performance data informs governance decisions. One gap remains in the AIMS architecture: G9 — supply chain governance — which addresses third-party and self-hosted AI systems where the procurement layer creates additional obligations. That gap is the subject of the next article in the series.

Quick check — test your understanding
Question 1 of 3 — A fee earner reads an AI-generated chronology, finds a transposed date, corrects it, and files the output without telling anyone. Which oversight non-conformance has occurred?
Correct Answer: C The defining failure here is that the review — even though it happened — left no trace in the matter file. NF2 is the most damaging non-conformance because it disables NF3 and NF4. Without the HITL Sign-Off Form, the COLP cannot evidence the review occurred, the error cannot be escalated to REG-AIMS-RISK-001, and no pattern data accumulates for the quarterly performance report. The next fee earner who uses GPT-4o for a chronology will make the same date transposition mistake — and no one will know it has happened before.
Question 2 of 3 — A junior paralegal reviews a GPT-4o output against the Tier 1 checklist, finds a fabricated citation, and flags it on the sign-off form as Material. Under PROC-AIMS-HITL-001, what happens next?
Correct Answer: B Material and Serious errors both trigger formal escalation per Step 2 of the incident pathway — Material errors are not "below" the threshold, they are at it. The COLP's role at Step 3 is to assess three things: whether the error constitutes an AI incident under POL-AIMS-001, whether client notification is required, and whether it triggers T2 (incident) in PROC-AIMS-LIFE-001 — moving the affected system to Under Review. The Tier 1 sign-off is the trigger; the COLP's assessment is the gate. Both happen.
Question 3 of 3 — Which evidence package is most likely to satisfy a future SRA thematic review on AI oversight?
Correct Answer: C An SRA thematic review evaluates three things: does a procedure exist (the document), is it actually used (the sign-off forms in real matter files), and is the data being acted upon (the quarterly report showing pattern detection and trend analysis). Answer C satisfies all three. Answer A is a statement of intent without evidence. Answer B shows the procedure exists but does not show it operates. Answer D is external communication, not internal supervision. The combination of procedure + operational records + aggregated analysis is the evidence package that converts intent into evidence — which is precisely what the SRA's Principle 5 supervision obligation requires.
Section 08 — Completion Gate

Module 09 Completion Checklist

Tick every box below to complete the programme. This is the final module — closing it means every operational control is in place for Halstead & Cole's AI systems. Each item maps to a specific Clause or SRA obligation that audit will test.

Checklist progress
0 / 16
Understanding the Scenario
I can explain why "I reviewed it" is not a supervision standard and what makes a review evidenced
I can describe the Ayinde judgment and why a structured checklist would have caught the failure where an unstructured read-through did not
I understand why Gap 8 is rated High — the failure mode is invisible until a regulator or court asks the question that demands evidence
I can name the four oversight non-conformances (NF1–NF4) and explain how they are interconnected
Understanding the Standard
I can list the five "what" questions Clause 9.1 requires the firm to determine (what to monitor, methods, when, when analysed, who)
I can explain the distinction between Clause 9.1 (mandatory) and Annex A.6 (recommended control) and why both apply
I can describe why the "retain appropriate documented information" requirement is the operative phrase that converts intent into evidence
I can name the four AI output failure modes (F1–F4) and explain why each one requires a specific check, not generic "review"
Understanding the Three Tiers
I can describe each tier (1: 100% review / 2: spot-check / 3: audit) and identify which Halstead & Cole systems fall into each
I can name the 10 Tier 1 checklist items and explain why each one specifically catches a known AI failure mode
I can describe the 5-step incident escalation pathway and explain why every error (including Minor) is recorded
I can explain why the HITL Sign-Off Form is the integration point that makes PROC-AIMS-LIFE-001 (Module 06) actually function
Programme Readiness
I can describe the document-control block fields and why each one is mandatory under Clause 7.5
I can describe the five post-issue steps (briefing, first sign-off, Pemberton response, quarterly report, pattern detection)
I can explain why the evidence package for an SRA review is procedure + sign-off forms + quarterly report — not any one of these alone
I am ready to confirm the firm's commitment to issue PROC-AIMS-HITL-001, brief all fee earners, and complete the first Tier 1 sign-off within 14 days

Beyond Module 09 — Towards Module 10 (Course Finale)

With PROC-AIMS-HITL-001 in place, Halstead & Cole's AIMS is operationally complete for every internally deployed AI system. The remaining gap in the architecture — G9, supply chain governance — addresses third-party and self-hosted AI procurement where additional oversight obligations apply. Firms that procure AI capabilities through multiple vendors, aggregators, or self-hosted deployments will need REG-AIMS-SUP-001 (the supplier assessment regime) to close that final gap. That article is the next in the UNUS London gap series.