Automate the queue, not accountability
Altivate | Published 14 June 2026 | Updated 26 July 2026
1. Public-sector AI is already an operating question
The OECD’s Digital Government Outlook 2026 reports AI use in 35 of 36 surveyed OECD countries. Adoption is strongest in internal processes and public services, and lower in policymaking and accountability functions.
That denominator and geography matter. This is not a Gulf adoption statistic and it does not say that 35 governments have mature controls. In the same survey:
- 14 of 36 require pre-deployment risk assessments;
- 12 of 36 have internal review committees;
- 11 of 36 conduct post-deployment audits;
- 10 of 36 report any impact measurement.
The useful finding is the gap: use is widespread among the surveyed countries, while operational review and impact measurement are less common.
That changes the public-sector decision. The question is no longer whether an organisation has an AI policy. It is whether a resident, caseworker, auditor and decision owner can understand what happened in one real case.
The paper uses one consequence boundary:
The same model may safely assist in one step and require strict control in the next.
Context: Altivate’s work with the public sector and AI technologies.
2. The public-service consequence ladder
| Level | Example | What AI may do | Accountability boundary |
|---|---|---|---|
| Inform | Explain service requirements | Retrieve and summarise approved information | Cite source, language and effective date |
| Assist | Draft a case note or translation | Prepare work for an official | Official reviews; original evidence retained |
| Route | Direct a request to the right team | Classify and prioritise | Service-level monitoring; no rights decision |
| Recommend | Flag possible eligibility or fraud issue | Present evidence and a proposed outcome | Authorised official decides and records reason |
| Decide | Grant, refuse, suspend, pay or enforce | Only inside a legally and operationally approved design | Named decision owner, traceability, notice, review and appeal |
This ladder is not a promise that every service should reach “decide.” For consequential cases, recommendation with human determination may be the correct permanent architecture.
The dividing line is sometimes subtle. Prioritising emergency inspections affects who is seen first. Translating a resident’s answer can change how evidence is interpreted. A fraud-risk score can become a de facto decision if staff are expected to accept it. Governance should follow the real effect, not the label in the system.
Ask five questions:
- Can this output delay, deny or reduce a service?
- Can it initiate investigation or enforcement?
- Can a person know that AI materially influenced the case?
- Can the responsible official disagree without penalty or technical friction?
- Can the person obtain review by someone with authority to change the outcome?
If the first two are “yes,” the final three need concrete answers before deployment.
3. Automate the queue
Public services contain large volumes of work that are necessary but not themselves exercises of public authority:
- identify the service or form relevant to a request;
- check whether required documents are present;
- extract fields into a draft case record;
- translate or simplify approved information;
- find the policy and evidence relevant to an official;
- route a case by location, subject or urgency;
- detect duplicates;
- draft routine correspondence;
- summarise a case history for review.
These are strong starting points because they shorten administrative queues while preserving the decision boundary.
They still need controls. A missing-document checker should not silently reject the application. A translation should preserve the original. A routing classifier needs monitoring across languages and service populations. A case summary must link to the underlying record so an official can inspect what was omitted.
The operating measure should follow the service outcome:
- time to first useful response;
- cases routed correctly the first time;
- incomplete applications resolved rather than abandoned;
- rework created by extraction or translation errors;
- staff time returned to decision and support work;
- performance by language, channel and service;
- complaints, corrections and appeals associated with the AI-supported step.
“Number of AI interactions” is an activity measure, not public value.
4. Do not automate away the reason
When a system recommends or makes a consequential decision, preserve a reason that is meaningful at the case level.
A model explanation is not necessarily that reason. It may be a plausible description generated after the output, rather than a faithful account of the factors that produced it. A complete case record should identify:
- the rule or policy in force at the time;
- the evidence used and its source;
- the data fields and versions relied on;
- the model, prompt, workflow and threshold version;
- the recommendation or action;
- the official who reviewed or approved it;
- overrides and their reasons;
- the notice given to the person;
- the review or appeal route;
- the final outcome after any challenge.
This record serves several audiences. The resident needs a reason they can act on. The caseworker needs the evidence and rule. The service owner needs to see patterns in error and override. The auditor needs to reconstruct the process.
Do not make access to the record depend on replaying the model. Models, prompts and source indexes change. Store the material facts of the case when the action occurs.
5. Human review must have authority
The UAE AI Ethics Principles and Guidelines provide a useful regional control direction. They call for human case evaluators where significant decisions are challenged, external audit for critical decisions, human validation for life-and-death decisions, and traceability in critical contexts. This is national guidance, not legislation, and this paper does not turn it into a legal obligation.
Its practical lesson is that a human checkpoint has four requirements:
- information – the official can see the evidence, uncertainty and policy;
- time – the workflow allows meaningful review;
- authority – the official can change the outcome;
- independence – performance incentives do not make agreement automatic.
A button labelled “approve AI recommendation” does not establish human control. If rejecting takes five extra screens, if the official cannot see the source record, or if overturning the score requires manager escalation, the system is designed for confirmation.
Measure override patterns, but do not punish officials merely for disagreement. A cluster of overrides can reveal model drift, a policy change, missing local context or insufficient training. It is an operational signal.
For a challenge, route the case to a reviewer with genuine power to alter the result and access to the original evidence. Sending the same inputs through the same model again is reconsideration by the same system, not independent review.
6. Procurement is part of governance
The OECD’s 2026 framing groups trustworthy government AI around enablers, guardrails and engagement, and identifies procurement, data, transparency and evaluation as practical conditions for scale.
That matters because many control failures are fixed or foreclosed in the contract.
Before procurement, require answers to:
- Which data is used for service delivery, product improvement or model training?
- Where are prompts, records, embeddings, logs and backups processed and retained?
- How are resident and caseworker permissions enforced in retrieval?
- Can the authority export case-level traces and evaluation results?
- Who may change the model, system prompt, tools, thresholds or source index?
- What notice and revalidation follow a material change?
- Can the authority suspend one function without losing access to its records?
- How are subcontractors and hosted-model providers disclosed?
- What is the process and maximum time for security, bias and accuracy incidents?
- Does termination return the data, configuration, audit history and evaluation set in usable form?
- How will accessibility be tested with disabled residents and assistive technologies?
- Which population groups and service channels will be monitored for disparate error, delay, exclusion or escalation rates?
Avoid a contract that promises “explainable AI” without naming the record available for one decision. Avoid a generic accuracy commitment without service-specific error costs and population breakdowns. Avoid a right to audit that excludes the evidence needed to exercise it.
The goal is not to prescribe one hosting model. It is to make accountability portable across hosting choices.
7. A Gulf operating context
Regional guidance is developing, but it should not be flattened into one Gulf rule.
The UAE AI Ethics Principles and Guidelines provide detailed ethics guidance, including challenge, audit, human validation and traceability concepts. They are guidance.
Saudi Arabia’s Digital Government Authority describes SDAIA’s AI Ethics Principles as a way to strengthen data and AI governance and mitigate economic, psychological, social, security and political harms. SDAIA’s official publication index separately lists Generative Artificial Intelligence for Government as guidance for safe and reliable use aligned with Saudi regulation, privacy and human rights. Those indexed descriptions are directional evidence only; confirm the current operative guidance and applicable law for the service before deployment.
These sources support a direction, not a regional legal conclusion:
- govern the specific service and consequence;
- preserve human accountability for significant decisions;
- make critical activity traceable;
- assess risk before deployment and impact after it;
- connect AI governance with data, security, procurement and service management.
The applicable legal position will depend on the country, authority, service, data and sector. Confirm it with the responsible legal and policy owners rather than copying an ethics principle into a compliance matrix.
For each proposed service, require jurisdiction-specific legal review before deployment and after a material change. The review should address equality and non-discrimination duties, accessibility, data protection, administrative-law obligations, notice, reason-giving, review and appeal. Test outcomes by relevant population group and channel, using privacy-preserving methods, and give the service owner authority to suspend deployment when disparate impact or accessibility failure crosses an agreed threshold.
8. A service-level governance packet
For each AI-supported service, publish internally one short operating packet:
Purpose
The service outcome, affected population, AI-supported steps and explicitly prohibited uses.
Consequence
The highest ladder level, possible harm, affected rights or payments, and named accountable owner.
Evidence
Data sources, lawful and policy basis as determined by the authority, quality limitations, retention and access.
Evaluation
Task success, error costs, language and population breakdowns, baseline, acceptance thresholds and known unresolved cases.
Human authority
Who reviews, what they see, how they disagree, when a second reviewer is required and how a person challenges the outcome.
Operation
Model and workflow version, monitoring, change control, incident response, suspension and case-reconstruction method.
Impact
Queue time, completion, abandonment, corrections, complaints, appeals, reversals and distribution of benefits and errors.
This packet should be reviewed before launch and on a fixed cadence afterward. Re-open it when the model, data, policy, population, supplier or delegated action changes.
9. A release sequence
Phase one: internal assistance. Start with retrieval, drafting and summarisation for officials. Retain original records and require review. Establish the evaluation set from real cases.
Phase two: resident-facing information and routing. Provide cited answers, form support and status navigation. Measure failure by language and channel. Give a clear route to a person.
Phase three: recommendation. Let the system surface evidence or risk for an authorised official. Preserve reason, disagreement and case-level trace. Audit whether recommendation becomes automatic in practice.
Phase four: bounded decision, only where justified. Use deterministic eligibility rules where possible. Any AI contribution to a consequential decision needs explicit authority, impact evidence, notice, review, appeal and a tested suspension path.
Not every service should enter phase four. A programme that stops at phase two after materially shortening a queue has delivered public value.
10. Accountability is service-specific
The OECD figures describe 36 surveyed OECD countries, not Gulf governments or all countries. The UAE source is ethics guidance rather than legislation, while the Saudi sources support only the high-level governance framing used here. This paper does not determine when automated decision-making is lawful in a particular service.
A human checkpoint is not sufficient when the reviewer lacks information, time, authority or independence. Equality, accessibility and disparate-impact monitoring are operating controls, not one-time ethics checks. The framework prescribes no cloud, model or procurement route. Government guidance was checked 26 July 2026 and must be refreshed together before revision.
11. The executive agenda
- Choose one public service and map every AI-supported step onto the consequence ladder.
- Mark where output first affects priority, eligibility, payment, enforcement or access.
- Name the decision owner, case-level reason, notice, challenge route and remedy at that boundary.
- Audit whether recommendations become decisions in practice through workload or interface design.
- Automate queue work first and measure service improvement without delegating accountability.
Then look backward for the queue work that can be improved without delegating that authority: retrieval, document checking, translation, routing, drafting and case preparation. Those steps often hold substantial service friction and generate the evidence needed for safer later decisions.
Altivate works across public-sector cloud, data, enterprise applications and AI in Saudi Arabia, the UAE, Jordan and India. Public value does not require maximum decision automation. It requires less friction, better evidence and an accountable institution at the point of consequence.
Apply the charter to the public sector, or return to White Papers.
Sources and verification status
Checked 26 July 2026.
International public-sector evidence – [read]
- OECD, Digital Government Outlook 2026 – Adopting and governing AI in government
- the 35-of-36 adoption finding, the four operational-control denominators and the enablers, guardrails and engagement framework.
- OECD, Governing with Artificial Intelligence: Are governments ready?
- government-AI governance, transparency, data, procurement and evaluation context.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile – system-level risk and evaluation guidance.
Regional public guidance
- UAE AI Office, AI Ethics Principles & Guidelines – [PDF text read]; challenge review, critical-decision audit, human validation and traceability. December 2022. Guidance, not legislation.
- Saudi Digital Government Authority, AI Ethics Principles – [official indexed description; direct fetch blocked]; governance and harm-mitigation scope only.
- SDAIA, publications index including Generative Artificial Intelligence for Government – [official indexed description; direct fetch blocked]; safe-use scope only.

