White paper · Edition 5

The Context Dividend

How externalizing governance returns context to your AI agents, measured against the live system.

H2Om Technologies · Edition 5 · September 2026 · Updated

Measured against the live system, a self-governing agent on Claude Opus 5.5 spends about 14,700 tokens of context per routine governed action. With GaaS it spends 0–74 tokens through a framework plugin, or about 2,470 through MCP: an 83–100% reduction.

context-dividend · edition 5
Opus 5.5

Tokens in the agent's context, per governed action

Self-governed~14,700
GaaS MCP server~2,470
GaaS plugin0–74
Reduction
83–100%
Returned at 500,000 actions a month
6.1–7.4B tokens
Self-governed cost at that volume
~$73,000 / month
Decision latency, no deliberation
24 ms median
MEASURED · MODELED
Updated 2026-09-29

For an agent on Claude Opus 5.5, from the paper below. Cost at Opus 5.5 list prices, without prompt caching.

Contents
  1. How to read this edition
  2. What changed since Edition 4
  3. What Edition 4 changed from v3
  4. 1. Executive Summary
  5. 2. The Context Window Problem
  6. 3. The Governance Tax, Measured
  7. 4. The Context Dividend: What GaaS Returns
  8. 5. What Agents Do With Reclaimed Context
  9. 6. Financial Impact
  10. 7. Scaling: Governance Cost Is Flat for the Agent
  11. 8. The Liability Dividend
  12. 9. The Compliance Dividend
  13. 10. The Intelligence Dividend
  14. 11. The Deliberation Dividend
  15. 12. The Trust Dividend
  16. 13. The Visibility Dividend
  17. 14. The Learning Dividend
  18. 15. Who Benefits
  19. 16. The Compliance Imperative
  20. 17. Separation of Concerns
  21. 18. Conclusion
  22. Appendix B — Method

Download the PDF

How to read this edition

Every number in this paper carries a tag that points to its entry in the evidence register (Appendix A of the full paper), which records the number's value, its type, its source and the date it was taken. The register includes counts of GaaS's own production traffic, so it is kept internally and is not part of the web or PDF edition. The tag tells you which kind of evidence each number rests on.

TagTypeMeaning
[M#]MEASUREDRead from the production system or its code as deployed, or counted exactly with Anthropic's token counter
[T#]MEASURED (token count)An exact token count of a real GaaS artefact, per Claude model
[D#]MODELEDAn estimate calculated from measured parts; the assumptions are stated next to it
[X#]EXTERNALA fact published by someone else (model prices, the law), cited to its source

"Production" means the hosted service at api.gaas.is, as deployed on 2026-09-29 [M1]. Where a feature exists in the code but is not switched on in that service, this paper says so.

Token counts differ between Claude models: models from Opus 4.7 onward, including Opus 5.5 and Sonnet 5, use a tokenizer that produces about 30% more tokens for the same text [X2]. Every token figure is therefore given for Opus 5.5 (the headline model; identical counts on Sonnet 5) and Opus 4.6 (identical counts on Haiku 4.5).

What changed since Edition 4

Edition 4 (2026-09-25) named the features that were built but not running. Between 2026-09-25 and 2026-09-29 they were switched on or replaced, and each was checked against production [M1]:

  • Signed Governance Proof Tokens are issued for every live decision, and audit records are signed (ECDSA P-256). Anyone holding a token can check it, without an API key, at GET /v1/verify/proof/{token_id}; the public key is published [M12].
  • Plain-English explanations are on, generated after the verdict so they add no latency [M15].
  • Behavioural anomaly baselines and session trust decay run and enforce, with their state in GaaS's own database. The scoring was rebuilt so that an agent's normal mix of actions is never flagged; a person's approval restores an agent's trust [M15, M27].
  • SIEM integration is through per-organization webhooks and a guide; the single global push was retired [M15, M21].
  • Public timestamps replace "ledger anchoring". Once a day each organization's audit records are stamped into Bitcoin through the free OpenTimestamps service. The old anchoring had never anchored anything outside the process and was removed [M15].
  • Connectors no longer simulate. The hosted service stopped filling context with simulated data; a category with no connected system is reported as missing context and treated as risk [M14].
  • Customers can connect their own systems. Since 2026-09-29 an organization can give GaaS its own HTTPS context endpoint, which answers, during every decision, the facts the policies check [M14].
  • The audit chain verifies. The older records Edition 4 reported as failing were an artefact of re-reading them through today's data model; every production record matches its stored hash, and verification now reports only the caller's own records [M13].
  • A person can approve a blocked action: from the block email or, as an operator or admin signed in with two-factor, in the dashboard. The agent's retry of the same action within 24 hours is then approved once, on the record [M27].
  • Framework plugins fail closed: if GaaS gives no decision, the tool does not run, unless the customer opts out [M19].
  • One figure moved: the MCP server's cost is ~2,470 tokens per action (Edition 4: 2,350), an 83.2% reduction (84.0%). Live verdicts now carry a signed proof token; the MCP server shows it by its ID only, so the model reads about 100 more tokens than before rather than about 480 [T1, T3, D5]. The framework plugins' figures are unchanged.

What Edition 4 changed from v3

v3 (February–March 2026) was written before its figures were measured. Edition 4 replaced them.

  • The token model is rebuilt from measurements. v3 estimated 26,000–73,000 tokens per governance cycle from industry reports. This edition sizes GaaS's own artefacts with Anthropic's token counter: about 14,700 tokens per routine governed action for a self-governing agent on Opus 5.5, and 23,400–40,600 when the action needs deliberation [D3, D4].
  • What GaaS costs the agent is measured, per integration. v3 said 900–2,000 tokens. The framework plugins cost 0 tokens on an approved action and 74 on a blocked one; the MCP server costs about 2,350 [D5]. The reduction is 84–100%, not 92–97% [D6].
  • "Reclaim 30–60% of your agent's working memory" is now conditional. That share is reached only in long sessions: after about 14 and 28 routine governed actions on a 200K-token model, or 50 and 102 on a 1M-token model [D7].
  • Product facts corrected to production: 59 built-in policies, not 33 [M4]; a deliberation panel of at most 5, not 6 [M6]; deliberation time limits of 150 seconds (180 for critical cases), not 200 ms to 10 s [M7]; 29 connector integrations in code (18 in v3; gaas.is lists 27), none switched on in the hosted service [M14]; 9 published SDK packages, and no Java SDK [M19]; the plan prices [M20].
  • Features described as live that are built but not running in the hosted service: signed Governance Proof Tokens [M12], third-party context connectors [M14], behavioural anomaly baselines and session trust decay [M15], SIEM push, ledger anchoring and plain-English explanations [M15].
  • There is no "Security SKU." The security features are part of every plan [M20].
  • EU AI Act facts corrected. Maximum fines are €35M or 7% of turnover for prohibited practices, and €15M or 3% for operator obligations [X4], not €30M or 6%. The high-risk obligations now apply from 2 December 2027, not 2 August 2026 [X5].
  • Fixed during this review, all on 2026-09-25:
    • A customer's own agent was refused (HTTP 422) for any non-public action, and the framework plugins then let the action run ungoverned. Fixed in production [M17].
    • Deliberation did not complete: its model calls timed out or were cut off, so every case escalated [M8, M9]. A fix moved every Claude call to Opus 5.5 and fixed the deliberation path; it has completed in every lab, staging and production check since [M2, M9, M26].
    • The SDKs gave up after 5–30 s, shorter than a deliberation. New SDK releases wait up to 240 s [M19].
  • Refreshed the same day. The figures first published at 08:36 PT used deliberation answers measured on the pre-fix code, most of them cut off. This version re-measures them on the working code, which raises the self-governed figures (≈12,800 → ≈14,700 tokens per action).
  • The cost model uses current Claude prices. Opus 4.6 costs $5 / $25 per million input / output tokens, not $15 / $75 [X1].

1. Executive Summary

Every token an AI agent spends on governance is a token it cannot spend on its task. When an agent governs itself, it carries its rules, its safety checks, the context it looked up, its reasoning about each consequential action, and a record of that reasoning, all inside its own context window.

GaaS moves that work out of the agent. The agent declares what it intends to do and receives a verdict. The difference is the Context Dividend.

Measured on the live system, for an agent running on Claude Opus 5.5 [D3, D5, D6]:

Per governed actionSelf-governedGaaS framework pluginGaaS MCP server
Tokens in the agent's context~14,7000 (approve) · 74 (block)~2,470
Reduction—99.5–100%83%
  • For a deliberated action (legal but risky, novel, or high-stakes), self-governance costs 23,400–40,600 tokens; the GaaS cost to the agent does not change [D4].
  • At 500,000 governed actions a month, externalizing governance returns about 7.4 billion tokens (framework plugin) or 6.1 billion (MCP). At Opus 5.5 list prices, self-governance would cost about $73,000 a month in tokens, or about $56,000 with prompt caching [D8].
  • How much of the window this is depends on the session. The standing overhead of self-governance is 0.9% of a 1M-token window and 3.0% of a 200K window. It grows with every governed action: self-governance fills 30% of a 200K window after about 14 routine governed actions in one session, and 60% after about 28 [D7].

The Context Dividend is the most measurable return, but not the only one. GaaS also enforces 59 built-in policies before an action runs [M4], writes a hash-chained audit record for every decision [M13], signs it, and issues a proof token anyone can check for every live decision [M12], and routes the hard calls to a deliberation panel or a human. Section 8 onward covers these returns, and says plainly which of them are live in the hosted service today.

2. The Context Window Problem

2.1 Context is finite

A context window is an agent's working memory. It must hold the task, the conversation, tool outputs, reasoning, and any governance logic the agent carries. Research on long-context performance reports that model accuracy degrades as input length grows, and that it degrades unevenly [X6]. Larger windows raise the ceiling; they do not remove the trade-off.

2.2 Today's prices and windows

ModelInput $/MTokOutput $/MTokContext windowEarliest retirement
Claude Opus 5.5$4.00$20.001Mnot before 2027-09-22
Claude Sonnet 5$2.00$10.001Mnot before 2027-06-30
Claude Haiku 4.5$1.00$5.00200Knot before 2026-10-15
Claude Opus 4.6$5.00$25.001Mnot before 2027-02-05

Source: Anthropic pricing, models overview and model deprecations pages, fetched 2026-09-24 [X1, X3]. Opus 4.6 is listed because GaaS's deliberation used it until 2026-09-25, and agents still run on it; GaaS now uses Opus 5.5 [M2]. Other providers are not re-priced in this edition.

2.3 Agents spend context fast

An agentic workflow makes many model calls, and each call re-sends the accumulated context. Any governance content the agent carries is paid for on every call, and any governance record it writes stays in the window for the rest of the session.

3. The Governance Tax, Measured

To size what self-governance costs, this edition asks a concrete question: what would an agent have to carry and write to do, by itself, the governance GaaS does? Each component below is an exact token count of a real GaaS artefact [T#]. Treating it as the agent's cost is the modeling assumption [D#].

3.1 Standing overhead: sent with every request

ComponentWhat the agent carriesOpus 5.5Opus 4.6 / Haiku 4.5
S1 · Policy instructionsGaaS's 59 built-in policies, written out as instructions [T5]7,7455,236
S2 · Safety checksGaaS's 10 intent checks and 17 prompt-injection patterns [T6]1,126808
Standing subtotal[D1]8,8716,044

3.2 Per governed action: added to the context each time

ComponentWhat the agent reads or writesOpus 5.5Opus 4.6 / Haiku 4.5
S3 · Context lookupsThe context payload GaaS assembled for a decision, measured in Edition 4 on the simulated sources the hosted service then used [T7]. It stands as a lower bound for what an agent would read to look the context up itself: real system data would be larger365273
S4 · Statement of the actionWhat the agent intends to do, in structured form [T8]307208
S5 · Governance reasoningOne independent risk assessment: the mean output of a round-1 GaaS deliberation call, including the model's thinking [T10]2,5421,751
S6 · Audit recordThe median production audit record [M11, T9]2,6221,859
Per-action subtotal[D2]5,8364,091

S5 counts the model's thinking as well as its written answer, because both are generated and billed; how much earlier thinking an agent keeps in its window varies by model and harness. The Opus 4.6 / Haiku 4.5 value of S5 is converted from the Opus 5.5 measurement with the tokenizer ratio measured on GaaS's own artefacts (0.689, D10), since deliberation no longer runs on Opus 4.6.

3.3 Totals

Opus 5.5Opus 4.6 / Haiku 4.5
Routine governed action (standing + per action) [D3]14,70710,135
Deliberated action: S5 is replaced by a whole panel's prompts and answers, measured over 5 deliberations of 2–4 reviewers [D4]23,445–40,64516,153–28,000

v3 estimated 26,000–73,000 tokens per cycle. The measured components are smaller: GaaS's policies, checks and verdicts are compact structured data. The earlier figure is withdrawn.

4. The Context Dividend: What GaaS Returns

4.1 What GaaS costs the agent

The agent's cost depends on how it connects to GaaS.

  • Framework plugins (LangChain, CrewAI, OpenAI Agents, Pydantic AI, Vercel AI) wrap the agent's tools. The plugin builds the intent in code, so the model writes nothing and reads nothing when the action is approved. When an action is blocked, the model sees one short error line [T4]. The signed proof token stays in the plugin's error object and the API response; the model never reads it.
  • The MCP server gives the model four tools. The model writes the intent itself, as tool-call arguments, and reads the verdict back [T2, T3]. From version 0.1.4 it shows a live verdict's proof token by its ID and where to check it, not the whole signed token. The tool definitions ride along in every request [T1].
IntegrationStanding (every request)Per actionTotalReduction vs. self-governed
Framework plugin00 on approve · 74 on block (Opus 5.5) / 60 (Opus 4.6)0–7499.5–100% [D6]
MCP server, Opus 5.51,4091,065 (198 written + 867 read)2,47483.2% [D6]
MCP server, Opus 4.61,042849 (146 written + 703 read)1,89181.3% [D6]

If the agent has no other tools, Anthropic's tool-use system prompt adds 286 tokens (Opus 5.5) or 497 (Opus 4.6) to the MCP standing cost [X1].

4.2 The verdict

A live verdict is about 2.2 KB of JSON: about 1,040 tokens on Opus 5.5 and 770 on Opus 4.6 as compact JSON, of which the signed proof token is about 420 [T3]. Before proof tokens were issued a verdict was about 1.4 KB (median 1,390 bytes) [M11]. The MCP server returns it pretty-printed with the token shown by its ID, at 867 and 703 tokens [T3]. A verdict carries the verdict, the reason, any modifications or conditions, the blocking policies and suggested alternatives when blocked, a six-dimension risk assessment, governance metadata, an audit reference and the proof token. The full audit record stays in GaaS [M11].

4.3 What GaaS does outside the agent's context

StageWhat it doesStatus in the hosted service (2026-09-29)
1 · Intent validation10 checks, including 17 prompt-injection patterns [M5]Live. Since 2026-09-25, a customer's own agent is validated and passed on to policy evaluation [M17]
2 · Context enrichmentLooks up context by category and flags contradictionsLive, without simulation. No vendor integration is switched on, so a category no system answers reports missing context, which counts as risk; since 2026-09-29 an organization can answer six categories from its own context endpoint; history and behaviour come from GaaS's own database [M14, M15]
3 · Policy evaluation59 built-in policies plus customer policies; six-dimension risk score [M4]Live
4 · DeliberationA panel of up to 5 model-based reviewers for high-risk or conflicting cases [M6]Live since 2026-09-25 on Claude Opus 5.5. It did not complete before that (model calls timed out or were cut off, so every case escalated) [M8, M9]
5 · Decision and auditVerdict, hash-chained audit record, proof tokenLive. Records are signed and every live decision carries a proof token anyone can check [M12]; plain-English explanations are on, generated after the verdict [M15]

Measured latency without deliberation, on the smoke-test account's 831 decisions in the hosted service: median 24 ms, 95th percentile 36 ms [M16]. A decision that goes to deliberation takes about 40–60 s end to end (staging and production checks, 2026-09-25) [M26].

5. What Agents Do With Reclaimed Context

5.1 When the dividend matters most

The standing overhead of self-governance is small against a large window, but the per-action cost accumulates for the whole session:

Agent model (window)Standing overhead30% of window reached after60% reached after
Opus 5.5 / Sonnet 5 (1M)0.9%~50 governed actions~102 governed actions
Haiku 4.5 (200K)3.0%~14 governed actions~28 governed actions

[D7] Routine actions, in one session. Deliberated actions reach these points sooner.

So the dividend is largest for long-running agents, high-action sessions, and smaller-window models, which is where v3's "30–60%" applies.

5.2 Multi-agent systems

Every agent in a multi-agent system would carry its own copy of the standing overhead (8,871 tokens each on Opus 5.5) [D1]. With GaaS, each agent carries none (plugins) or 1,409 (MCP). Anthropic has reported that its multi-agent research system outperformed a single agent while using many times more tokens [X7]. Any fixed per-agent overhead multiplies with the agent count.

5.3 Less context rot

Removing governance content keeps the agent further from the lengths at which long-context performance degrades [X6].

6. Financial Impact

6.1 Token cost at 500,000 governed actions a month [D8]

Assumptions: routine actions (no deliberation), one model request per governed action, list prices, the standing overhead either uncached or read from the prompt cache.

Agent modelSelf-governed tokens / monthReturned (plugin)Returned (MCP)Self-governed cost, uncached…with prompt cachingMCP cost to the agent
Opus 5.57.35 B7.35 B6.12 B$73,182$56,327$6,532
Sonnet 57.35 B7.35 B6.12 B$36,591$28,607$3,266
Opus 4.65.07 B5.07 B4.12 B$63,518$49,918$6,188
Haiku 4.55.07 B5.07 B4.12 B$12,704$9,984$1,238

The framework plugins cost the agent no tokens on approved actions. The GaaS subscription is not deducted (see 6.3). Deliberated actions would raise the self-governed figures (Section 3.3).

6.2 What deliberation costs GaaS

Deliberation runs on GaaS's own model account, not the customer's. Measured on the production deliberation code (Claude Opus 5.5, effort medium) over 5 lab deliberations of 2–4 reviewers, 32 model calls, every one completed [M9, D9]:

LowestMeanHighest
Cost per deliberation$0.13$0.24$0.31
Tokens per deliberation (prompts + answers)11,280—28,480

6.3 GaaS plans

Free ($0, 1,000 decisions a month), Developer ($99), Starter ($500), Growth ($2,500) and Enterprise ($10,000+) a month. Overage is $0.002 per routine decision, $0.05 per deliberation and $0.25 per escalation [M20].

7. Scaling: Governance Cost Is Flat for the Agent

Governance complexitySelf-governed, Opus 5.5GaaS pluginGaaS MCP, Opus 5.5
Routine (policies pass or block)14,7070–742,474
Deliberated (measured, 2–4 reviewers)23,445–40,6450–74~2,530

[D3, D4, D5] A deliberated verdict is about the same size as a routine one: 829 tokens pretty-printed on Opus 5.5 in Edition 4, about 925 with the proof token shown by its ID [T3]. The panel's transcript stays in GaaS.

8. The Liability Dividend

What it means. When an agent's decision is questioned, you need a record of what governance was applied, when, and with what result, in a form that cannot be quietly edited afterwards.

What is live. Every decision writes an audit record that includes the intent, the policy evaluations, the risk assessment and the verdict. Each record is hash-chained to the previous one (SHA-256), so an edit breaks the chain. Records can be exported (GET /v1/audit/export, /v1/audit/export/stream), and the chain can be checked (GET /v1/audit/chain/verify) [M13, M18]. Every production record matches its stored hash. Edition 4 reported a set of older records as failing; they had been re-read through a data model that had since gained fields, and now verify as stored. Verification reports only the caller's own records [M13].

Live since 2026-09-29. Audit records are signed (ECDSA P-256), and every live decision carries a Governance Proof Token: a signed receipt of the verdict, the agent, the organization and the audit record's hash. Anyone holding a token can check it at GET /v1/verify/proof/{token_id} without an API key, against the public key GaaS publishes at /.well-known/gaas-audit-keys.json [M12, M18].

Public timestamps. Once a day, each organization's audit records for the previous day are combined into one digest and stamped through the free, public OpenTimestamps service into Bitcoin. The customer downloads the proof and checks it with the standard ots tool, without trusting GaaS; no other organization's data is in it. This replaced "ledger anchoring", which had never anchored anything outside the process. The daily job was deployed on 2026-09-29. Its first run, at 03:02 UTC that day, stamped the previous seven days, and those proofs are confirmed in Bitcoin, from block 969093 [M15]. Co-signing with a customer key is built and switched off.

Regulatory exposure. Under the EU AI Act, fines reach €35M or 7% of worldwide annual turnover for prohibited practices, €15M or 3% for breaches of operator obligations, and €7.5M or 1% for supplying misleading information [X4].

9. The Compliance Dividend

GaaS enforces policy before an action runs, and records the result.

59 built-in policies are loaded in production: 20 at Tier 1, 36 at Tier 2 and 3 at Tier 3 [M4]. They cover:

  • EU AI Act Articles 9, 10, 13, 14 and 15
  • Privacy: HIPAA (2 policies), PCI, GDPR, CCPA and FERPA
  • Finance: SOX, and agent payments (AP2): mandate validity, conditions and spend limits, PSD2 strong authentication, and anti-money-laundering velocity (7)
  • Communications: TCPA (6) and Florida's FTSA
  • Security frameworks: NIST CSF (5), NIST 800-53 (5), FedRAMP (5) and CMMC (4)

Tier 1 covers universal safeguards, among them prompt-injection blocking, minor-data protection, irreversible-action confirmation and delegation limits [M4].

Other compliance capabilities that exist in production:

  • Policy packs, installable with one call (POST /v1/policy-registry/{pack_id}/install). All nine resolve every policy they list, after two entries with no policy behind them were dropped. The underlying policies are loaded for every organization regardless [M4, M24]
  • Plain-language policy authoring (/v1/policy-authoring/generate) [M18]
  • EU AI Act status and report endpoints (/v1/compliance/eu-ai-act, /report) [M18]
  • A model inventory of every agent for model risk management (/v1/model-inventory, with export). It was built for SR 11-7, which SR 26-2 rescinded on 2026-04-17; the new guidance places generative and agentic AI models outside its scope [M4, M18]
  • Policy epochs, recording which policy set was active for each decision (/v1/policy/epoch, /history) [M18]

EU AI Act timing. Regulation (EU) 2026/1744 (the Digital Omnibus on AI), in force since 27 July 2026, moved the high-risk obligations to 2 December 2027 for stand-alone (Annex III) systems and 2 August 2028 for systems embedded in regulated products (Annex I). Transparency obligations have applied since 2 August 2026 [X5].

10. The Intelligence Dividend

The idea. Governance decisions should be made against reality, not against whatever the agent claims. GaaS looks up context (identity, account state, environment, security signals), compares it with what the agent declared, and treats missing context as a risk.

What exists. GaaS's code can build 29 third-party integrations: Twilio, Slack, Microsoft Teams, Google Workspace, Zendesk, Asana, Salesforce, Stripe, Workday, ShipStation, GitHub, Okta, Datadog, PagerDuty, Jira, Vanta, Alexa, SmartThings, Google Nest, Honeywell Home, Philips Hue, Ring, Canvas LMS, Clever, AEMP 2.0, Leaf Agriculture, Tesla Fleet, SolarEdge, and a SIEM reader for Splunk, QRadar and Sentinel [M14].

What is live in the hosted service. None of these integrations is switched on, and the hosted service no longer simulates their data (since 2026-09-28). A category with no connected system is reported as missing context and treated as risk; history and behaviour come from GaaS's own records. On past production decisions, this change altered 32 of 325 verdicts and sent more to deliberation; none became a block.

Connecting your own systems. Since 2026-09-29 an organization can give GaaS one HTTPS context endpoint of its own, set up in the dashboard or the API. During every decision GaaS sends it a signed request describing the action (never its free-text content) and reads back the facts the policies check, in six categories: environment, account state, regulatory, organizational, identity and security. The endpoint's secrets are encrypted with AWS KMS. If it fails or is slow, the decision still happens and the missing facts count as risk. In production, an endpoint answering that the channel was encrypted let an action through that pol_t1_001 otherwise blocks [M14].

11. The Deliberation Dividend

The design. Rules settle the clear cases. For risky, novel or conflicting ones, GaaS convenes a panel of model-based reviewers [M6]:

ReviewerWeightVeto
Compliance1.0Absolute: a block is final
Ethics0.9A block with confidence above 0.8 is final
Risk0.8—
Domain expert0.7—
Cost / efficiency0.5—

Compliance and Risk always sit. Domain expert, Cost and Ethics join when the action warrants it. Panels are capped at 2, 4 or 6 members for routine, elevated and critical urgency, but only five reviewers can be selected: the Precedent reviewer is not enabled [M6]. Round 1 is an independent assessment, Round 2 a cross-examination (only when reviewers conflict), and Round 3 a final vote. A unanimous, high-confidence Round 1 ends early. Dissent is recorded. Time budgets are 150 seconds, and 180 for critical cases [M7].

Status in the hosted service. Deliberation runs on Claude Opus 5.5 at effort medium, and completes, since 2026-09-25 [M2, M26]. Before that it did not. When production last saw deliberation traffic (22–27 February 2026), every Opus 4.6 call timed out and the circuit breaker then refused further calls [M8]. Lab runs of that code showed routine deliberations still timing out and a quarter of answers cut off at a 1,024-token reply limit [M9]. The fix read replies by type, raised the reply limit to 8,192 tokens and the time budgets to 150–180 s, and moved to Opus 5.5. Since then every panel vote has completed: 5 of 5 lab deliberations, 3 of 3 on staging (22 votes) and a production test-mode check (8 votes), taking about 40–60 s end to end [M9, M26]. If a reviewer does fail, GaaS records an escalate vote, so the case goes to a human: the failure mode is conservative, not permissive.

In production traffic to date, about one live decision in four was routed to deliberation [M10].

12. The Trust Dividend

Live:

  • An Agent-to-Agent (A2A) v1.0 endpoint on the official A2A SDK (JSON-RPC at /a2a/v1), with a public agent card (8 skills); the official A2A conformance suite passes every MUST-level requirement it tests [M23]
  • AP2 payment mandates, transactions and spend summaries, governed by the 7 AP2 policies [M23]

Live since 2026-09-28:

  • Behavioural baselines. Each agent's recent live decisions form its baseline, kept in GaaS's database. An action type or sensitivity is unusual only if it is new, or under 2% of the agent's last 200 live decisions, with at least 15 on record; a large jump in amount counts too. Blocked attempts are never learnt, so a retry cannot teach itself in [M15].
  • Session trust decay. Blocks lower an agent's trust budget most (more for riskier ones), escalations and modified approvals less, and approvals slowly restore it. At the floor (0.10), the agent's actions are blocked. A person's approval restores the budget to 0.5, and an operator can reset it [M15, M27].

Built, not active:

  • Trust tiers (Registered, Verified, Certified) exist as a field on an agent profile, but no code moves an agent between tiers. Every customer agent is Registered [M22].
  • Customer co-signing of audit records is built and switched off.

13. The Visibility Dividend

Live in production:

  • A conversational dashboard, running on Claude Opus 5.5, for questions about governance activity [M3]
  • An escalation queue: pending escalations with their full context, statistics, review, reassignment and cancellation [M18]
  • Webhooks, signed with HMAC-SHA256 (X-GaaS-Signature), per organization, for decision.approved, decision.blocked, decision.escalated, decision.overridden, escalation.decided, escalation.cancelled, escalation.reassigned, escalation.timed_out and quota.exceeded. They are also how customers feed a SIEM (Splunk, Sentinel, QRadar), with a published guide [M21]
  • Block notifications and approvals. A live block emails the organization's contact address (at most once per agent per hour), and a person can approve it from the email or, signed in with two-factor, in the dashboard; an organization can require the dashboard [M27]
  • Decision metrics: a snapshot and history, and a decision stream filterable by verdict, risk, agent, action type, deliberation and mode [M18]

The single global SIEM push was retired in favour of the per-organization webhooks. Three Sigma detection rules ship as documentation (docs/siem/sigma/) [M15].

14. The Learning Dividend

The learning API is deployed: feedback, incident reports, calibration and its history, backtesting, knowledge-base updates, and learning metrics (/v1/learning/*) [M18]. This edition did not measure its effect on decision quality in production, and makes no claim about it.

15. Who Benefits

  • CTO / engineering. Framework plugins for five agent frameworks, plus a Python SDK, a TypeScript SDK and an MCP server: 9 published packages. There is no Java SDK yet [M19]. Shadow mode evaluates without enforcing. The plugins fail closed by default: if GaaS gives no decision, the tool does not run, unless the customer opts out; they can also wait for a person's approval of a block and retry once. Every SDK waits up to 240 s, long enough for a deliberated decision [M19]. Measured latency: median 24 ms, 95th percentile 36 ms without deliberation [M16]; about 40–60 s with it [M26].
  • Compliance / GRC. 59 enforcement policies across the frameworks in Section 9, audit export, EU AI Act reporting, a model inventory, signed audit records and proof tokens anyone can check, daily public timestamps, and audit retention the organization sets (Section 8).
  • Operators. An escalation queue, webhooks and decision metrics (Section 13).
  • Policy managers. Plain-language policy authoring and installable policy packs (Section 9).
  • Agencies serving SMEs. Organization-scoped policies and API keys.
  • CFO / board. The token economics in Section 6, against plans from $0 to $10,000+ a month [M20].
  • Legal. A hash-chained, signed, exportable audit record and a checkable proof token for every live decision (Section 8).

16. The Compliance Imperative

Self-governed reasoning lives in a context window and disappears with the session. Record-keeping, human oversight and continuous risk management need something that persists outside the agent. That requirement is architectural. It does not depend on window size, and it is the case for externalized governance even where the token savings are small.

The EU AI Act's transparency obligations apply now. Its high-risk obligations apply from 2 December 2027 (Annex III) and 2 August 2028 (Annex I) [X5].

17. Separation of Concerns

With governance externalized, policy changes, new frameworks and model upgrades roll out once, in GaaS, and no governed agent needs to change. The 2026-09-25 move of GaaS's own models to Claude Opus 5.5 (Sections 6.2 and 11) is an example: it changed nothing in any customer's agent.

18. Conclusion

Measured against the live system, a self-governing agent on Claude Opus 5.5 spends about 14,700 tokens of context per routine governed action, and 23,400–40,600 per deliberated one. With GaaS it spends 0–74 tokens through a framework plugin, or about 2,470 through MCP: an 83–100% reduction. At 500,000 actions a month, that returns 6.1–7.4 billion tokens, worth $56,000–$73,000 a month at Opus 5.5 list prices.

The dividend is real, and smaller than v3 claimed. The larger case is the rest of the portfolio: enforcement before action, a signed and publicly checkable record, and escalation to people. Edition 4 named the parts that were built but not running; Edition 5 reports them live. What remains off is named where it appears: third-party connectors, trust tiers and customer co-signing.

Appendix B — Method

  1. Production facts were read from the running system: ECS and ECR (what is deployed), the live API (/openapi.json, /v1/health/ready, /.well-known/agent-card.json), CloudWatch logs, and the smoke-test account's own decisions (read-only GETs).
  2. Production traffic was measured with one read-only SQL query in a one-off task (the api container only, inside a read-only transaction, returning aggregates only). The live service was unchanged: still one task, same revision, health 200.
  3. Token counts used Anthropic's count_tokens endpoint per model. Framing tokens were subtracted, and no text was generated.
  4. The pipeline was run locally with production's code (before and after the fix) and settings: connector mode production, no state store.
  5. Lab deliberations used 10 synthetic intents per model with no customer data, running the production deliberation code with a usage-recording provider: first on the pre-fix code ($0.58), then on the fixed code at the production setting ($1.93).
  6. Every derived figure comes from one calculation, so a new measurement changes every dependent figure consistently.
  7. Edition 5 re-counted only what changed (T1, T3), as the difference between the old and new artefact counted the same way, added to Edition 4's value. As a check, today's verdict with its new parts removed counted within 1% of Edition 4's figures (620 / 777 against 615 / 771). No other measured figure changed: the plugins' block message is the same text, and the self-governed components are GaaS's policies, checks and records, whose sizes this edition did not re-measure.
  8. Production checks for Edition 5 used the production smoke-test account (read-only GETs, and live test decisions of its own), CloudWatch, ECS and ECR, and one read-only SQL count in a one-off task.

About this paper: Edition 5, September 2026. It replaces Edition 4 (2026-09-25), which replaced v3 (February–March 2026). The next edition should re-measure after real connectors go live.

Contact: gaas.is/contact

Measure the dividend on your own agents.

The Free plan is $0 for 1,000 decisions a month. Start in shadow mode: GaaS evaluates every action without enforcing, so you can see its decisions before you switch enforcement on.