01
Pipeline Architecture
Every agent action traverses a 5-stage governed pipeline. Stages 1–3 are deterministic compute. Stage 4 (Deliberation) is conditional and only activates on high-risk actions. Stage 5 assembles the final verdict and writes the tamper-evident audit record.
STAGE 01
Intent Validation
Routine path
10-point structural validation: prompt injection detection (17 pattern signatures + imperative heuristic), timestamp drift checks (±60s), financial exposure quantification, regulatory domain mapping, data category classification (PII/PCI/PHI), reversibility assessment.
Always runs
→
STAGE 02
Context Enrichment
Routine path
Parallel async fetch across the systems connected for your organization (today, your own context endpoint) and GaaS’s own records. 6-step sub-pipeline: relevance analysis, source selection, parallel fetch, contradiction detection, gap analysis, confidence scoring. Results in a 0.0–1.0 confidence score.
Always runs
→
STAGE 03
Policy Evaluation
Routine path
59 built-in policies (plus your own) across the 4-tier hierarchy, evaluated against the enriched intent. 6-dimension weighted risk scoring. Scope pre-indexing for sub-millisecond policy filtering.
Always runs
→
STAGE 04
Deliberation
~40–60s
Multi-agent debate panel of up to 5 reviewers, powered by Claude Opus 5.5 (effort medium) with a 1M-token context window. Consensus-based verdict with veto authority. Only triggered when composite risk score ≥ 0.65.
Conditional
→
STAGE 05
Decision & Audit
Routine path · 24ms median
Final verdict assembly (APPROVE / APPROVE_MODIFIED / ESCALATE / BLOCK), tamper-evident audit record with SHA-256 hash chain linkage, signed (ECDSA P-256), timestamped daily in Bitcoin (OpenTimestamps); every live decision carries a governance proof token anyone can check.
Always runs
Routine decisions resolve in Stages 1–3 and 5 in a median of 24ms end to end (p95 36ms, measured in production; per-stage times are not measured separately). Only high-risk, novel or conflicting actions enter Deliberation, which takes about 40–60s end to end. Actions a human must decide are escalated for review.
| Stage |
Name |
Latency |
Trigger |
| 1 |
Intent Validation |
Routine path |
Every action |
| 2 |
Context Enrichment |
Routine path |
Every action |
| 3 |
Policy Evaluation |
Routine path |
Every action |
| 4 |
Deliberation |
~40–60s |
Risk score ≥ 0.65 only |
| 5 |
Decision & Audit |
Routine path: 24ms median, p95 36ms |
Every action |
02
Policy Engine
59 built-in policies (Tier 1: 20, Tier 2: 36, Tier 3: 3), plus your own, organized into four layered tiers, evaluated in order. Every action passes through all applicable tiers. Risk scores from each tier are aggregated into a single 0.0–1.0 composite score that determines pipeline routing.
Non-disableable
Tier 1 — Universal
20 policies
The non-negotiable floor. Sensitive data protection, unverified channel detection, delegation depth limits, contradiction detection, irreversibility guards, rate anomaly detection, self-governance modification prevention, prompt injection blocking, session trust enforcement, behavioral anomaly blocking. Cannot be disabled by any organization. (Session trust and behavioral anomaly checks are live and enforcing since 2026-09-28.)
Regulatory
Tier 2 — Regulatory
36 policies
GDPR, EU AI Act (Articles 9, 10, 13, 14 and 15; high-risk obligations from 2 Dec 2027), data residency enforcement, consent compliance, right-to-explain requirements. Automatically activates based on the regulatory domains mapped during Stage 1 intent validation.
Configurable
Tier 3 — Organizational
3 policies, plus your own
Policies defined by your compliance team. Natural-language policy authoring — describe a rule in plain English and GaaS generates, smoke-tests, and deploys it to your policy scope. Overrides and extensions to Tier 2 where regulations permit.
Sandboxed
Tier 4 — Experimental
Built, not loaded
Emerging governance frameworks and research policies. Evaluated in sandboxed mode — verdicts are logged and observed but do not affect routing. Graduated to Tier 3 after validation thresholds are met. (The Tier 4 set is built but not loaded in the hosted service yet.)
AP2 Payments
AP2 Payment Governance
7 policies, within Tiers 1–2
Mandate validation, spend limits and velocity controls, PCI DSS compliance enforcement, PSD2 Strong Customer Authentication (SCA) for agent-initiated payments, AML velocity detection. Applied automatically to all actions classified as financial.
6-Dimension Risk Scoring
| Dimension |
Weight |
What it measures |
| Reversibility |
20% |
Can the action be undone? |
| Financial Exposure |
20% |
Cost if the action is wrong |
| Regulatory Density |
20% |
How many regulations apply? |
| Audience Impact |
15% |
Who is affected, and how many? |
| Context Confidence |
15% |
How certain is the enriched context? |
| Novelty |
10% |
How unprecedented is this action? |
03
Deliberation Engine
When an action clears the risk threshold, a structured multi-agent debate panel convenes. Up to five specialized reviewers with defined roles, weights, and authority levels (a sixth role, Precedent, is built but not enabled) reach a consensus verdict in up to three rounds of structured deliberation; round 2 runs only when reviewers conflict.
5
reviewers at most (Precedent built, not enabled)
3
debate rounds at most; round 2 only when reviewers conflict
55%
weighted consensus threshold
1M
token context window (Claude Opus 5.5)
Compliance Agent
Regulatory & legal compliance evaluation
Veto Authority
Ethics Agent
Ethical implications & fairness assessment
Conditional Veto
confidence > 0.8
Risk Agent
Operational & financial risk quantification
Standard
Domain Expert
Industry-specific context & precedent
Conditional Include
Precedent Agent
Historical decision pattern analysis (built, not enabled)
Conditional Include
Cost / Efficiency Agent
Operational cost & workflow impact
Conditional Include
Panel scaling: panels are capped at 2, 4 or 6 members for routine, elevated and critical urgency; with the Precedent reviewer not enabled, at most 5 reviewers sit. Time budgets: 150 s (180 s for critical cases). Anthropic ephemeral prompt caching (5-min TTL) keeps deliberation cost efficient.
3 retry attempts with exponential backoff
Per-provider circuit breaker: 5 failures → 60s cooldown
30s timeout per agent per round
Graceful fallback when LLM unavailable
04
Data Connectors & Enrichment
29 connector integrations built (27 listed on /connectors, plus Ring and a SIEM reader); none is switched on in the hosted service yet, and a category no system answers is reported as missing context and treated as risk. Since 2026-09-29 each organization can connect its own systems through one HTTPS context endpoint (/v1/context-endpoint; guide), which answers six categories: environment, account state, regulatory, organizational, identity and security. Connected sources are queried in parallel during Stage 2 with per-source resilience. Enrichment failures never block decisions — they inflate the risk score instead.
Communication, support & collaboration
Twilio, Slack, Microsoft Teams, Google Workspace, Zendesk, Asana
6 built
CRM, payments, HR & fulfillment
Salesforce, Stripe, Workday, ShipStation
4 built
Engineering, security & operations
GitHub, Okta, Datadog, PagerDuty, Jira, Vanta
6 built
Smart home & facilities
Alexa Smart Home, SmartThings, Google Nest, Honeywell Home, Philips Hue
5 built
Education
Canvas LMS, Clever
2 built
Fleet, field & energy
AEMP 2.0, Leaf Agriculture, Tesla Fleet, SolarEdge
4 built
Also built (not listed)
Ring; SIEM reader for Splunk, QRadar and Sentinel
2 built
Your own systems
One HTTPS context endpoint per organization, six categories
Live since 2026-09-29
Resilience of the built connectors (none switched on yet) and of your context endpoint
| Resilience Spec |
Value |
| Request timeout |
5s |
| Retry policy |
3 retries, exponential backoff (0.5s base) |
| Circuit breaker |
5 consecutive failures → OPEN for 60s |
| Enrichment cache |
Redis-backed, 30-min TTL, successful results only |
| Failure behavior |
Graceful degradation — failures inflate risk score, never block |
| Your context endpoint |
One attempt per decision within its time limit (1.5s by default, 200ms to 3s); no retries, redirects not followed; if it fails, the decision still happens and the missing categories count as risk |
Context Confidence Score: Each enriched context carries a 0.0–1.0 confidence score. Missing critical data applies a −0.20 penalty. Stale data (>24h) applies a −0.10 penalty. Low confidence directly inflates the composite risk score.
05
Audit & Cryptographic Integrity
Every governed action generates a 7-stage tamper-evident audit record, cryptographically signed and chained, and each organization's records are timestamped daily in Bitcoin. Any modification to any record breaks the entire downstream chain — tampering is detectable and provable.
1
Intent Declaration
Raw agent intent captured verbatim. Timestamp, agent ID, session ID, and source fingerprint recorded.
2
Validation Record
Stage 1 validation results: all 10 check outcomes, risk flags triggered, injection signatures matched.
3
Context Snapshot
Enriched context at time of decision: sources queried, data retrieved, confidence score, gaps identified.
4
Policy Evaluation Log
Each policy evaluated, its verdict, and the weighted contribution to the composite risk score.
5
Deliberation Transcript
Full agent panel transcript (if triggered): initial positions, cross-examination, final verdicts with confidence scores per agent.
6
Decision Record
Final verdict (APPROVE / APPROVE_MODIFIED / ESCALATE / BLOCK), rationale, and any modifications applied to the intent.
7
Outcome Record
Execution outcome reported by the agent, outcome verification status, and any deviation from the governed decision.
Hash Algorithm
SHA-256 per record, chained to prior hash
Proof Token Signature
ECDSA P-256 (live since 2026-09-29)
Default Retention
Kept until your organization sets a period (at least 180 days); the same on every plan
Export Formats
JSONL streaming, CSV, HTML (print-ready)
Non-Repudiation
Anyone can check a proof token, no API key: /v1/verify/proof/{token_id}; public key at /.well-known/gaas-audit-keys.json
Tamper Detection
Modification breaks entire downstream chain
Public Timestamps
Daily per organization, OpenTimestamps → Bitcoin. List with
GET /v1/audit/timestamps, download
/v1/audit/timestamps/{day}/ots, check with
ots (
guide)
06
Compliance Frameworks
GaaS ships with built-in coverage for 12 regulatory frameworks and 59 built-in governance policies. Compliance status is queryable via API and exportable for auditor review.
| Framework |
Status |
Coverage |
| EU AI Act |
5 policies |
Articles 9, 10, 13, 14 and 15. Compliance status API + reporting endpoint. Transparency obligations since 2 Aug 2026; high-risk obligations from 2 Dec 2027 (Annex III) and 2 Aug 2028 (Annex I). |
| GDPR |
4 policies |
Article 17 (erasure), Article 20 (portability), Article 22 (automated decision explanation), Article 28 (sub-processor management, 30-day advance notice). |
| SR 11-7 (Federal Reserve) |
Inventory |
Model inventory with validation status, decision statistics, and delegated authority limits: a record of every agent for your own model risk management. CSV/HTML export. Built for SR 11-7, which SR 26-2 (Federal Reserve, OCC and FDIC) replaced on 17 April 2026; SR 26-2 places generative and agentic AI models outside its scope. Not one of the 12 enforced frameworks (there is no model-risk policy). |
| PCI DSS |
Enforced |
Agent payment compliance via AP2 policy layer. Cardholder data detection and blocking. |
| PSD2 SCA |
Enforced |
Strong Customer Authentication enforcement for all agent-initiated payment actions. Part of the agent-payment (AP2) policies. |
| CCPA |
Enforced |
California consumer privacy rules, evaluated on agent actions that touch personal information. |
| NIST 800-53 |
Mapped |
AC, AU, IA, SC control family mapping for AI agent governance. Agent-specific interpretations of access control, audit, identification, and system protection controls. |
| FedRAMP |
Mapped |
FedRAMP Moderate baseline alignment. Control inheritance documentation for GaaS as a cloud governance provider. |
| CMMC |
Mapped |
Level 2 practice mapping for defense contractor agent governance. CUI handling controls and access enforcement. |
| NIST CSF |
Mapped |
Identify, Protect, Detect, Respond, Recover functions mapped to governance pipeline stages 1–5. |
| SOX |
Enforced |
Sarbanes-Oxley controls, evaluated on agent actions that touch financial records and reporting. |
| HIPAA |
Enforced |
PHI detection and blocking. Minimum necessary standard enforcement. BAA-ready audit trail with full chain of custody. |
| FERPA |
Enforced |
Education record protection. Role-based access enforcement for student PII. Directory vs. non-directory information classification. |
| TCPA + Florida FTSA |
Enforced |
Consent and contact rules for agent-initiated calls and texts: 6 TCPA policies plus Florida's Telephone Solicitation Act. |
07
Security
Defense-in-depth from the API boundary to the audit record. Cryptographic integrity and multi-layer injection detection are non-disableable Tier 1 controls; behavioral profiling and session trust decay run and enforce.
Prompt Injection Detection
17 regex pattern signatures + imperative heuristic. Evaluated at Stage 1 (reject before enrichment) and re-scanned at Stage 3 against enriched context.
Behavioral Anomaly Detection
Live and enforcing since 2026-09-28. An action type or sensitivity counts as unusual only if it is new, or under 2% of the agent’s last 200 live decisions (with at least 15 on record); blocked attempts are never learnt.
Session Trust Decay
Per-agent floating trust budget (1.0 → 0.10). Decays with risky decisions. Session blocked when budget depleted (≤0.10). Live and enforcing since 2026-09-28; a person’s approval restores trust.
Rate Anomaly Detection
Flags actions exceeding 3× baseline frequency within a 1-hour sliding window. Automatic escalation at 5× baseline.
Self-Governance Prevention
Agents cannot modify their own governance policies. Tier 1 non-disableable policy. Any attempt is auto-blocked, and the block reaches your SIEM through your organization’s signed webhooks.
SIEM Integration
Through signed per-organization webhooks and a published guide for Splunk, Microsoft Sentinel and QRadar; the old global CEF push was retired. Three Sigma detection rules ship as documentation.
API Key Management
gsk_ prefixed keys, SHA-256 hashed at rest, 90-day maximum lifetime, atomic rotation endpoint with zero-downtime key swap.
Rate Limiting
Sliding window: 600 req/min global. Tiered per-key limits: 120/20/10/5 by endpoint class. PostgreSQL-backed for distributed deployments.
08
API & Integration
167 REST paths under /v1 in the published OpenAPI spec (v0.2.17). Full OpenAPI spec, idempotent submissions, bulk processing, and field-level response filtering.
Idempotency: Idempotency-Key header with 24h TTL and request hash deduplication
Bulk submission: Up to 50 intents per request, 10 concurrent pipeline executions
Field filtering: Sparse responses via ?fields=verdict,risk_score
ETag caching: SHA-256 based, per-endpoint cache control headers
Webhook delivery: HMAC-SHA256 signed payloads with retry logic
A2A v1.0: JSON-RPC endpoint at POST /a2a/v1 (header A2A-Version: 1.0), Agent Card at /.well-known/agent-card.json, API-key sign-in (cross-org federation is built, not yet enabled)
SDKs
Python (async, Pydantic v2)
TypeScript (full OpenAPI types)
Java 17+ (maven.gaas.is)
gaas-langchain
gaas-langchain provides native LangChain integration: govern_tools() for tool-use wrapping, @govern_node decorator for LangGraph nodes, and GaaSCallbackHandler for automatic pipeline instrumentation.
09
Infrastructure & Availability
Production-hardened stack with full observability, load-tested against sustained and spike traffic profiles, and validated across three Python runtime versions.
| Component |
Specification |
| Runtime |
Python 3.11 / 3.12 / 3.13 |
| Database |
PostgreSQL 15 with read replicas |
| Cache |
Redis (enrichment results, agent profiles) |
| Observability |
Prometheus + OpenTelemetry |
| Logging |
Structured (structlog) with request ID correlation across all 5 stages |
| Load Testing |
k6 profiles — sustained, spike, and soak |
| CI Matrix |
Python 3.11, 3.12, 3.13 — ruff + mypy + pytest |
| Test Coverage |
4,300+ automated tests (unit, integration, API and E2E) plus 479 dashboard tests, run on every change in the CI matrix above |
| Tier |
Uptime SLA |
Response SLA |
Credit on Breach |
| Starter |
99.5% |
< 24h email |
10% monthly credit |
| Growth |
99.9% |
< 4h priority |
25% monthly credit |
| Enterprise |
99.99% |
< 1h dedicated |
100% monthly credit |
10
Cost Efficiency
Governance shouldn't become a bottleneck — financially or operationally. The pipeline is designed so the vast majority of actions never reach the expensive deliberation stage.
Routine Path
Stages 1–3 and 5
$0.002
Most actions
Deliberation Path
Stage 4 triggered
$0.05
High-risk actions
Escalation Path
Human review routed
$0.25
Human-review actions
Per decision beyond your plan’s included actions, at the published rates; a decision is billed once, in its highest class. See pricing.
Deliberation triggers at a ≥0.65 risk score, so routine actions never reach it. Prompt caching (Anthropic ephemeral, 5-min TTL) further reduces deliberation cost for repeated context patterns.
11
AI Model & Protocol Stack
GaaS runs on the latest best-performing prime foundation model. Today that is Claude Opus 5.5 (since 25 September 2026), with a 1-million-token context window. As better models emerge, GaaS upgrades automatically so your governance layer never falls behind the frontier.
Claude
Opus 5.5
prime foundation model
Always
latest frontier model
Included
deliberation inference in every plan
Always on the frontier: GaaS tracks frontier model performance and upgrades automatically. You don't configure model versions — you govern agents, and GaaS handles the rest.
Framework plugins & MCP
Governance where your agent runs, not just on the wire.
Blocked before it happens. Drop-in plugins for LangChain, CrewAI, the OpenAI Agents SDK, Pydantic AI, the Vercel AI SDK and Microsoft Agent Framework (AutoGen’s successor) wrap the tools you give your agent. Each call is checked before it runs, and a blocked action never runs.
Any tool, not only MCP tools. The plugins govern whatever you give the agent: your own functions, internal APIs and MCP tools alike.
Native MCP server. Any MCP-capable agent can ask GaaS before it acts, get a verdict, and look up the full audit record. There’s no custom integration code: npm install @governancehq/mcp.
Proof, not promises. Every live decision carries a signed proof token that anyone can check, no key needed.
Schema-compatible. Governance verdicts are returned as structured MCP tool results — no response parsing required by the agent
Agent-to-Agent Protocol (A2A)
A2A v1.0: GaaS is an A2A v1.0 agent, built on the official A2A SDK and tested on every change with the official A2A conformance suite (TCK): every MUST-level requirement it can test passes — SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask and SubscribeToTask over JSON-RPC at
/a2a/v1, skill discovery via
/.well-known/agent-card.json; verdicts become task states (approve → completed, block → rejected, escalate → input-required). Push notifications are not available yet, and an API key is the only sign-in. A2A, created by Google and now hosted by the Linux Foundation, reached v1.0 in 2026.
A2A guide
Governance proxy layer: A2A messages sent to GaaS’s A2A endpoint go through policy evaluation before task handoff; forwarding them on to the target agent is built, not yet switched on
Agent Trust Registry: A2A agents are registered with GaaS, each with a trust score computed from its interaction history (/v1/agent-registry)
Cross-org federation: built, not yet enabled. Delegation chains are audited end to end today: every record in a chain is hash-verified (/v1/audit/{audit_id}/delegation-chain)
How GaaS governs A2A, in detail
MCP gives agents tools. A2A gives agents colleagues. GaaS gives agents accountability. Both protocols are governed by the same 5-stage pipeline — same audit trail, same compliance posture, same cryptographic integrity.