Enterprise AI agent compliance means proving that a platform’s controls, contracts, deployment architecture, and operating model fit your regulated workflow. For SOC 2, HIPAA, and bank reviews, request the current Type II report, BAA scope, subprocessor data flow, deployment boundary, immutable evidence logs, and a human approval gate before production changes.
In their 2024 guide for community banks, the Office of the Comptroller of the Currency (OCC), Federal Reserve and Federal Deposit Insurance Corporation say reliance on third parties reduces a bank’s direct operational control and may introduce new risks or increase existing ones. Banks should identify, assess, monitor and control third-party risk throughout the relationship lifecycle, with the level of review shaped by the relationship and risk profile (OCC, 2024).
What is enterprise AI agent compliance?
Enterprise AI agent compliance is the set of technical, contractual and governance controls that allow an artificial intelligence (AI) agent to use enterprise data and take workflow actions within an organization’s security, privacy, regulatory and change-management requirements.
The key procurement question is not whether a vendor displays a security badge. It is whether the exact service you will buy, configure and connect is covered by the vendor’s controls and contracts.
Three distinctions prevent common review errors:
- Service Organization Control 2 (SOC 2) Type II is evidence that an auditor assessed specified controls over a review period. It is not blanket approval for every product, integration, subprocessor or customer configuration.
- The Health Insurance Portability and Accountability Act (HIPAA) is not a vendor badge. If a vendor acts as a business associate handling protected health information (PHI), the covered entity needs satisfactory assurances in a written business associate agreement (BAA). The U.S. Department of Health and Human Services (HHS) describes the BAA as the contract or written arrangement that sets out the business associate’s obligations to safeguard PHI (HHS).
- Bank security review is broader than certifications. It can cover third-party risk, operational resilience, data access, subcontractors, monitoring and the bank’s ability to control the relationship.
Use this distinction to prevent procurement teams from treating different forms of evidence as interchangeable.
Start with a gating decision, not a vendor scorecard
A vendor should not reach detailed scoring if it cannot satisfy the legal, data or deployment conditions that are non-negotiable for your use case. This is consistent with the OCC’s third-party risk guidance, which treats identification, assessment, monitoring and control as activities that continue throughout the relationship lifecycle (OCC, 2024).
Use four gates before comparing features:
- Data gate: Will PHI, nonpublic personal information, payment data or regulated decision data enter prompts, retrieval systems, transcripts, tool calls or logs?
- Contract gate: Will the vendor sign the needed agreement, and does it cover the exact service, add-ons, channels and subprocessors?
- Architecture gate: Can data stay in the required region, private cloud, virtual private cloud (VPC), on-premises environment or network boundary?
- Operations gate: Can your team reconstruct each action and stop or approve changes before production promotion?
Treat any “no,” “not publicly documented” or “we can discuss that after signature” answer as a procurement event. The result should be a documented contract exception, a compensating control or rejection.
This approach also separates a vendor’s general readiness from your own configuration responsibilities. A platform may provide access controls and audit logs, while your team still determines which credentials, data sources, retention rules and production permissions are appropriate.
BAA and SOC 2 due diligence: what public documentation proves
Public documentation can show that a vendor has published certain artifacts or describes certain controls. It does not replace reviewing the underlying report, contract language or service scope.
Zendesk
Zendesk states that customers with Advanced Compliance can enter into a BAA or other healthcare agreement for accounts that may store PHI. Its documentation limits that coverage to expressly listed covered services and eligible add-ons. Early Access Programs, Marketplace applications and other third-party services are not covered. The listed HIPAA-enabled add-ons include AI Agents - Advanced (Zendesk Advanced Compliance).
Zendesk’s Trust Center states that it undergoes routine SOC 2 Type II audits and makes reports available on request under a nondisclosure agreement (Zendesk Trust Center).
The procurement question is not simply whether Zendesk can sign a BAA. It is whether the agreement covers the service plan, agent feature, integration, channel and data path in your proposed design.
Sierra
Sierra’s Trust Center lists a Sierra SOC 2 Type II Audit Report for 2026 and a Sierra HIPAA Audit Report for 2026. The underlying documents are available through an access request. The same public materials list model-service subprocessors and locations, including third-party cloud and model providers (Sierra Trust Center).
Sierra’s public materials show SOC 2 Type II and HIPAA artifacts, but buyers should obtain written confirmation that a BAA will cover their contracted products, channels, data flow and subprocessors before PHI processing.
Decagon
Decagon’s Trust Center publicly lists SOC 2 Type II reports and HIPAA among its compliance items. The reports require an access request (Decagon Trust Center).
Its security page says that Decagon maintains production-system audit logs, uses role-based access control and encrypts web traffic and stored data (Decagon Security and Compliance). The public materials reviewed do not state that Decagon will sign a BAA or define BAA-covered services. That is an evidence gap for a healthcare buyer, not proof that a BAA is unavailable.
For any vendor, request the current report, opinion letter, scope statement, exceptions and bridge letter if the report period has ended. Do not treat a public “HIPAA” listing as proof that a particular product configuration is covered, and do not describe a vendor as “HIPAA certified.”
Private cloud, on-premises and data-boundary options
Deployment documentation describes where a platform or model can run. It does not determine whether a particular implementation meets every regulated buyer’s policy. The OCC’s guidance makes the operational boundary important because third-party access and reliance can affect a bank’s direct control (OCC, 2024).
The following public documentation describes deployment capabilities, not a determination that a particular deployment meets every regulated buyer’s policy.
“Private cloud,” “VPC” and “on-premises” are not interchangeable. A private cloud may still give the vendor administrative access. A hybrid design may still require outbound content or monitoring connectivity. C3 AI’s documentation illustrates why buyers should inspect operational responsibility and access paths, not only the deployment label.
Ask vendors to map every data path, including:
- Prompts and retrieved content
- Model inference and tool calls
- Transcripts and application logs
- Backups and disaster-recovery copies
- Vendor support access
- Monitoring components
- Model-service subprocessors
- Egress controls and network dependencies
The production-change test: audit logging plus human approval
A buyer should distinguish workflow actions from platform changes. A system may log what an agent did without providing a governed approval process for changes to prompts, tools, guardrails, models or workflows.
Require an evidence record for each production change that links:
- The trigger or observed issue
- The proposed change
- Test results
- Approver identity
- Approval time
- Deployment version
- Rollback path
- Post-release monitoring result
NuPlay provides a documented example of this operating model. NuLoop replays runs, proposes changes, versions every release and, in its approve-to-promote mode, holds changes in one queue with supporting evidence until a human approves or rejects them (NuLoop).
NuPlay runs enterprise workflows in production and improves them after every run through NuLoop, with agents, the systems they operate, and the context they draw on under one platform.
The relevant control is approve-to-promote, not unsupervised autonomy. NuLoop’s stages are Report, Diagnose, Propose, Try and Ship. A proposed change can be tested and prepared for release, and in approve-to-promote mode a human approves it before it reaches production. For a regulated buyer, confirm that approve-to-promote is the contracted mode and request evidence that the approval step applies to the exact workflow and environment under review.
Model governance and explainability for insurers
The National Association of Insurance Commissioners (NAIC) model bulletin says insurers remain responsible for actions and decisions supported by artificial intelligence systems. It expects insurers to maintain a written artificial intelligence systems program designed to mitigate adverse consumer outcomes. The bulletin also addresses transparency, explainability, governance, risk controls, testing, documentation and third-party involvement (NAIC Model Bulletin, adopted December 4, 2023).
The bulletin is a model: it applies in a state only when that state’s insurance department issues it. As of the NAIC’s Spring 2026 meeting, 24 states and the District of Columbia had adopted it (Mayer Brown summary). Check the status in each state where you write business.
For an insurance use case, ask the vendor and internal business owner for an evidence package covering:
- Model inventory: Model, version, provider, task, inputs, outputs and deployment location.
- Decision boundary: What the agent may recommend, execute, route or escalate.
- Human accountability: Named business owner, approval authority and override route.
- Explainability artifact: How the system records the policy, data, tools and rationale used in a material outcome.
- Testing evidence: Bias, error, failure-mode and drift testing appropriate to the use case.
- Change file: Version history, approval record, test evidence, rollback plan and issue remediation.
- Third-party chain: Subprocessors, model providers, data uses and contractual obligations.
A broad governance description does not prove explainability for a regulated insurance decision. Map each artifact to the specific action the agent can take and the consumer impact that could follow.
If the agent only retrieves claim status and drafts a handoff, the evidence burden differs from a system that recommends coverage, routes a claim based on risk or supports a claim decision. The decision boundary should be documented before technical testing begins.
Hypothetical example: an insurer applying bank-style review to a claims-status agent
A regional insurer wants an agent to answer claim-status questions, retrieve policy information and draft a handoff summary. It must not approve or deny claims.
The insurer applies the same control discipline expected in a bank third-party review: define the data boundary, test vendor access, document sub-processors and require traceable production changes.
Allowed
- Read-only policy and claim lookup
- Pre-approved answers
- Human handoff
- Event logs
- Escalation when the agent cannot answer within its permitted scope
Blocked pending evidence
- PHI entry before BAA execution
- Model-training use of conversations
- Unmanaged model subprocessors
- Agent write access to claim status
- Any change to production behavior without human approval
Required controls
- Least-privilege integration credentials
- A transcript retention schedule
- Escalation rules
- Current SOC 2 Type II scope
- A BAA-covered service schedule
- Private deployment evidence if internal policy requires it
- Sampled audit logs
- A human-approved production-change process
Success check
Security should be able to select an interaction and trace it from user input through retrieval, agent output, tool activity, human escalation and any related configuration version.
This test separates low-risk assistance from regulated decision-making. It also gives security and procurement a concrete demonstration to request instead of accepting a general platform presentation.
State the line of business first. If the insurer is a health plan, HIPAA applies and a BAA must be in place before PHI enters the workflow. For other lines, different privacy rules apply and the BAA row in the checklist may not. Have counsel confirm which applies.
**What does a review of a claims-status agent look like in practice? **
Imagine a regional insurer’s claim-status agent on an ordinary Tuesday. It can read a claim record, answer from approved language, and hand off to a person. It cannot approve or deny a claim, and it cannot write to the claim.
2:14pm - A policyholder asks where her claim stands. The agent verifies her with the insurer’s existing step and logs the event.
2:14pm - The agent looks up the claim with a least-privilege, read-only credential. Nothing is written back.
2:15pm - The agent gives the status and the next step from approved language. She asks whether the repair estimate will be approved, which is outside the agent’s scope, so it hands off to an adjuster with the transcript and claim number.
2:16pm - The adjuster picks up with the context. The log records which agent configuration version produced the answer.
Next Monday - A security analyst picks this interaction from the log and traces it from the policyholder’s message through retrieval, the agent’s answer, the tool call, the handoff and the configuration version, then checks that the version had an approved change record.
No single step here is remarkable. What a regulated buyer needs is that every one of them can be reconstructed by someone who was not there.
What to send vendors before security review
Send the following request list before a formal security review:
- Current SOC 2 Type II report, scope, exceptions and bridge letter.
- BAA template and covered-service schedule where PHI is involved.
- Data-flow diagram covering prompts, logs, retrieval, model calls, backups, support access and subprocessors.
- Subprocessor list with each party’s role, location and notification terms.
- Deployment architecture and vendor-administrator access model.
- Data-retention, deletion and customer-data training terms.
- Sample audit log and production-change approval record.
- Incident response, breach notification, business continuity and exit/deletion procedures.
- Model inventory and governance documentation for regulated decisions.
- Written answers identifying what is unavailable or not publicly documented.
Ask the vendor to mark each answer as documented, contractually available, available only under request access or not currently available. That classification prevents a questionnaire response from being mistaken for independent evidence.
The failure case to avoid is accepting “SOC 2 compliant” or “HIPAA ready” when the vendor cannot identify the exact service scope, legal coverage or data boundary. The same applies to “private deployment” when the vendor cannot explain support access, outbound connections, backup location or model-provider calls.
.gif)






