Questions enterprise security, platform and compliance teams ask when evaluating AI governance infrastructure — answered directly. Where a regulation is involved the answer states what the regulation actually says as of September 2026, including where it says nothing.
The category
What is runtime AI governance?
Enforcement of AI policy at the moment a call is made, rather than documentation written beforehand or logs reviewed afterwards. The decision point sits in the request path: a prompt is evaluated, an action is permitted or refused, and the record of that decision is written as it happens. Full definition: runtime governance.
How is an AI gateway different from an API gateway?
An API gateway's unit of work is the request — a method, a path, headers, a body, a consumer. An AI gateway has to reason about the content of the request: what was asked, what model will answer, what the response contains, and what tools the model decided to call. Token metering, semantic caching and prompt inspection have no equivalent in API management. See the enterprise AI gateway guide.
Is an AI gateway the same thing as an AI control plane?
Not quite. A gateway routes and observes. A control plane decides — it enforces policy before the action executes and binds each action to an identity. Many products called gateways do both; some do only the first. The test is whether a call can be refused, and whether the refusal is recorded with the policy version that caused it.
What is the difference between AI governance and AI security?
AI security asks whether an attacker can make the system do something it should not. AI governance asks whether the organization can prove what the system was permitted to do and what it actually did. They overlap at the enforcement point and diverge in what they produce: security produces prevention, governance produces evidence.
Where does APERION sit?
Smartflow operates at the model request path and the tool path, in one control plane, on-premises or air-gapped inside the customer's perimeter. It does not scan model artifacts, red-team agents or manage SaaS posture. The full field is mapped by control point in the AI security and governance landscape.
Evaluating vendors
Which vendors support per-team model allowlists and deny-lists for risky capabilities?
Most enterprise AI gateways support model allowlists scoped to a team, key or workspace — LiteLLM through team IDs, Portkey through workspaces, Kong through consumer identity, Smartflow through virtual keys carrying model_restrictions, rate limits and spend caps. Capability guardrails on what the model may then do are usually a separate product category. Two questions separate real implementations: is the allowlist evaluated before or after the gateway rewrites model choice for routing, and does a deny-list entry override an allowlist entry. Detail: model allowlist and deny-list.
What should I ask a vendor that I would not think to ask?
Four questions. Where exactly does this intervene, and what does that position not see? Can it run where our data has to stay — on-premises and air-gapped specifically, not single-tenant with customer-managed keys? Does an agent get an identity, or borrow one? And can it produce the artifact our regulator asks for, as distinct from a control-framework mapping? Ask for a sample of the last one.
What is the difference between on-premises, private cloud and air-gapped?
Private cloud means the workload runs in your cloud account, possibly while the control plane still calls a vendor service to decide what is permitted. On-premises means it runs on your infrastructure, possibly with outbound connectivity for licensing or policy updates. Air-gapped means no outbound path exists at request time. A product can be moved on-premises late; air-gapped has to be true from the beginning. See air-gapped AI deployment.
Do I need a gateway if I already have a CASB or secure web gateway?
They cover different traffic. A CASB or SWG sees what employees do in browsers and SaaS applications. An AI gateway sees what applications and agents do when they call models — including inside a VPC where nothing crosses a routable network path. Most enterprises eventually need both coverage models; which comes first depends on whether the immediate problem is employees using AI you cannot see, or applications making calls you cannot prove you governed.
Should we buy a governance platform or a gateway?
They answer different questions and are frequently deployed together. A governance platform produces the registry, risk assessments and conformity documentation an auditor reviews. A gateway produces the runtime record showing the documented policy was actually in force. A program with documentation and no enforcement describes controls that may not be operating; one with enforcement and no documentation can show what happened but not that it was intended. See the Credo AI comparison.
How do we compare vendors that all say "runtime governance"?
By control point, not by category language. Nine positions exist — model request path, network path, tool and MCP path, agent harness, agent action layer inside SaaS platforms, endpoint and browser, credential path, posture and supply chain, and out-of-band assessment. Two vendors using identical words can sit in entirely different parts of the stack and fail in entirely different ways. The map is here.
Regulation
What does SR 11-7 require now?
Nothing — it was superseded on 17 April 2026 by SR 26-2, with the OCC rescinding Bulletin 2011-12 the same day. The substance of model risk management carried forward, but with two changes most summaries miss: SR 26-2 states it does not set forth enforceable standards, and it is expected to be most relevant to institutions above $30 billion in total assets. See SR 11-7 and SR 26-2.
Does model risk guidance cover generative and agentic AI?
No, and this is the most consequential fact in US bank supervision right now. SR 26-2 states that generative and agentic AI models "are novel and rapidly evolving. As such, they are not within the scope of this guidance." The same passage instructs that the institution's own risk management practices should determine appropriate controls. The obligation did not disappear; it relocated onto the institution, without a framework to point at.
Is there federal AI guidance for banks?
As of September 2026, no US federal banking agency has issued AI-specific model risk or governance guidance. A request for information on bank use of AI was announced alongside SR 26-2 and has not been issued. The FFIEC has published nothing standalone on AI and sunset its Cybersecurity Assessment Tool on 31 August 2025 without a replacement. See FFIEC.
What does FINRA Rule 3110 require for AI?
Rule 3110 is not AI-specific and that is the point — it is an enforceable supervision rule, where model risk guidance is not. It asks whether the firm supervised the conduct. A model can be well validated and still used in a way nobody supervised, and an agent acting on a borrowed credential cannot answer the question of who supervised it. See FINRA Rule 3110.
Has the HIPAA Security Rule been updated for AI?
No. A proposal to strengthen it was published 6 January 2025 at 90 FR 898 and has not been finalized — the 2026 federal regulatory agenda classifies it as a long-term action with final action projected no earlier than July 2027. The operative text is unchanged since 2013, and OCR has issued no AI-specific guidance. An AI system handling ePHI is governed by the existing rule. See HIPAA Security Rule and AI.
Does "addressable" under HIPAA mean optional?
No. It means assess whether the specification is reasonable and appropriate in your environment, then either implement it or document why not and implement an equivalent alternative. Every addressable specification produces a mandatory documented decision. The flexibility is in the control chosen, not in whether to decide.
What are the current NYDFS Part 500 deadlines?
There are none outstanding. Every Second Amendment transitional deadline has passed, the last on 1 November 2025 carrying multi-factor authentication and the asset inventory requirement. Commentary written before that date still describes them as upcoming. See NYDFS Part 500.
Does NYDFS require anything specific for AI?
No. NYDFS has issued two AI letters — 16 October 2024 on cybersecurity risks arising from AI, and 21 May 2026 on frontier model risks — and both state explicitly that they impose no new requirements. AI is something the existing Part 500 framework must be applied to. In practice the obligations that bite are the ones already there: the asset inventory must include AI systems, MFA must cover access to them, and the 72-hour notification clock runs on an AI-related incident.
Does ISO/IEC 42001 certification satisfy the EU AI Act?
No, and this is routinely overstated. The European Commission has stated that although ISO/IEC 42001 helps establish an AI management system, "its goals and definitions are not aligned with the quality management system that is required under the AI Act," and it requested a separate European standard as a result. Presumption of conformity attaches only once a harmonized standard is cited in the Official Journal. Certification is credible evidence of governance maturity; it is not conformity. See ISO/IEC 42001.
Is the NIST AI RMF mandatory?
No. It is voluntary guidance organized around Govern, Map, Measure and Manage. It tells an institution what to be able to demonstrate; it does not generate the demonstration, which is where most programs stall — the mapping is complete and the evidence behind it is a spreadsheet maintained by hand. See NIST AI RMF.
Which GLBA rule applies to us?
Depends what you are. Banks, thrifts and credit unions fall under the Interagency Guidelines implemented by their prudential regulator. Non-bank financial institutions fall under the FTC Safeguards Rule at 16 CFR part 314. Broker-dealers answer to the SEC and insurers to their state regulator. Hospitals are generally not GLBA-covered at all. See GLBA.
What has to be reported, and how fast?
The clocks come from existing regimes, not AI-specific ones. NYDFS Part 500: 72 hours from determining an incident occurred, 24 hours for an extortion payment plus a written explanation within 30 days. FTC Safeguards Rule: 30 days for an event affecting 500 or more consumers. None have an AI carve-out. See AI incident response.
Agents
What is agent identity and why does it matter?
Most products attribute agent activity back to a human — correlating to the invoking user or mapping every action to a person's credential. That answers who is responsible. It does not let an agent be revoked without disabling an employee, does not let an information barrier be enforced against the agent as a party, and does not produce a supervision record naming the acting entity. See agent authority scope.
How do you govern agent-to-agent traffic?
By asking one question at every hop: is this task still inside the authority the original human held? That requires the agent to have an identity of its own, authority that narrows rather than widens as it delegates, and a decision made in the path rather than in a log reviewed later. See A2A.
What is the difference between governing MCP and governing the model?
The model produces text. The tools do things. An agent governed on model calls and not on tool calls is governed on neither, because the consequential action happens at the tool. MCP governance covers which servers are reachable, which individual tools are enabled, what credential the call carries, and whether an ungoverned call fails. See MCP security.
Can you stop prompt injection?
Not reliably by filtering, because the attack is instruction-shaped content arriving through a channel the model was told to trust. The defenses that hold are architectural: constrain what the model may do rather than what it may read, inspect the output path as well as the input, and treat retrieved documents and tool responses as untrusted regardless of origin. See prompt injection and indirect prompt injection.
How do you approve an MCP server safely?
Approve tools rather than servers, with deny overriding allow so a newly added tool is not automatically reachable. Record the approved tool definition and detect when what is served stops matching it — a one-time security review is not a control, because nothing in the protocol requires definitions to stay the same after approval. See tool poisoning.
Where should a human approval gate sit?
Gate by consequence and reversibility. An agent reading a record needs no gate. An agent writing irreversibly, moving money, contacting a customer or granting access does. The decision belongs in versioned policy rather than in an agent's prompt, because an instruction to ask permission is a request the agent can be talked out of. See human in the loop.
Architecture and operations
Does a gateway add latency?
Yes, and any vendor saying otherwise is describing something else. The figure worth asking for is measured end-to-end overhead on a cache-miss workload — every request traversing the full pipeline with no cache benefit — at several request rates, so you can see whether overhead is stable as load rises. Ask for the method alongside the number.
What does semantic caching actually save?
It removes calls rather than redirecting them. Routing sends a request to a cheaper model; semantic caching avoids the request when the question has already been answered in substance. On enterprise traffic, where the same question arrives phrased many ways, they are additive. The similarity threshold is a governance setting, not a tuning dial — set it loose and the system answers a question nobody asked. See semantic caching.
How do you detect that a model changed underneath you?
Capture the model identifier on every request, including what the gateway routed to after any rewrite, and run a small fixed probe set on a schedule so a change in the provider's artifact shows as a change in answers. A validated model can stop being the model that was validated with no event in your own change control. See model drift.
What belongs in an AI asset inventory?
Models and versions, data sources, frameworks, prompts and system instructions, and the tools and external services the system can reach. The last two are the ones most inventories miss and the ones that change most often. Deriving the inventory from observed traffic beats maintaining a document, because a declared list says what teams intended to use. See AI bill of materials.
What makes an audit log usable as evidence?
That alteration is detectable, usually through hash chaining, with append-only storage and periodic anchoring somewhere the log operator does not control. Each entry has to carry the actor — including when the actor is an agent — the policy version in force, the decision and the time. An entry without the policy version shows what happened but not what rule produced it. See tamper-evident audit log.
How long do we have to keep records?
Longer than most vendor defaults. HIPAA documentation runs to six years, SOX to seven, and several supervisory contexts longer. A vendor-imposed retention ceiling is a governance constraint rather than a pricing detail — worth checking, since documented ceilings in this market range from seven days to tiered plans.
Can this run air-gapped?
Smartflow can. Supporting it properly means accepting constraints elsewhere: threat intelligence and model updates arrive on a media transfer cadence, observability stays inside so vendor-side diagnosis is unavailable, and policy distribution needs signed offline bundles. Across the security-branded AI governance cohort, on-premises deployment is largely undocumented and air-gapped operation is documented by almost none of it.
Does red teaming replace runtime controls?
No. A red team result is a statement about a configuration on a date. Providers update models, the tool surface grows, and published techniques evolve continuously. The useful pattern is a confirmed finding becoming an enforced policy at the gateway, with the finding identifier carried into the enforcement record, so the institution can show both how the control was tested and what it stopped. See AI red teaming.
What can a runtime control not do?
It cannot inspect a training corpus, detect a backdoor in weights it did not produce, or scan a model artifact before deployment. Those controls sit upstream — provenance for training data, signing and verification for artifacts. What a runtime control contributes is bounding the consequence, recording behavior over time, and governing which models are reachable at all. See data poisoning.
Getting started
What is the fastest way to find out what AI is already in use?
Route AI traffic through a control plane and read the result. Shadow AI is invisible at the application layer and visible at the network path, because every call to an external model leaves the perimeter. Discovery is not a separate exercise — it is the first output of putting a gateway in the path. See govern shadow AI.
Where do most programs stall?
Between the framework and the evidence. The control mapping is complete, the policy is written, and the artifact a supervisor asks for is still assembled by hand from spreadsheets and log exports when the request arrives. Closing that gap means treating enforcement records as the source of the evidence rather than a parallel system. See examination readiness.
Who owns AI governance?
In practice it is contested between the CISO, the platform team and the risk function, which is why policy as code tends to be where it settles. It is the one artifact all three can read: the platform team gets something testable, the risk function gets something auditable with a change history, and the CISO gets a control that operates identically whether or not anyone is watching.
Related reading
Glossary · Comparisons · The landscape, mapped by control point · Trust Fabric architecture · Examination readiness
Regulatory positions verified against primary sources in September 2026. Several of these are live rulemakings. If something here is out of date, tell us and we will correct it.
Put this in the path of your own agents.
Policy enforced inline between your agents and every model and tool they reach, with a record bound to the human who owns it.
Request a Demo Read the docs