
The NAIC AI Evaluation Tool Pilot Closes Next Month. Here Is What That Actually Means.
The NAIC AI Evaluation Tool pilot (now formally referred to as the AI Risk Evaluation Supplement, or simply the supplement) launched in March 2026 across 12 participating states: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin. It closes next month. When it does, the industry's relationship with this framework changes in a way that matters to every insurer, not just the ones operating in those 12 states.
Between September and October, the NAIC's Big Data and Artificial Intelligence Working Group will refine the supplement based on pilot feedback, with an exposure draft targeted for early September and a comment period before the fall national meeting. In November, the updated version goes to a vote. If it is adopted, as the Working Group's stated timeline and the NAIC's own topic page both indicate is expected, every state insurance department in the country gets access to it. The supplement is built to slot into existing market conduct exams and financial examinations, and states can also deploy it as a standalone questionnaire outside the full examination cycle. The question shifts from whether your state will use it to when it shows up in your next examination workflow.
What the NAIC AI Evaluation Tool Actually Measures
The supplement was built around a proportionality principle: regulators concentrate examination resources on AI systems that carry real consumer or financial risk, rather than applying uniform scrutiny across every back-office application. In practice, that means the systems at the center of your business, underwriting, claims, pricing, fraud detection, are where examination attention concentrates.
The supplement also settled one question the industry spent considerable energy lobbying around. Vendor origin is not a governance defense. If an AI system affects policyholder decisions, the carrier owns the governance responsibility, regardless of whether that system was built internally or purchased from a third party. That expectation is built into the framework's design, not added as an interpretation, and carriers that have leaned on vendor relationships as a proxy for oversight are going to find the four exhibits clarifying on that point.
The Four Exhibits: Where AI Governance Auditing Gets Specific
The supplement runs on four exhibits, and each one surfaces a different kind of preparation gap.
Exhibit A asks insurers to document how extensively they use AI systems, in which functions, and affecting which types of decisions. The scope is wider than most teams expect on first read because it includes vendor-embedded models and machine learning features built into larger third-party platforms, not only systems your team built and maintains. Carriers with a current, maintained model registry can move through Exhibit A efficiently. Carriers without one learn how much time an AI usage inventory takes when it has to be built from scratch under a deadline.
Exhibit B addresses AI governance risk assessment: accountability structures, risk management policies, documentation practices, and evidence that governance operated on an ongoing basis. This is where the gap between a real program and a paper program becomes visible to regulators. Contemporaneous records matter here. Dated approval logs, test results tied to actual deployment timelines, and change documentation that aligns with your system's history read as evidence of a control that functioned. Documentation assembled in advance of an exam request does not read the same way, and experienced examiners know the difference.
Exhibit C covers high-risk AI systems in detail. For each system classified as high-risk, carriers are expected to provide documentation on model design, training data, validation procedures, performance metrics, and bias testing. The proportionality principle means Exhibit C is where examination attention concentrates, which also means it is where documentation gaps create the most exposure.
Exhibit D covers AI data specifics: sources, quality controls, representativeness, and potential for proxy discrimination in rate setting. The current version includes a field on reasonable accommodations and policy modifications, reflecting a regulatory focus on algorithmic fairness that is becoming more specific over time. The Foley & Lardner analysis of what to do when you receive a pilot request and the Fenwick breakdown of the 12-state program are the most practically detailed external resources available on approaching these exhibits operationally.
What the August 13 Meeting Revealed
The August 13 Summer National Meeting session added substance that goes well beyond pilot status. Florida noted during the session that the pilot has been "really effective," with strong cross-state collaboration and pilot states meeting nearly weekly to share findings. The exposure draft is targeted for early September, moving faster than earlier timelines suggested.
The session's main content, an AM Best presentation by Edin Izerovich, covered how AI governance needs to evolve as carriers move across three distinct AI types, each with its own regulatory demands.
Predictive AI uses structured data to produce scores and rankings. Governance questions here center on model testing, fairness, documentation, and performance drift over time. Generative AI uses unstructured data to produce summaries and extracts. The governance questions shift: output trust, how prompt changes affect results, sensitive data boundaries, and whether vendor model updates change what the system actually does. Agentic AI takes multi-step, tool-using action in the world. Its governance demands are the most intensive: permissions management, action logging, meaningful human override capability, and rollback procedures.
These risks compound. An agentic system carries predictive and generative risks on top of its own. Approving one layer does not cover the others.
The AM Best framing organized governance intensity around two factors: impact (what are the consumer and financial stakes of a mistake) and autonomy (how much does the system act independently). Systems with low impact and low autonomy, an underwriting research assistant for example, require limited governance. Systems with high autonomy and high impact, automated claims routing or payment execution, sit in the most demanding governance zone. Governance investment should follow where autonomy and impact intersect, not be allocated uniformly.
The session also drew a distinction worth internalizing before your next examiner conversation: documentary governance answers with policies, operational governance answers with evidence. Regulators are looking for evidence in four specific areas. First, an AI inventory that captures what systems exist, what decisions they affect, and how much autonomy each has. Second, a validation environment that shows what triggered a review, what tests ran before go-live, and what changed after a problem was identified. Third, real examples of human oversight, instances where a person challenged, changed, or stopped an AI output, not a policy that says humans are in the loop. Fourth, for agentic systems, demonstrated ability to see vendor changes, test their effects, stop the system, and roll back to a safe state.
The most common failure mode observed across the industry: AI tools added to existing workflows without redesigning the operating model or reassigning accountability. Pilots succeed, usage spreads, governance never catches up.
Two emerging issues raised at the session have direct implications for AI inventory and vendor risk management. The AM Best presentation cited research indicating that increasing compute budget, without any change to the underlying model, can materially improve performance. The implication is that approving a model is not the same as approving its deployment, and capability changes may not be visible through standard model versioning controls. The session also referenced UK AI Security Institute data suggesting the share of AI tool usage involving action-taking agents grew from roughly 25 percent to 65 percent in 16 months, and that payment-executing servers grew from approximately 46 to more than 1,200 in a single year. Worth verifying against the primary source, but if directionally accurate, the inventory question is no longer just about what models a carrier uses. It is about what tools those models can call and what actions they can take. Vendor concentration was flagged as a related systemic risk: when many carriers depend on the same provider, the exposure is correlated across the industry, not isolated.
What AI Compliance Looks Like After the Pilot Closes
The final version of the supplement will be more refined than the current one. It will not be less rigorous. Once it is in the hands of every state insurance department, the examination environment changes for all carriers simultaneously, regardless of size or product line. Small mutual carriers face the same exhibit requirements as national P&C carriers. The framework does not distinguish.
The carriers in the best position when November arrives are the ones that treated the pilot window as a chance to build toward the standard the four exhibits represent. Monitaur's early overview of the NAIC AI systems evaluation tool pilot covers what the supplement was designed to accomplish and where carriers should start.
Where Monitaur Fits in Your AI Governance Assessment
Monitaur's platform maps directly to what the four exhibits require. Exhibit A is answered from a maintained, auditable model registry rather than a manual inventory built under exam pressure. Exhibit B documentation, the contemporaneous governance records that distinguish a real program from an assembled one, is generated continuously, not reconstructed when a request arrives. Exhibit C documentation for high-risk systems is maintained as an ongoing record. Exhibit D data controls are tracked at the system level.
The distinction the August 13 session drew between documentary and operational governance is exactly what Monitaur is built to address: operationalizing policies with verifiable evidence, objective testing of critical systems, and documented critical failure analysis. The longer-term goal, raised in discussions at the session, is for AI governance to become part of general risk evaluation rather than a standalone compliance function. That is the direction the supplement is moving.
Monitaur's AI governance assessment maps your current program against what the four exhibits require so you can identify gaps while there is still time to close them.
The NAIC AI Evaluation Tool pilot closes next month. A nationwide rollout is expected in November. The carriers that treat September as a deadline for getting their AI governance program in order will be in a materially different position from those that treat it as someone else's problem until their state adopts the framework. At that point, the runway to prepare outside of exam pressure is gone.