The NAIC AI Evaluation Tool Pilot Closes in September

Regulation & Legislation

The NAIC AI Evaluation Tool Pilot Closes Next Month. Here Is What That Means.

The NAIC AI Evaluation Tool pilot launched in March 2026 across 12 states: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin and it closes next month. When it closes, the industry's relationship with this framework changes in a way that matters to every insurer, not just the ones operating in those 12 states.

Between September and October, the NAIC's Big Data and Artificial Intelligence Working Group will refine the tool based on pilot feedback. In November, the updated version goes to a vote at the Fall National Meeting. If it is adopted, as the Working Group's stated timeline and the NAIC's own topic page both indicate is expected, every state insurance department in the country gets access to it. The question stops being whether your state is participating in the pilot. It becomes,  when will your examiner schedule their first AI governance audit using this framework?

What the NAIC AI Evaluation Tool Measures

The framework was built around a proportionality principle: regulators concentrate examination resources on AI systems that carry real consumer or financial risk, rather than applying uniform scrutiny across every back-office application. That sounds reasonable in theory. In practice, it means the systems at the center of your business, underwriting, claims, pricing, fraud detection, are where regulators are looking most closely.

The tool also settled one question the industry spent considerable energy lobbying around. Vendor origin is not a governance defense. If an AI system affects policyholder decisions, the carrier owns the governance responsibility, regardless of whether that system was built internally or purchased from a third party. That expectation is part of the framework's design, not added as an interpretation, and carriers that have leaned on vendor relationships as a proxy for oversight are going to find the four exhibits clarifying on that point.

The Four Exhibits: Where AI Governance Auditing Gets Specific

The evaluation framework runs on four exhibits, and each one surfaces a different kind of preparation gap.

Exhibit A asks insurers to document how extensively they use AI systems, in which functions, and affecting which types of decisions. The scope is wider than most teams expect on first read because it includes vendor-embedded models and machine learning features built into larger platforms, not only systems your team built and maintains. Carriers with a current, maintained model registry can move through Exhibit A efficiently. Carriers without one learn how much time an AI usage inventory takes when it must be built from scratch under a deadline.

Exhibit B addresses AI governance risk assessment: the accountability structures, risk management policies, documentation practices, and evidence that governance operated on an ongoing basis. This exhibit is where the gap between a real program and a paper program becomes visible to regulators. Contemporaneous records matter here. Dated approval logs, test results tied to actual deployment timelines, and change documentation that aligns with your system's history read as evidence of a control that functioned. Documentation assembled in advance of an exam request reads differently, and experienced examiners know the difference.

Exhibit C covers high-risk AI systems in detail. For each system classified as high-risk, carriers are expected to provide documentation on model design, training data, validation procedures, performance metrics, and bias testing. The proportionality principle means Exhibit C is where examination attention concentrates, which means it is also where documentation gaps create the most exposure.

Exhibit D covers AI data specifics: sources, quality controls, representativeness, and potential for proxy discrimination, particularly in rate setting. The current version of the tool includes a field on reasonable accommodations and policy modifications, which reflects a regulatory focus on algorithmic fairness that is becoming more specific, not less, over time. For a detailed breakdown of how to approach these exhibits operationally, the Foley & Lardner analysis of what to do when you receive a pilot request and the Fenwick breakdown of the 12-state program are the most practically detailed external resources currently available.

What the August Working Group Meeting Signals

The Working Group's August 13 Summer National Meeting session included two discussions worth flagging for carriers thinking about where AI governance regulations are heading beyond the current pilot.

AM Best presented how governance needs to evolve as insurers move from predictive AI into generative and agentic systems, emphasizing governance scaled proportionally to each use case's risk level, along with validation, human oversight, monitoring, documentation, and rollback procedures for more autonomous applications. The current four-exhibit framework was built around the AI systems most insurers are running today. The AM Best framing signals that regulators are already thinking about what comes after that.

There was also a panel discussion on how generalized linear models fit within the tool's scope relative to more complex AI and machine learning systems. That question has not been publicly resolved, and the outcome will affect how actuarial systems get classified under Exhibit C. Carriers with significant GLM-based pricing or reserving models should watch for how the Working Group addresses this in the October refinement period.

What AI Compliance Looks Like After the Pilot Closes

The final version of the tool will be more refined than the current version. It will not be less rigorous. And once it is in the hands of every state insurance department, the examination environment for AI governance changes for all carriers simultaneously, regardless of size or product line. Small mutual carriers face the same exhibit requirements as national P&C carriers. The framework does not distinguish.

Monitaur's early overview of the NAIC AI systems evaluation tool pilot covers what the framework was designed to accomplish. The short version: regulators wanted a consistent, structured way to assess AI governance across the industry. They built one. It is about to go live nationwide.

Where Monitaur Fits in Your AI Governance Assessment

Monitaur's platform maps directly to what the four exhibits require. Exhibit A is answered from a maintained, auditable model registry rather than a manual inventory built under exam pressure. Exhibit B documentation, the contemporaneous governance records that distinguish a real program from an assembled one, is generated continuously by the platform's oversight controls, not reconstructed when a request arrives. Exhibit C documentation for high-risk systems is maintained as an ongoing record. Exhibit D data controls are tracked at the system level.

The companies that will move through an AI governance audit most cleanly are the ones that have been running a program all along, not preparing for an exam. Monitaur's AI governance assessment maps your current program against what the four exhibits require so you can identify gaps while you still have time to close them.

The NAIC AI Evaluation Tool pilot closes next month. A nationwide rollout is expected in November. The carriers that treat September as a deadline for getting their AI governance program in order will be in a materially different position from the ones that treat it as someone else's problem until their state adopts the framework. At that point, the runway to prepare outside of exam pressure is gone.