Blog · September 27, 2026

Evaluating AI-based proxy research quality

Why it is time to establish a framework for a nascent industry

By Alexander Kaltenböck and Nicolaas Koster, Proxywise AI

Also published on LinkedIn: read it there →

How do you evaluate quality when AI generates proxy research?

This is a question that has been top of mind for us and many partners we have been working with as we build our AI engine for proxy research. Fiduciary duty is taken seriously in corporate governance, and mistakes are seen as failures of a system built carefully over decades.

We believe each investor should be able to apply their own voting policy at scale. LLMs have only recently become powerful enough to do this within guardrails, with every recommendation resting on facts from publicly available information and the investor's policy. AI enables policy application at scale; it does not replace the stewardship team's judgment.

But how do you know such an AI-powered engine is doing a good job? What is the KPI? So far, nobody publishes a consistent one. Glass Lewis's AI series, from “AI and the Fiduciary Test” to “The Architecture of Investor-Grade AI”, focuses on governance and oversight. It describes quality-control processes but publishes no evaluation results: no accuracy rates, benchmarks, or measurement methodology. Human proxy research never had a quality KPI either: the only public record of its errors are issuers' complaints in supplemental filings, which the advisors mostly dispute.

A backtest against actual votes is a proxy, not a quality KPI

A backtest against past votes mixes two questions: did the engine apply the policy, and did the organization follow it? Most investors vote the vast majority of items with management, so an engine that simply voted with management would score above 80% at most organizations. Academic work that measures agreement with an established proxy advisor policy (79% with ISS on shareholder proposals, in one study) has the same problem. Actual votes help calibrate a policy; they are never the answer key.

An industry without agreed KPIs for the quality of its work cannot be taken seriously by its users. We found no existing standard, so we took our own shot at an evaluation framework, and we invite others to contribute.

The evaluation framework and first results

Our framework compares the engine's output to a golden dataset: a human-built, human-checked answer key. For each meeting it records every ballot item, the facts needed to decide it, the rules that should fire (and near-miss rules that must not), the expected recommendation, and the arguments an analyst would expect.

The headline number is the F1 score of precision (was what the engine raised right?) and recall (did it raise everything that deserved it?): 2 × precision × recall ÷ (precision + recall). Unlike an average, it stays low when either side is low: an engine that is always right but raises only half of what it should scores 67%, not 75%. A critical error, such as a vote in the wrong direction, scores the item wrong: a good rationale does not rescue a wrong vote.

The evaluation framework: a composite F1 score of precision (was what we raised right?) and recall (did we raise all that deserved it?), above five steps scored one by one: D1 Ballot, D2 Facts, D3 Rules, D4 Recommendation and D5 Rationale.
Our attempt at a comprehensive quality framework

Below the composite sit five dimensions, one per step the engine performs, so every error is traced to where it started. Only the rationale (D5) needs judgment to score: a language model grades it against a human-written checklist of must-have arguments, with human spot checks.

Because it scores every step, the framework catches what a backtest cannot, such as a right vote for the wrong reasons or an outdated policy version. It also tests the policy itself: every golden dataset so far has led to drafting improvements.

Each golden dataset is built from the issuer's public filings. A human drafts the expected answers against the investor's policy, and an independent reviewer audits them. No one looks at engine output while building it.

One example from our full-policy dataset: Pfizer's 2026 annual meeting (16 ballot items, 135 rule checks) scored 100% on ballot / fact extraction and rule detection and 94% on recommendations; the one error resulted from the engine not finding the right information, which we could fix with one update.

Our single-rule dataset on unequal voting rights (114 director elections at 11 meetings, five of them at dual-class companies) reached 94% recall with no false flags and 99% recommendation accuracy.

The samples are small, so read these as a first indication. But we are working on expanding the golden dataset as well as the scope of these evaluations, using our methodology.

Implications and next steps

We hope that sharing these measures publicly will support a transparent discussion about quality in a nascent industry. But we are also aware that any provider of AI-based proxy research should not only assess their own work. Legal research providers went through this two years ago: an independent Stanford study found vendors' “hallucination-free” claims overstated, and an independent benchmark followed within a year. We would rather start this than wait for it, so we propose a working group of AI proxy research providers to agree on common metrics, a common golden dataset format and a common report card.

In the coming weeks we will publish our framework in a public repository, with scoring scripts and a complete golden dataset for a public U.S. annual meeting, and invite anyone building AI-based proxy research to test their engine against it and join.

Alexander Kaltenböck and Nicolaas Koster are co-founders of Proxywise AI.