Blog · April 23, 2026

Volatility in AI Proxy Voting Is an Architecture Problem

A response to the Kekst CNC and Glass Lewis reports on AI in proxy voting

By Alexander Kaltenböck and Nicolaas Koster, Proxywise AI

Two reports on AI in proxy voting landed in April 2026: both with merit, both making assumptions and getting to conclusions we are challenging.

Kekst CNC's AI as the New Proxy Advisor tested four general-purpose LLMs on 49 contested U.S. director elections. Glass Lewis's AI and the Fiduciary Test, published the same week, set out what fiduciary-grade AI should look like. Neither answers what stewardship teams are actually solving: how do you apply your own voting policy, consistently and defensibly, to every ballot, without outsourcing the judgment?

What Kekst measured, and what the framing misses

Kekst's findings are striking yet unsurprising: across 49 contests the four LLMs backed activists or split tickets about 63% of the time (versus 41% actual), flipped their recommendation in 38% on repeat prompts, and cited press releases in nearly a third of sources.

Comparing four LLMs on the same ballots treats the model as the key variable. For naive prompting, that holds. In a production workflow, architecture matters more than model: what policy constrains it, what data it sees, what it must cite. Volatility, press-release reliance, and pro-activist tilt are predictable when a probabilistic model freestyles without a framework. "AI is pro-activist and volatile" really means "naive LLM prompting is."

Some pro-activist skew is real. Calibration is the remedy: examples, overrides as rule-linked precedents, back-testing. The 38% flip rate, though, is exactly what the right architecture eliminates. As we argued in the Harvard Law School Forum in March, stable output is a property of traceable policy-first systems. Across our 2026-season pilots, repeat runs of routine items produce identical recommendations: effectively zero variance. Repeat-run stability, how often a system reaches the same conclusion on the same ballot, should become an industry KPI. Kekst set the baseline at a 38% flip rate across naive LLM prompting. We will publish ours transparently, and invite other providers to do the same.

The Glass Lewis framework

Glass Lewis draws a line the industry has needed. Fiduciary-grade AI cannot rest on human review at the output stage. Approach A, AI-first with output review, inherits every LLM risk: hallucination, broad-prompt drift, opacity. Approach B, in which domain experts design methodology, data schema, and decision rules up front and AI operates inside that governed structure, is the correct standard. We agree. Regulation converges the same way: the EU AI Act, U.K. and revised Japan Stewardship Codes, and SEC guidance all require oversight by design.

A policy-first system, with policy codified into machine-executable rules, AI applying them to source disclosure, every assessment traceable, is Approach B. The policy is the foundation; the rules, refined by stewardship experts, are the architecture. The AI operates inside that frame, eliminating most of the "black box" behavior that worries people about LLMs. We also separate evidence-first analysis from the vote decision: the system identifies policy criteria and evidence before synthesizing a recommendation, checking premature-conclusion bias. Both AI operations, extraction and reasoning, run inside machine-executable rules and the asset manager's codified policy.

Where our view parts from Glass Lewis is not on architecture, but on whose expertise governs it. Their implicit answer: outsourced governance experts. An expected pitch for a major proxy advisor defending the status quo. Institutional stewardship teams want the opposite: AI that applies their own policy and analyst judgment to every ballot, informed by external research where it adds value. This is the purest definition of fiduciary duty, and AI finally makes it workable at portfolio scale.

A workflow-level view: three buckets

Ballot volume breaks into three buckets, and the split depends on the maturity of the investor and its stewardship expertise. Across more than 88,000 analyzed US ballots between 2023 and 2025, from one of our most policy-mature, sophisticated clients, the split has run 80% routine, 18% non-routine requiring some analyst calibration, 2% hard cases the analyst owned end to end.

  1. Routine items (uncontested director elections, standard comp, auditors). Against a policy harness, the architecture is close to deterministic. Smaller, cheaper models match frontier models here. Oversight is upstream in the rules.
  2. Non-routine with a principled answer: proposals needing context from prior engagement, voting history, or precedent. AI retrieves stored context and proposes a recommendation with evidence. The analyst engages before the vote.
  3. The genuinely hard cases: contested elections, high-stakes say-on-pay, first-of-kind proposals, complex M&A. The analyst owns the judgment; AI supports research and evidence retrieval. Every contest in Kekst's dataset lives here, and no stewardship team we have met would delegate these to ChatGPT or any other chatbot.

What good looks like: two more questions to ask

Glass Lewis's five questions (on data governance, human role, investment-grade data, designed scope, and exception handling) are the right ones. We would add two.

Sixth: whose policy is the AI applying? For a fiduciary, the answer must be the asset manager's own codified policy: inspectable, editable, versioned. A proxy advisor's house view with a client overlay is not the same thing. Fiduciary accountability is not portable, and cannot be outsourced.

Seventh: is the triage itself governed? A system that cannot tell you which bucket a ballot landed in, and why, is a system whose boundaries you cannot defend.

Closing

Kekst measured one configuration: naive prompting. Glass Lewis defined one standard: fiduciary duty cannot just mean "AI quality assurance post-prompting". Between them sits the real question: whose policy, whose expertise, whose accountability, and what architecture makes AI usage in proxy voting workable at portfolio scale.

Alexander Kaltenböck and Nicolaas Koster are co-founders of Proxywise AI, which builds policy-first voting infrastructure for institutional stewardship teams. Their earlier piece from the 2026 proxy season pilots appeared in the Harvard Law School Forum on Corporate Governance in March 2026.