On August 26, we hosted the third module of our Roundtable Series, which closed the first cycle of our roundtable conversation series.
This session focused on governance and the reality of responsible agentic AI adoption in higher education: specifically, that responsible adoption requires a clear understanding of its true operational costs and a realistic blueprint for how institutions sequence implementation.
Why governance for Agentic AI is a design problem
Before exploring responsible adoption, we defined governance itself: deciding in advance what an AI system is allowed to do, who is accountable when it acts, and how you know if something goes wrong. Its core pillars are authority, accountability, and visibility.
Most AI governance in higher education today addresses teaching and learning, specifically policies on how faculty and students use AI in coursework. Far less attention goes to administrative operations, where the institution executes its own processes. This is precisely where agentic systems create exposure.
The distinction matters because the nature of risk changes:
Traditional AI Governance evaluates output risk: Is the output correct?
Agentic AI Governance evaluates action risk: What is the system authorized to do, and who owns the consequences when it acts?
This shift makes governance an architectural design decision rather than a post-hoc compliance audit. An end-of-process review only reveals what already happened; it cannot show how augmented decisions were reached.
In practice, checkpoints must sit within each step of the workflow. Using artifact scoring for Assurance of Learning (AoL) as an example, responsible design requires explicit approval points and clear ownership of AI-generated evaluation outputs:
Human Confirmation: An AI-generated score never enters the official record until a faculty member confirms or adjusts it.
Scoped Access: Access is restricted strictly to the artifact set selected for the active review cycle.
Traceable Logs: Every action is logged with user attribution, justification, and timestamping.
Dynamic Scope: System permissions are reassessed every cycle.
Verification is "New Work" in Agentic AI
An AoL committee participant summarized current practice concisely:
You have a committee, you have artifacts, you have a rubric, someone assesses them, and you’re done.
Standard processes today rarely confirm scoring consistency across assignments or courses.
This creates a critical realization for anyone evaluating AI in assessment: an AI review checkpoint does not replace an existing task—it introduces entirely new work.
In the words of one participant, adding that checkpoint brings in “100% of new involvement.” Another participant highlighted the mirror risk: institutions expect AI to save time, only to discover faculty end up with more checks to perform than before.
That has a direct consequence for AI-assisted scoring, where the AI evaluates each criterion, provides its rationale, and routes only contestable points to a human. Efficiency claims have to be benchmarked against a baseline of zero pre-existing verification, because early implementations will rely on faculty and rubric owners verifying scores regardless of how the guardrails are designed.
The true value of our upcoming pilots will be measuring how calibrated AI scores are against human graders, and whether this model ultimately yields actual time savings.
Sequencing a pilot to build evidence and secure leadership endorsement
Participants offered strategic guidance for securing endorsement from accreditation committees and senior leadership, which is worth borrowing if you are considering a similar solution in your own work:
Run in Parallel: Begin with full human scoring, generate AI scores alongside for direct comparison, and start with a single faculty member on the AoL committee before expanding to a second class and then broadening access.
Convert Trust to Data: Parallel execution shifts the conversation with skeptical faculty and leadership from an emotional debate over trust to an empirical evaluation of measurement validity.
Align with Power Structures: The committee accumulates its own record of where AI and human judgment agree and diverge, then presents that record with the Dean behind it. At one participant’s school, the Dean sits on the AoL committee, allowing them to establish institutional trust from day one.
Budget the Verification Burden: Weak sponsorship is a leading cause of digital transformation failure. The initial verification workload in Cycle 1 is an intentional design input, not a platform failure—it must be budgeted upfront rather than argued away.
Where AI earns trust first
Two low-friction applications emerged that require far less initial trust than automated scoring:
Pattern Flagging & Anomaly Detection: The AoL cycle is used as much to revise rubrics as to generate scores. A participant shared that during one cycle, ratings repeatedly came back marked “out of scope,” which revealed the committee had been assessing the wrong course entirely. AI is well suited to flagging anomalies like that and capturing rubric modifications alongside their justification, which is exactly the material an accreditation report asks for.
Supporting Non-Subject-Matter Evaluators: At one institution, committee members rather than course faculty perform the assessment, meaning the person evaluating a communication outcome is often not a communications expert. AI can provide contextual rubric guidance to support non-experts working outside their primary field.
A Related Validity Caution: One participant noted their institution’s writing measurements predated generative AI. On subsequent cycles, nearly all students exceeded benchmarks because they used AI to write. When a direct measure no longer discriminates performance, it creates a validity issue that must be addressed before introducing new AI assessment tools.
Critical structural constraints:
Peer Rating Data: Peer rating surveys from team assignments can be primary evidence for a certain competency, and should not be treated as edge cases.
Rubric Separation: Program-level AoL rubrics are distinct from grading rubrics used by course faculty; conflating the two will lead to incorrect conclusions.
What to Prepare Before Adoption Review
Legal and privacy reviews dictate the implementation timeline. Participants described privacy hurdles that can halt projects outright, ranging from restrictions on using public faculty LinkedIn data for CV collection to mandatory multi-tiered approvals for any tool touching student data.
When evaluating an AI tool that handles student work, requesting these key deliverables upfront will save time:
Vendor Security Assessment: Completed within your IT department’s required framework.
Data Flow Diagram: Clear mapping of artifact data routing, including external model providers.
Data Retention & Disposal Terms: Explicit policies detailing what happens to data and access when the pilot ends.
Model Training Exclusions: Written commitments that institutional data will never be used to train foundational models.
Scoped Access Control: Granular permissions restricting tool access strictly to selected cycle artifacts.
Comprehensive Audit Trails: Detailed logs tracking action attribution by user, step, and model version.
Responsibility ultimately rests with the deploying institution, which holds process oversight and accreditation obligations. Vendor solutions must carry these requirements through their platforms and model providers. General AI frameworks are helpful, but because none are written specifically for higher education, we are working out these requirements directly with schools.
The unresolved question: peer review visibility
Can external peer reviewers detect excessive AI involvement lacking human judgment, and what does that look like in practice?
The room agreed on the general direction: explicit rigor with human checkpoints mapped out clearly.
However, a key operational question remains: How would a peer reviewer actually detect human involvement, and is naming the report author sufficient? While AACSB is currently developing guidance on this topic, the standards remain unpublished.
That open question is why we treat the evidence trail as a core feature rather than a byproduct of scoring. If the record of who confirmed a score — and why an adjustment was made — cannot be exported in a form a peer review team accepts, the governance was never really there.
Where we're headed
This session concluded our first roundtable cycle, which we will re-run in upcoming months for new cohorts. Based on these insights, we are making three key adjustments in our own work:
Instrumenting Agreement/Divergence: Tracking AI vs. human score alignment directly to generate empirical pilot data.
Decoupling Pattern Flagging: Scoping anomaly detection and change capture as standalone capabilities exportable into accreditation reports.
Pre-empting Security Requirements: Completing full vendor security documentation prior to initial pilot discussions.
We are building this framework in the open alongside faculty, assessment directors, and accreditation leaders. If you are interested in testing these governance design principles against your institution’s Assurance of Learning workflows, we’d be interested to hear how you are approaching them.






