- Full title: Enforceable Trust for Enterprise AI Agents: A Framework for Establishing, Measuring, and Enforcing Agent Trustworthiness
- Published in: International Journal of Scientific Research in Engineering and Management (IJSREM), Vol. 10, Issue 9, September 2026
- Authors: Karthik Rajkumar Kannan, Bipin Chandra
- DOI: 10.55041/IJSREM67554
- Keywords: agentic AI, trust management, zero trust, AI governance, autonomy levels, enterprise security
Read the full paper (PDF) · View on IJSREM
#Abstract
Enterprises are moving large language model (LLM) agents from answering questions to taking actions in IT service management, finance and regulated life-sciences operations, yet trust in these agents is still granted statically. This paper proposes Enforceable Agent Trust (EAT), a framework that treats trust as a computed, decaying, context-scoped state that directly limits agent authority. EAT models five dimensions: identity, provenance, behavioural competence, authorization and intent alignment, and compliance. Identity and provenance act as cryptographic gates; the other dimensions are estimated from attested evidence with a Beta-reputation formulation that adds exponential decay and explicit uncertainty. A pessimistic composite score assigns agents to five autonomy tiers that decide which credential scopes may be issued, while regulatory ceilings and circuit breakers override the score. A simulation of 50 agents over 180 days shows that EAT reduces unsafe actions by hijacked agents by 99.6% relative to static role-based access and contains hijacks in a median of 4 h, but reacts weakly to silent behavioural degradation; a per-dimension behaviour floor reduces degradation-driven exposure by a further 66%. An OWASP Agentic Top 10 coverage analysis positions EAT as an authority layer that makes agent autonomy earned, measurable and revocable.
#The paper in plain terms
#The problem: trust that never changes
Agents no longer just answer questions. They call enterprise tools, change tickets, move money and delegate work to other agents. The risk has moved from "the model said something wrong" to "the agent did something wrong with real privileges".
Yet in most enterprises an agent is trusted the way a service account is: it is registered, given a fixed set of permissions, and trusted until someone revokes it. The paper identifies three flaws in that model:
- It ignores changing evidence. An agent's reliability shifts with model updates, prompt edits, new tools and adversarial inputs.
- It confuses identity with trustworthiness. A strongly authenticated agent can still be hijacked by a prompt injection hidden in the data it reads.
- It cannot grade autonomy. The same agent may be fit to read incident tickets alone, but not to change a validated production system.
Frameworks such as the NIST AI Risk Management Framework define what trustworthy AI means. What they do not say is how to turn that into a decision at the moment an agent calls a tool. That is the gap this paper fills.
#Five dimensions of trust
Enforceable Agent Trust (EAT) measures an agent along five dimensions, split into two kinds:
- Gates — verified cryptographically. Identity (is this really the agent it claims to be?) and provenance (is it running the model, prompt and tools that were approved?). Fail either and the agent gets nothing, whatever its history.
- Graded — estimated from evidence. Behavioural competence, authorization and intent alignment, and compliance. These are scored from attested records of what the agent has actually done.
A configuration change — a model upgrade, a prompt edit — closes the provenance gate until the agent is re-attested. The agent that earned the trust is, in a meaningful sense, no longer the agent being trusted.
#A score that fades, and knows what it does not know
Each graded dimension is estimated with a Beta-reputation model: good and bad outcomes are counted as evidence. EAT adds two things. Evidence decays exponentially, so last month's good behaviour counts for less than today's. And the score carries explicit uncertainty: enforcement uses a pessimistic lower bound, so an agent with little history cannot reach high autonomy no matter how clean that short history is.
The dimensions are then combined pessimistically into one composite score.
#Five tiers of autonomy, each enforced
The score maps to an autonomy tier, and each tier is bound to concrete enforcement — above all, which credentials the agent can be issued. An agent at the supervised tier cannot obtain a write-scoped token at all.
| Tier | What the agent may do | Human role |
|---|---|---|
| T0 Quarantined | Nothing; existing tokens are revoked | — |
| T1 Read-only | Read non-sensitive resources; drafts are never committed | Operator |
| T2 Supervised | Propose writes; every write needs explicit human approval | Approver |
| T3 Bounded autonomy | Reversible low- and medium-risk writes, within rate and value budgets | Collaborator |
| T4 Autonomous | All in-domain actions, up to the regulatory ceiling | Observer |
Three rules sit on top of the score:
- Trust is lost quickly and earned slowly. Demotion is immediate; promotion happens one tier at a time, only after the score has held above the threshold for a sustained window.
- Circuit breakers override the score. A confirmed injection attempt, an unexplained configuration change or a human stop command drops the agent's ceiling at once, regardless of how good its record looks.
- Regulation caps autonomy independently. A change to a validated life-sciences system, for example, can never exceed the supervised tier, because the regulation requires a human electronic signature. Delegation cannot amplify trust either: a child agent never outranks the agent that delegated to it.
#What the simulation showed
The framework was evaluated with a simulation of 50 agents over 180 days, repeated across 30 random seeds, against three baselines including static role-based access.
- Unsafe actions by hijacked agents fell by 99.6% — from 1,548 per run under static access to about 5 — and every hijack was contained in a median of 4 hours. Without circuit breakers, containment took 29 days.
- Scores alone react too slowly to a single severe event when an agent has a long, positive history. That is why the circuit breakers are essential rather than optional.
- A naive average score with no decay did worse than static access, because it promoted average agents too readily.
#Where it fits
Mapped against the OWASP Top 10 for Agentic Applications (2026), EAT gives strong coverage of identity and privilege abuse, supply-chain risk, insecure inter-agent communication and rogue agents, and partial coverage of risks that are semantic or human in nature, such as goal hijack detection and human-agent trust exploitation.
The paper positions EAT as the authority layer beneath content-level defences: it does not try to detect every prompt injection, but it limits what a compromised agent can do before anyone notices. Autonomy becomes something an agent earns, that can be measured, and that can be taken away.
#Cite this paper
K. R. Kannan and B. Chandra, "Enforceable Trust for Enterprise AI Agents: A Framework for Establishing, Measuring, and Enforcing Agent Trustworthiness," International Journal of Scientific Research in Engineering and Management, vol. 10, no. 9, Sep. 2026, doi: 10.55041/IJSREM67554.@article{kannan2026enforceable,
title = {Enforceable Trust for Enterprise AI Agents: A Framework for Establishing,
Measuring, and Enforcing Agent Trustworthiness},
author = {Kannan, Karthik Rajkumar and Chandra, Bipin},
journal = {International Journal of Scientific Research in Engineering and Management},
volume = {10},
number = {9},
year = {2026},
month = sep,
doi = {10.55041/IJSREM67554}
}