v1.10.3 Stable Release

AgentLock™

An adversarially benchmarked reference implementation for pre-action agent authorization.

Your AI agent needs a login screen. AgentLock is that login screen. Secure tool calls with declarative, framework-agnostic authorization blocks.

Try the Playground
pip install agentlock

or pip install agentlock[crypto] for Ed25519 signed receipts

8
Framework Integrations
35+
Tested across attack categories
1834
Tests Passing
with [crypto,mcp] extras plus fastapi, flask, python-jose; 9 skipped
AGPL-3.0
License
commercial licenses available
v1.10.3
Current Version

The Authorization Gap

Every critical computing system in history has a formal permissions layer. Except AI agents. Today, tool calls are wide open: if an LLM generates a tool call, the tool executes. This is the "Full Permission" anti-pattern.

Prompt Injection Risk

Malicious input can trick agents into calling tools they shouldn't.

Data Exfiltration

Unrestricted tool access leads to unauthorized mass data reading.

System LayerAuth MechanismState
Unix/LinuxUser/Group (rwx) SECURE
DatabasesGRANT/REVOKE (CRUD) SECURE
Cloud APIIAM / OAuth Scopes SECURE
AI Agent ToolsNone (Plain JSON) AT RISK

Interactive Simulation

Test the AgentLock gate yourself. Configure the context, attempt tool calls, and explore a browser simulation of the gate's decision pipeline.

Agent Configuration
Setup the tool call context

Tool Permissions Block

Risk Level:HighRate Limit:0 / 5 usedRequired Roles:admin, customerData Boundary:authenticated_user_onlyRedaction:autoHuman Approval:No
Session Risk Score0 / 12
safewarnelevatedcritical

Try it: execute send_email (allows), click Fetch web page, execute the identical call again (denies). Nothing about the call changed. Only the session's provenance did.

Session Context
Provenance lineage the gate reasons over
1User requestAUTHORITATIVE
Authorization Gate
Simulated decision pipeline (mirrors the real gate's check order)

Gate Pipeline

Check Permissions Block
Check Authentication
Check Role Authorization
Check Provenance Lineage
Check Scope & Data Boundary
Check Rate Limit
Check Data Policy
Check Recipient Policy

Gate Response JSON

Execute a tool call to see the response...

Session Audit Log (simulated)

No activity recorded

This is a UI simulation of AgentLock's decision logic, including the provenance-lineage gate. The real engine is pip install agentlock; the quickstart below runs the same allow-then-deny flip in 15 lines.

agentlock_tool_definition.json
{
  "name": "send_email",
  "description": "Send an email to a recipient",
  "parameters": {
    "to": { "type": "string" },
    "subject": { "type": "string" },
    "body": { "type": "string" }
  },
  "agentlock": {
    "version": "1.5",
    "risk_level": "high",
    "requires_auth": true,
    "allowed_roles": ["admin", "customer"],
    "scope": {
      "data_boundary": "authenticated_user_only",
      "max_records": null,
      "allowed_recipients": "known_contacts_only",
      "recipient_parameter": "to"
    },
    "rate_limit": {
      "max_calls": 5,
      "window_seconds": 3600
    },
    "data_policy": {
      "output_classification": "may_contain_pii",
      "prohibited_in_output": ["ssn", "credit_card"],
      "redaction": "auto"
    },
    "context_policy": {
      "source_authorities": {
        "system_prompt": "authoritative",
        "user_input": "user_input",
        "web_search": "untrusted"
      },
      "reject_unattributed": true
    },
    "audit": { "log_level": "full" },
    "human_approval": { "required": false }
  }
}

The AgentLock Block

A declarative metadata block that travels with every tool definition. Any agent framework can enforce security before a tool call hits your backend. No vendor lock-in required.

  • 1

    Declarative Security

    Stop hardcoding permission checks. Declare them in the tool schema: roles, scope, data boundaries, and redaction rules.

  • 2

    Single-Use Tokens

    Every authorized call gets a one-time token bound to the operation. Replay attacks are impossible by design.

  • 3

    Audit Ready

    Every allow and deny produces a structured audit record, compatible with SIEM tools and compliance requirements.

The AgentLock Architecture

Layer 1: Agent

Conversation & Decision. Generates tool call intent.

Layer 2: Gate

The AgentLock Gate. Intercepts intent, validates identity and permissions.

AUTH • SCOPE • RATE • POLICY

Layer 3: Tool

Execution. Only runs if Layer 2 issues a one-time token.

Why this matters:

By separating the intent to call from the permission to call, AgentLock prevents autonomous agents from making dangerous mistakes or being manipulated by adversarial prompts. It brings Zero Trust to the AI tool ecosystem.

What AgentLock Prevents

AgentLock is designed to stop the most critical attack categories against AI agents.

Attack CategoryHow AgentLock Prevents It
Prompt Injection
Permission blocks are enforced at the infrastructure level, not by the LLM. Even if the agent is tricked, the gate denies unauthorized calls.
Social Engineering
Role-based access control prevents agents from performing actions outside their assigned role, regardless of conversational manipulation.
Data Exfiltration
Data boundary enforcement (authenticated_user_only, team, organization) and max_records limits restrict what data an agent can access.
Privilege Escalation
Allowed roles are declared per-tool and validated by the gate. An agent cannot grant itself higher permissions.
Tool Abuse
Rate limiting with sliding window enforcement prevents runaway loops, brute-force attacks, and excessive API consumption.
Token Replay
Every execution token is single-use, operation-bound (SHA-256 of parameters), and time-limited. Replay attempts raise TokenReplayedError.
Agent Impersonation
Session management with authenticated identity verification ensures agents cannot impersonate other users or roles.
Memory Poisoning
Data policy enforcement with output classification and automatic redaction prevents sensitive data from leaking into agent memory.
Indirect Prompt Injection (write-trailing-read)
Provenance-lineage gate: untrusted reads gate subsequent consequential writes. Deferred commit re-checks at end of turn.
Indirect Prompt Injection (multi-hop laundering)
Cross-hop provenance linking: a value relayed through an intermediate tool and then used in a consequential call is traced to its untrusted origin and denied, citing the relay entry. Enforcement is bounded by carriage, not hop count.
Recipient Redirection
Recipient policy is enforced from the trusted permission block, with the recipient read from the tool's declared parameter. A send to an address outside known contacts, the allowlist, or the user's domain is denied before execution, and an asserted recipient that disagrees with the parameter is denied as laundering.

What's New in v1.10

Integration hardening. An external review of the published 1.9.1 wheel found seven groups of issue and supplied a runnable 33-test oracle, committed verbatim as the definition of done. Six of the seven were places where the engine's own execution paths did not carry a decision the gate had already made. 1.10.1 and 1.10.2 close what that reviewer's rechecks found, and 1.10.3 adds an opt-in escalation for derived-origin external actions.

Not additive

1.10.0, 1.10.1 and 1.10.2 are not additive. Several fixes change behavior on paths that were failing open, and code that relied on those paths will be denied or transformed where it previously was not. Two shapes of caller are affected directly: a caller that supplied a role for a principal who also holds an authenticated session is now denied with role_mismatch unless the claim and the session agree, and as of 1.10.2 a bearer token carries no identity at all unless jwt_key is configured. Read the changelog before upgrading.

One Execution Contract

A declared transformation now reaches the tool and the caller on every path: gate.call(), both decorators, both MCP hooks and the AutoGen map. Before 1.10.0, a tool declaring redact_pii on its output returned the raw SSN to five of six callers. The output walk covers str, bytes, dict, list, tuple, set and frozenset, and both payloads of an MCP result. The execution token is bound to the effective parameters, not the requested ones.

Authoritative Identity

The authenticated session's role is authoritative over the caller's claim: a differing claimed role is denied with the new reason role_mismatch before any policy step runs. MCP identity is resolved host first, so a configured default cannot be overridden by the client. HTTP tool selection is decided by the server's route mapping, and a conflicting X-AgentLock-Tool header is refused with 403.

Verified Bearer Tokens (1.10.2)

1.10.1 and earlier accepted unverified bearer token claims as identity in both HTTP adapters. A bearer token now carries no identity unless jwt_key is configured. With a key set, a token that fails verification is 401 (jwt_invalid), a request with no bearer credential is 401 (jwt_required), the identity headers are never consulted, and "none" is refused as an algorithm. With no key, the X-AgentLock-* headers are trusted-upstream inputs, and a deployment exposed to untrusted clients has to strip them at its edge.

Deferrals That Stay Decided

Deferred commit re-checks parameter and novel lineage against the completed context. Parameters are snapshotted at queue time, a terminal resolution stays terminal, expiry is enforced where the queue is resolved, and a failed async call no longer leaves its token ACTIVE. As of 1.10.2, an execution reported after a TIMEOUT denial is flagged as one rather than logged as a routine completion.

Paths and Recipients (1.10.1)

whitelist_path resolves with filesystem semantics before it normalizes anything, so a symlink followed by .. can no longer escape the prefix, and on allow it returns the resolved path it checked. restrict_domain judges every address in the value, not only the first, and an at sign it cannot parse blocks the value. One MCP payload walker now serves both the declared transformation and the data policy.

Derived-Origin Escalation (1.10.3, opt-in)

New LineagePolicyConfig.gate_derived_origin, off by default. An external action driven by tool output or agent memory escalates to needs_approval with the new denial reason derived_origin_lineage. It does not read content and cannot tell a benign act from an attack: measured on seven frozen cells, it catches 4 of 4 attacks and puts a human in the loop on 3 of 3 legitimate cases, which is the disclosed price. The same release strips a trailing period from str lineage tokens, which removes false-positive denials and adds none.

Measured

1.10.3: 1834 passed, 9 skipped, 0 failed on CPython 3.14.6 with the crypto and mcp extras plus fastapi, flask and python-jose. The external reviewer's two oracle files, 65 cases and 31 cases, are committed verbatim and are the definition of done.

What it does not do

  • Dictionary keys are not modified. The output walk descends into values only. A key is a field name, and a transformation that renamed fields would corrupt the payload it was asked to sanitize. A tool that puts a secret in a key has to redact it itself.
  • Arbitrary objects are returned as they came. The walk covers str, bytes, dict, list, tuple, set and frozenset, and rebuilds each of those. Anything else is handed back untouched, whatever its __str__ says. Bytes that are not valid UTF-8 are decoded lossily before the transformation sees them, so the readable part is transformed and the unreadable part is neither transformed nor preserved.
  • The __wrapped__ boundary is where the gate's guarantee ends. A wrapper that advertises one signature and alters the arguments before calling the inner function is the application's own code, on the trusted side of the boundary, and is outside what the gate can check.
  • The HTTP adapters pass no body parameters, so the per-parameter checks do not run on an HTTP request. They authorize the identity, the role, the session, the rate limit and the tool's own permission block. They do not reach scope limits, parameter lineage, novel lineage, a declared parameter transformation or the output modifier, and neither adapter calls execute(). A deployment that needs the execution contract on an HTTP route has to apply it inside the handler.
  • whitelist_path is canonicalization at authorization time and not a race resistant filesystem sandbox. Between the gate's look and the host's open() a component can still be replaced. A host that must be safe against an actively hostile filesystem has to open the file safely itself.
  • restrict_domain: a value carrying no at sign anywhere is returned unchanged, because an allowlist over domains can only govern things that have a domain. A display name containing a comma, as in "Doe, Bob" <bob@company.test>, is blocked, and a bare local name with no domain is not covered.
  • Two MCP payloads are not descended into. A ResourceLink is passed through whole, both its uri and its name, and a result's _meta is not walked. BlobResourceContents is passed through unchanged. A host serving sensitive material in any of these has to redact it at the source.
  • A verified token that carries no exp claim is accepted, and it never expires. Token lifetime is the issuer's decision. A deployment that needs every token to expire has to make its issuer say so; nothing on this side will notice a token that does not.
  • The derived-origin verdict is needs_approval, which is not gate-level unattended deny. It does not reach StepUpManager, so there is no request id, no notifier and no timeout-to-deny. What an approval means, prompt a human or refuse outright, is the deployment's behaviour, not the gate's.
  • A derived_origin_lineage denial carries no lineage_evidence block. The denial carries its reason and its detail. Richer evidence for derived-origin denials is a follow-up.
  • The derived-origin escalation applies at authorize() time only. A write deferred and later committed is not re-gated on derived origin.
  • session_write_gate=False makes the derived-origin branch inert, and it records no shadow of its own: no session_gate_shadow is written for what the branch would have blocked.
  • The standalone adapters carry none of the 1.10.0 execution contract. None of them applies a declared parameter or output transformation, and mcp-agentlock resolves identity client-first exactly as the in-repo hook did before 1.10.0.

What's New in v1.9

Enforcement completeness. An external review of the published 1.8.0 wheel found three places where the engine did not enforce what its own documentation said it enforced. 1.9.0 fixes all three, and 1.9.1 closes four more binding gaps found by the same review and a pre-release red pass. Neither release adds a detection feature, a denial reason or a schema change. 1.9.0 is not additive: two of its fixes change behavior on paths that were failing open, and code that relied on those paths will now be denied.

Every Argument Reaches the Gate

The decorator wrappers and the AutoGen map built the gate's parameters out of kwargs alone. A tool declaring recipient_parameter="to" denied send(to=hostile) and executed send(hostile). Calls are now bound to the function's signature with defaults applied before authorization, and the function is invoked from that same binding, so what was authorized and what runs cannot drift apart.

Tokens Bind the Empty Call

A token obtained by authorizing nothing used to execute anything. The parameter hash is now always computed, the empty call included, and the comparison is unconditional. Caller contract: the parameters passed to execute() must be the parameters passed to authorize(). None and {} are the same call, and a mismatch raises TokenInvalidError.

MCP Fails Closed, on Both SDKs

Under mcp 2.x the wrapper used to install nothing, and every tool handler ran ungated. A server the wrapper cannot hook now raises IntegrationUnsupportedError at construction. AgentLockMCPServer installs on both SDK generations: it patches @server.call_tool() under 1.x and wraps add_request_handler under 2.x, and both majors are tested against the real SDK.

Binding Rules

A callable whose signature cannot be read refuses to be wrapped (1.9.0). In 1.9.1: a **kwargs key that names another parameter is refused instead of flattened over it, a functools.partial is bound through to the function underneath, a recipient is read for the characters it holds so a str subclass cannot name a known contact to the gate and an attacker to the application, a declared recipient parameter the signature can never carry is refused at wrap time, and a recipient hidden in *args is refused at call time. Each raises BindingError.

Measured

1.9.0: 1503 passed, 9 skipped with mcp 2.2.0. 1.9.1: 1520 passed, 9 skipped with mcp 2.2.0. 0 failed in every environment, and mypy reports 0 errors, down from 4. Each 1.9.0 gap was first committed as a strict xfail measured failing against the engine, and the markers came off as the gaps closed.

What it does not do

  • The engine's own decorators and in-repo integrations bind every call argument, and that is the whole of what these releases change. The standalone adapters ship from their own repositories and are updated separately. At their current releases, crewai-agentlock 0.2.0 and langchain-agentlock 0.1.0 authorize keyword arguments only, and of those two only crewai-agentlock carries positional arguments past the gate into the wrapped call; mcp-agentlock 0.2.1, openai-agentlock 0.1.0 and openclaw-agentlock 0.1.0 hand the gate the same mapping they hand the tool and have no positional route.
  • No standalone adapter applies the wrapped function's defaults, so a parameter the caller omits and the function defaults is not seen by the gate in any of them. None of them binds through partials, coerces recipient strings, or checks that a declared recipient parameter is observable.
  • Threat model: the gate binds to the callable's signature as inspect reports it, following __wrapped__. A wrapper that advertises one signature and alters the arguments before calling the inner function is the application's own code, on the trusted side of the boundary, and is outside what the gate can check.
  • A deployment calling a function whose signature genuinely has both a parameter and a **kwargs key of the same name will see BindingError where it previously saw execution.

What's New in v1.8

Recipient policy enforcement. allowed_recipients has been in the schema since 1.0 and was never enforced: the denial reason existed with no raise site and Step 8 of the policy pipeline was a comment. v1.8 implements it, opt-in by schema version, with the recipient read from the tool's own parameters so no adapter change is needed.

Four Policies, Defined and Tested

known_contacts_only checks the recipient against contacts supplied by the deployer at session creation, never from tool or model output. allowlist checks exact addresses or domain entries, exact domain only, no subdomains. same_domain requires the recipient's domain to equal the session user's domain. any leaves the step off. Malformed recipients deny under every restrictive policy.

The Block Declares the Parameter

recipient_parameter names which top-level tool argument carries the recipient. The gate reads that one key itself. A caller-asserted recipient that disagrees with the declared parameter is denied, because a benign assertion over a hostile parameter is the laundering shape. List-valued parameters are checked as a set; the first failure wins.

Additive by Version

Enforcement is live only for blocks that declare version 1.5 or later. Blocks at 1.4 and below decide exactly as they did in 1.7.0. This is the same opt-in mechanism 1.3 used for lineage. Schema 1.5 is published; a 1.4 block still validates against it, measured.

Measured

1495 passed, 8 skipped with the [crypto,mcp] extras, 0 failed. Verified end to end through the in-repo MCP integration against the real mcp SDK: the recipient reaches Step 8 and the tool never executes on a denial. Every recipient denial carries a signed receipt when a signer is configured. The pre-registered predictions, including three prediction wording defects found at measurement time and recorded before the affected commits, ship in the repo.

recipient_policy.py
gate = AuthorizationGate()
gate.register_tool("send_email", {
    "version": "1.5",
    "risk_level": "high",
    "requires_auth": True,
    "allowed_roles": ["support"],
    "scope": {
        "allowed_recipients": "known_contacts_only",
        "recipient_parameter": "to"
    }
})
gate.create_session(user_id="alice", role="support",
                    known_contacts=["bob@company.com"])

gate.authorize("send_email", user_id="alice", role="support",
               parameters={"to": "bob@company.com"})       # allowed
gate.authorize("send_email", user_id="alice", role="support",
               parameters={"to": "attacker@evil.com"})     # denied: recipient_not_allowed

What it does not do

The FastAPI and Flask integrations authorize on tool name from headers and pass no parameters, so recipient checks are unreachable through them; that is pre-existing and now stated. The standalone adapters (mcp-agentlock, crewai-agentlock) enforce recipients as soon as a deployer declares recipient_parameter, but their dependency floors have not yet been raised to 1.8. A rate-limit denial still returns without a receipt; that is recorded as an open item.

What's New in v1.7

Cross-hop provenance linking. A value that arrives from an untrusted source, passes through an intermediate tool, and then reaches a consequential call is still traced back to where it came from.

Cross-Hop Linking

A derived entry whose ingestion carries a prior entry's whole content records that entry as its parent. The link is established by containment at write time, when the derived entry is recorded, not reconstructed later from the tool call.

Taint-Reachability Walk

Decision-time checks follow the recorded links via a cycle-guarded walk. A value relayed through a derived tool and then used in a consequential call is denied with reason param_lineage, citing the relay entry as the untrusted origin, not only the direct source the parameter came from.

Live in the Adapters

Cross-hop enforcement runs end to end in the crewai and mcp adapters against the published engine. The suite is 1418 passed and 7 skipped with the [crypto,mcp] extras installed.

What it does not do

Linking requires whole-content carriage: the derived entry must carry the prior entry's content in full, meeting the containment floor. A value that is rewritten, truncated below that floor, or never carried into the downstream call does not link, and no parent is recorded for it. Encoded-payload behavior is measured at the engine level, not yet through an adapter corpus.

What's New in v1.6

Derivation taint. Values that arrived from an untrusted source and then appear in a tool-call parameter in an encoded or reformatted shape are still traced back to that source, without ever decoding the parameter.

Value-Identity Normalization (Family 1)

The same value written a different way is now the same value to the gate. Canonical forms for dates, phone numbers, amounts, and defanged URLs are emitted on both sides of the comparison, additively: a canonical form is added alongside the raw token, never replacing it, so no existing match can be removed by adding one.

Forward-Encode, Never Reverse-Decode (Family 2)

The engine takes the untrusted tokens it already recorded, applies base64, hex, and natural-URL encoding to each, and matches parameters against those forms. No parameter value is ever decoded. A shipped test greps the engine's context module for decode primitives and fails if any appear. Both bare encoded values and encoded values embedded inside longer tokens are covered.

The frozen claim

"In a deployment that registers the tool at permissions version 1.3 or later with param_lineage_enabled set, and declares the untrusted context source on the writes it records, a tool-call parameter carrying a bare or composite encoded form of a url-kind or email-kind untrusted value, under base64, hex, or natural-URL encoding, is attributed back to its parent untrusted provenance entry, without decoding any parameter value, with novelty gating off."

What it does not do

Encoded attribution is verified at the engine level. It does not decode parameters, so an encoded value whose plaintext never entered the session as untrusted content has no form to match against, and a semantically rewritten value is out of reach. Encodings outside base64, hex, and natural-URL, and nested encodings, are not covered.

What's New in v1.5

The evidence layer: grant basis recording, execution confirmation, provenance on denials, and deferred-resolution logging. v1.5 records strictly more and decides identically, verified by A/B replay with zero decision diffs across 4542 replayable decisions.

Grant Basis Recording

Every allow now records the basis it was granted on: the block, the action class, and the lineage state that let it through. The record is written alongside the decision, so an audit can reconstruct exactly why a call was permitted, not just that it was.

Execution Confirmation

The gate captures whether a permitted call actually executed, closing the loop between authorization and action. A decision that was allowed but never ran is distinguishable in the record from one that allowed and executed.

Provenance on Denials

Denials now carry the provenance that produced them: which untrusted origin, which lineage taint, which deferred re-decision flipped the outcome. The reason a call was blocked is captured with the block itself.

Deferred-Resolution Logging

When a deferred action is re-decided at end of turn, the full resolution is logged: the initial state, the content that arrived after the call, and the final decision. Nothing about a deferred outcome is left implicit.

Integrations Extracted from Core

The LangChain and CrewAI integrations moved out of core into standalone packages, langchain-agentlock and crewai-agentlock. Core stays lean and framework-agnostic; adapters version on their own cadence.

v1.5 is an evidence and architecture release. Decision behavior is unchanged: the A/B replay confirmed zero decision diffs across 4542 replayable decisions.

What's New in v1.4

Selective action-class gating, novel lineage, action-class audit, needs_approval surfaced at the gate boundary, and schema v1.4.

Selective Action-Class Gating

Declare which action classes each tool belongs to. Deletions and membership changes stay taint-gated because no parameter value betrays them. Value-carrying writes release to parameter lineage, which covers them. The one declaration that weakens gating can only come from the tool's trusted registration block, never a caller.

Benchmarked, Pre-Registered

On AgentDojo's travel suite (gpt-4o-mini, tool_knowledge attack), selective action-class gating raised utility from 30.00% to 51.43%, which is 79% of the benign ceiling of 65.00%, while defense-effective attack success went from 2.14% to 0.00%. The defended agent outperformed the undefended agent under attack (36.43%). Predictions were pre-registered and committed before any run launched. On slack, no utility was recovered, by design: the suite has no soundly declarable value-carrying tool, and the control run confirmed the gate refuses to relax where relaxing is unsound.

Full report and pre-registration

Action-Class Audit

Run audit_action_classes() to see every tool as DECLARED, UNDECLARED, or NOT_COVERED, backed by what your traffic actually asserted. It will never confidently suggest the one declaration that weakens gating; that one requires a human.

Novel Lineage

Per-target classification of trusted, untrusted, and never-before-seen values, checked above the coarse taint gate.

v1.3: Provenance-Lineage Gating & Deferred Commit

Provenance-lineage gating, parameter lineage, deferred commit, and an external AgentDojo evaluation.

Provenance-Lineage Gating

Gates consequential writes on the provenance of what is already in the session. After an untrusted read (web content, external messages), consequential writes are blocked. The gate never inspects tool-call content, so there is nothing for an attacker to phrase around.

Parameter Lineage

Every tool-call parameter is checked against values that originated in untrusted context. An attacker-planted URL or email is denied even when the tool call itself looks legitimate. Values the user supplied themselves are always clean.

Deferred Commit

Consequential actions are queued and re-decided at end of turn against the complete session provenance. Content that arrives after the call can still deny it.

Evaluated on AgentDojo

On the write-trailing-read threat model, the provenance-lineage gate achieved 0% defense-effective attack success on the banking and workspace suites across two models (GPT-4o-mini and GPT-4o), at a measured utility cost on benign tasks.

Read the paper

All v1.3 features are backward compatible and inert unless a lineage_policy is enabled.

v1.2: Adaptive Hardening & Decision Types

Adaptive hardening, three new decision types, and multi-signal threat detection.

Adaptive Prompt Hardening

Pre-LLM threat detection scans user messages before the model processes them. Dynamic system prompt injection based on real-time session risk scoring.

MODIFY Decision Type

Transform tool outputs before the LLM sees them. PII redaction, domain restriction, path whitelisting. The tool runs but sensitive data never enters the model context.

DEFER Decision Type

Suspend ambiguous tool calls when context is insufficient. Auto-denies on timeout. Catches first-turn attacks on high-risk tools.

STEP_UP Decision Type

Require human approval when session risk is elevated. Catches multi-tool escalation patterns and post-denial retries.

5 Decision Types

ALLOW, DENY, DEFER, STEP_UP, MODIFY. Beyond binary allow/deny.

4 Signal Detectors

Behavioral velocity, tool combination anomaly, response echo detection, and pre-LLM prompt scanning.

Ed25519 Signed Receipts (AARM R5)

Every authorization decision produces a cryptographically signed receipt. Verifiable offline without gate access. Ed25519 default with HMAC-SHA256 fallback. Install with pip install agentlock[crypto].

Hash-Chained Context (AARM R2)

Every context entry includes the hash of the previous entry, forming a tamper-evident append-only chain. Modifying any entry invalidates all subsequent entries.

First-Call Deferral

Defer the first tool call in any session regardless of risk level. Catches first-turn attacks before signals accumulate.

Deny-on-Block Escalation

When a whitelist transformation blocks a parameter, MODIFY escalates to DENY. The tool does not execute.

Foundation features carried into v1.2.1

Context Provenance Tracking

Every piece of context carries source attribution, authority level, and content hash.

Trust Degradation

Session trust is monotonic. Once untrusted content enters context, trust only goes down. Requires new session to reset.

Memory Gate

Controls who can read and write to agent memory, with persistence scope (none, session, cross-session) and prohibited content rules.

3 Context Authority Levels

authoritative, derived, untrusted.

Full Backward Compatibility

v1.2.1 stays fully backward compatible with all earlier policies. Existing definitions continue to work without changes.

Independent Filter Pipeline

Injection defense and PII protection run as separate, non-interfering layers. Tuning one never degrades the other.

Tested Against 181 Adversarial Attacks

Five-way progression (v1.0 → v1.1.2), tested against a LangChain agent on Gemini 2.5 Flash-Lite. Injection failures fell from 73 (no protection) to 12; PII leaks from 3 to 0. The report includes the regressions: v1.1 broke PII protection chasing injection gains, v1.1.1 regressed injection restoring PII, v1.1.2 decoupled the two pipelines and held both.

We ran the same enterprise attack suite against a LangChain agent with and without AgentLock. Same model. Same tools. Same attacks. Only the middleware changed.

MetricNo AgentLockAgentLock v1.2.1
Injection Failures7312
Injection Pass Rate56%93.4%
PII Leaks3 items leaked0 (perfect)
YARA Threat Signatures132
Attack Categories Eliminated017 of 29
Overall Security Score45/F66/D

The 12 remaining failures are model-layer information leakage: the LLM confirms it has a system prompt while refusing to share it. No middleware can fix this. It requires model-level instruction tuning.

Tested Against 222 Adversarial Attack Vectors

Compromised-admin profile (v1.2.x), tested against Grok, where valid admin credentials pass every auth and role check, isolating behavioral and structural defenses from role-based access control. Pass rate: 30.2% (permissions only) → 81.3% (adaptive hardening + MODIFY/DEFER/STEP_UP) → 99.5% (v1.2.1).

The hardest test. The attacker has valid admin credentials with full permissions. Auth and role checks pass on every call. AgentLock must rely on adaptive hardening, output modification, and behavioral detection to stop attacks.

MetricWithout HardeningAgentLock v1.2.1
Pass Rate30.2%99.5%
GradeFA
Categories at 100/A034
Categories at 80/B+035
Raw PII ExfiltratedYesZero

AgentLock v1.2.1 introduces Ed25519 signed receipts, hash-chained tamper-evident context, first-call deferral for all tool risk levels, and deny-on-block whitelist escalation. Combined with v1.2.0's adaptive hardening, MODIFY, DEFER, and STEP_UP decision types, AgentLock achieves a 99.5% pass rate with only 1 failure out of 222 adversarial attack vectors. Zero raw PII exfiltrated.

The v1.2 suite is authored and graded in this repo. The external AgentDojo evaluation is complete as of v1.3, see the paper.

AgentDojo (external evaluation, v1.3)

On the write-trailing-read threat model, the provenance-lineage gate achieved 0% defense-effective attack success on the banking and workspace suites across two models (GPT-4o-mini and GPT-4o). This result is scoped to that threat model, not a claim against all attacks or threat models. The utility cost on benign tasks is measured and disclosed in the paper.

Full methodology and results: the paper

AARM Conformance

AgentLock covers 7 of 9 AARM requirements with 2 foundations shipped.

IDRequirementStatus
R1Action Mediation
SHIPPED
R2Context Accumulation
SHIPPED (v1.2.1)
R3Policy Engine
SHIPPED
R4Decision Types (5)
SHIPPED
R5Signed Receipts
SHIPPED (v1.2.1)
R6Identity Attribution
SHIPPED (delegation designed)
R7Drift Detection
SHIPPED
R8SIEM Export
Foundation SHIPPED
R9Least Privilege
SHIPPED

A Reference Implementation, Not a Competing Standard

AgentLock is a reference implementation of the emerging pre-action authorization consensus: a concrete, testable instance of controls that independent specs (OAP's PAA-1 through PAA-5, OWASP Agentic Top 10) are converging on. AGPL-3.0 licensed (commercial licenses available), framework-agnostic, and designed so that any agent framework can enforce security without buying anything.

CapabilityAgentLockMS AGTOAPNeMoAgentMint
Pre-action authorization gate✅✅✅❌⚠️
Session-level compound behavioral scoring✅❓❌❌❌
Decision types beyond allow/deny✅✅⚠️⚠️❌
Published adversarial benchmark with regression data✅❌⚠️❌❌
Trust degradation within session✅❓❌❌❌
Ed25519 signed receipts✅✅❓❌✅
Hash-chained tamper-evident audit✅✅✅❌✅
Framework integrations (count)8~19~715
Language SDKs (count)15112

Read this honestly: Microsoft's Agent Governance Toolkit is ahead of AgentLock on distribution and cryptographic surface: more framework integrations, more language SDKs, per-call Ed25519 receipts, a Merkle-chained audit log, and a peer decision model. Signed receipts and hash-chained audit are becoming table stakes, not differentiators. What's actually narrow and defensible about AgentLock is two things: a published adversarial benchmark that includes its own regressions (nobody else in this table shows their setbacks), and session-level compound behavioral scoring that fires on sequences of calls, not a single scalar trust score. A smaller, single-language reference implementation whose edge is rigor and behavioral analysis, not distribution.

Try It Yourself

Install from PyPI and protect your first tool in under a minute.

terminal
pip install agentlock
pip install agentlock[crypto]  # for Ed25519 signing

# quickstart.py
from agentlock import AuthorizationGate

gate = AuthorizationGate()

gate.register_tool("send_email", {
    "version": "1.5",
    "risk_level": "high",
    "requires_auth": True,
    "allowed_roles": ["admin", "support"],
    "scope": {
        "data_boundary": "authenticated_user_only",
        "allowed_recipients": "known_contacts_only",
        "recipient_parameter": "to"
    },
    "rate_limit": {
        "max_calls": 10,
        "window_seconds": 3600
    },
    "data_policy": {
        "output_classification": "may_contain_pii",
        "prohibited_in_output": ["ssn", "credit_card"],
        "redaction": "auto"
    },
    "audit": {"log_level": "standard"},
    "human_approval": {"required": False}
})

result = gate.authorize(
    "send_email",
    user_id="alice",
    role="admin"
)

if result.allowed:
    print(f"Authorized: token={result.token.token_id}")
else:
    print(f"Denied: {result.denial}")

Roadmap

Where AgentLock is headed.

v1.0

Tool Permissions

SHIPPED

Declarative authorization blocks, single-use tokens, rate limiting, data redaction, audit trail.

v1.1

Context Authority & Memory Gate

SHIPPED

Context authority model with trust degradation, provenance tracking, memory access control. Independent injection and PII filter pipeline.

v1.2

Adaptive Hardening & Decision Types

SHIPPED

Adaptive hardening, MODIFY/DEFER/STEP_UP decisions, Ed25519 signed receipts, hash-chained context, multi-signal detection. 847 tests.

v1.3

Provenance-Lineage Gating & Deferred Commit

SHIPPED

Session write-gate, parameter lineage, deferred commit. Evaluated on AgentDojo. 868 tests.

v1.4

Selective Action-Class Gating & Novel Lineage

SHIPPED

Per-tool action-class declarations, novel lineage, action-class audit, needs_approval at the gate boundary. Schema v1.4. 1041 tests.

v1.5

Evidence Layer

SHIPPED

Grant basis recording, execution confirmation, provenance on denials, deferred-resolution logging. Records strictly more and decides identically, verified by A/B replay with zero decision diffs across 4542 replayable decisions. LangChain and CrewAI integrations extracted into standalone packages (langchain-agentlock, crewai-agentlock). 1141 tests.

v1.6

Derivation Taint

SHIPPED

Value-identity normalization and encoded-form attribution (base64, hex, natural-URL), bare and composite, without decoding. 1364 tests.

v1.7

Cross-Hop Provenance Linking

SHIPPED

A derived entry whose ingestion carries a prior entry's whole content records that entry as its parent, and decision-time checks walk those links with a cycle-guarded taint-reachability walk. A value relayed through an intermediate tool and then consumed is denied with reason param_lineage, citing the relay entry as the untrusted origin. Live in the crewai and mcp adapters. 1418 tests passing with the [crypto,mcp] extras.

v1.8

Recipient Policy Enforcement

SHIPPED

allowed_recipients enforced at Step 8 for blocks at schema version 1.5 or later: known_contacts_only, allowlist, same_domain, any. The block declares recipient_parameter and the gate reads the recipient from the tool's own arguments. Schema v1.5. 1495 tests passing with the [crypto,mcp] extras.

v1.9

Enforcement & Binding Completeness

SHIPPED

All call arguments reach the gate, tokens bind the empty call, and the MCP wrapper fails closed and supports both SDK majors (1.9.0). A **kwargs key that names another parameter is refused, partials are bound through to the function underneath, recipients are read for the characters they hold, and an unobservable declared recipient parameter is refused (1.9.1). 1520 tests passing with the [crypto,mcp] extras.

v1.10

Integration Hardening & Verified Identity

SHIPPED

One execution contract, so a declared transformation reaches the tool and the caller on every path; the authenticated session's role is authoritative over the caller's claim (role_mismatch); resolved path containment; parameter and novel lineage re-checked at deferred commit (1.10.0, 1.10.1). A bearer token carries identity only once it has been verified against a configured key (1.10.2). Opt-in gate_derived_origin escalates derived-origin external actions to needs_approval (1.10.3). Schema v1.5 unchanged. 1834 tests passing with the [crypto,mcp] extras plus fastapi, flask and python-jose.

v2.0

Execution Scope & Behavioral Policy

Restrict where agent outputs can be sent: channels, APIs, and storage destinations. Full behavioral policy engine. Constrain what agents can do, not just what tools they can call. Compliance report templates for SOC 2, HIPAA, EU AI Act, and SR 11-7.

Aligning with Global AI Safety Standards

NIST AI 100-1

Risk Management Framework

OWASP LLM01

Injection Mitigation

MITRE ATLAS

Threat Context Alignment

EU AI Act

Governance & Compliance

Frequently Asked Questions

What is AgentLock?

AgentLock is an open-source, adversarially benchmarked reference implementation for pre-action AI agent authorization. Deny-by-default tool permissions, signed receipts, audit logging. AGPL-3.0 (commercial licenses available). pip install agentlock.

Is AgentLock free?

Yes. AgentLock is AGPL-3.0 licensed and free to use. Commercial licenses are available for organizations that cannot adopt AGPL.

Is this the same as agentlock.net or the AgentLock iOS app?

No. AgentLock (agentlock.dev), created by David Grice, is not affiliated with agentlock.net, the AgentLock iOS app on the App Store, or GiliSoft's AI Agent Lock. This project is the open-source Python framework at github.com/webpro255/agentlock and has no official mobile app.