Skip to content
EdgeLex

The Standard · Version 1.0 · July 2026

The Legal AI Accountability Standard

AI that answers questions is a research tool. AI that does work is a colleague — and colleagues are accountable. This is a vendor-neutral standard for evaluating any AI system that performs legal work: eight principles, each with the questions a firm should put to any vendor. Including us.

Published by EdgeLex. Use it against every vendor you evaluate — EdgeLex is scored below like anyone else.

01

Provenance

Every factual claim the AI presents is traceable to a source the lawyer can open: a document span, an email, a docket entry, a statute, a verified case citation, or a ledger record. The trail is attached to the answer itself, not reconstructable on request.

Lawyers are personally responsible for what they file and advise. An answer whose basis cannot be inspected transfers risk from the vendor to the lawyer while removing the means to manage it.

Ask any vendor

  • ?Can I click from any claim in an answer to the exact source behind it?
  • ?Are sources typed (document, email, authority, ledger) or just a list of links?
  • ?Is provenance recorded for every answer, or only when a feature is switched on?
02

Refusal

When the system cannot support a claim from evidence, it says so — visibly. Unsupported material is withheld or flagged, never silently blended into confident prose.

The most dangerous AI output is the fluent paragraph with one unsupported sentence in the middle. A system that cannot show you what it removed will eventually make you its co-signer.

Ask any vendor

  • ?What happens, concretely, when the AI cannot prove a claim — is it blocked, flagged, or just written anyway?
  • ?Can the vendor show you a refusal on screen, live?
  • ?Does the system distinguish “no evidence found” from “the evidence says no”?
03

Citation integrity

Citations are verified against an authoritative corpus before they are relied on — deterministically first, with probabilistic judgment only on top, and human review for what neither can settle.

Fabricated citations are the signature failure of legal AI, and sanctions for them are now routine. Verification that depends on the same model that wrote the citation is not verification.

Ask any vendor

  • ?Is citation checking deterministic lookup against a real corpus, or the model grading itself?
  • ?What corpus, how complete, and where does it run — inside my deployment or a third party?
  • ?What happens to citations that cannot be verified — review queue, or best guess?
04

Authority

The AI acts only under explicit grants. Each task or standing instruction carries an allowlist of what it may touch; anything consequential — sending, filing, changing deadlines, moving money — pauses for human approval, every time.

An assistant that can act is an agent, and an agent without defined authority is an unsupervised signatory. Law firms already know how to grant authority — engagement letters, signing authority, supervision rules. AI should meet that bar, not lower it.

Ask any vendor

  • ?Can a task's tool access be scoped — can a summarizer be structurally unable to send email?
  • ?Which actions require human approval, and can that set be configured — or is it marketing language?
  • ?Can an approval be inspected later: who approved what, when, and exactly what executed?
05

Versioned instructions

Standing instructions are immutable once active: edits create a new version, history never mutates, and every run records exactly which version it executed.

When an automated process does something surprising, the first question is “what was it told, at the time?” If instructions can be silently edited, that question has no reliable answer — and neither does a regulator's version of it.

Ask any vendor

  • ?If I edit a standing instruction, does the old version survive with its run history intact?
  • ?Can I see, for any past run, the exact instruction text it executed under?
  • ?Who can change instructions, and is the change itself audited?
06

Work of record

AI work product lands in the firm's systems of record — the task list, the document management system, the matter file — attributed, reviewable, and billable like any team member's work. Not in a side panel that vanishes with the browser tab.

Work that is not on the record cannot be supervised, cannot be billed, and cannot be defended. If the AI's output only exists in a chat transcript, the firm has an unsupervised shadow workforce.

Ask any vendor

  • ?Does AI work appear in the same task and matter systems as human work?
  • ?Is it attributed — can you tell what the AI did versus what a person did?
  • ?Can AI effort and cost be allocated to client matters for billing or absorption decisions?
07

Boundedness & audit

Autonomous work is bounded — iteration limits, budget ceilings, progress checks, a kill switch — and everything is written to an audit trail the firm can produce: decisions, actions, approvals, and cost, attributable down to the client matter.

“It keeps working while you sleep” is only a feature if it also stops. Bounds are what make autonomy insurable; the audit trail is what makes it defensible.

Ask any vendor

  • ?What stops a runaway task — hard limits, or hope?
  • ?Is there a single place to see and stop everything the AI is currently doing?
  • ?Can you produce a complete account of AI activity on one matter, including cost?
08

Sovereignty

The firm chooses where its data and models live — a managed cloud, a private cloud, or its own infrastructure — and can run local or firm-tuned models without surrendering any of the preceding principles.

Privilege and confidentiality obligations do not stop at a vendor's terms of service. Deployment choice is what makes every other principle enforceable rather than promised.

Ask any vendor

  • ?Is self-hosting or private-cloud deployment actually available, at what fidelity versus the cloud product?
  • ?Can the firm bring or run its own models, and do governance rules still apply to them?
  • ?What, precisely, leaves the firm's chosen boundary — telemetry, prompts, documents, embeddings?

Scoring a vendor

For each principle, answer the three questions with the vendor's product in front of you — demos, not datasheets. Score the principle Met when all three hold and the vendor can show them live; Partial when the capability exists but is optional, plan-gated, or cannot be demonstrated; Unmet otherwise. A system performing substantive legal work should meet all eight. A system that scores below that can still be useful — as long as the firm treats it as an unsupervised tool and staffs the supervision itself.

How EdgeLex scores: the standard describes our architecture because we built the platform to satisfy it — provenance and refusal via evidence-gated answers, citation integrity via a deterministic authority ledger, authority via per-delegation allowlists and approval gates, versioned instructions via frozen delegation versions, work of record via tasks on the firm's ledger, bounds and audit via run limits and matter-level cost attribution, and sovereignty via managed-cloud, private-cloud, or self-hosted deployment. Every one is demonstrable live — which is the test this standard asks you to apply to everyone.