Every legal AI demo looks the same. A question goes in, a confident answer comes out, and a salesperson explains that this will save your associates hours a week. The demo is real. It's also the wrong test. A demo shows you what the product does on its best day, with a friendly question, in front of an audience. It tells you almost nothing about what happens on a bad day — when the model is wrong, when a citation doesn't exist, when a draft goes out that shouldn't have, or when a client asks where their file has been.
Law firms already know how to evaluate the things that matter here. You would never hire an associate without asking who supervises them, what they're allowed to sign, and what happens when they make a mistake. Legal AI deserves the same interview. What follows is that interview: eight questions you can put to any vendor — including us — in writing, before a contract is signed. For each one, we've included what a good answer sounds like, and how EdgeLex answers it, honestly.
1. Where does my data live — and what leaves?
This is the foundation, and it has two halves that vendors like to blur together. The first is storage: where do the documents, emails, transcripts, and matter records physically sit? The second is movement: what leaves that boundary when the AI runs — and goes where, under whose terms? A vendor can truthfully say "your data is encrypted at rest in our cloud" while every privileged document you touch is streamed to a third-party model provider at inference time.
- Where is the system of record — your infrastructure, the vendor's, or a third party's?
- What data leaves that boundary when the AI answers a question, and to whom does it go?
- Can we deploy the whole platform inside our own infrastructure if we choose to?
- Is anything about our matters used to train models by default?
How EdgeLex answers: the firm chooses between two equal, first-class deployments — EdgeLex-managed cloud, where we host and operate the stack, or self-hosted on the firm's own infrastructure or private cloud. In both models the firm is the authority layer: identity, permissions, data, and audit belong to the firm, and per-firm isolation is enforced at the data layer. Nothing about your matters trains a model by default; your data trains a model only when your firm deliberately curates a dataset and runs its own fine-tune — and that model belongs to you.
2. Whose models are these?
Most legal AI products pick the model for you. "Model choice," where it's offered, usually means choosing which cloud provider your data is sent to. That's a real choice, but a narrow one. The fuller question is whether you can run models on your own hardware — where token cost drops to zero and data egress drops to nothing — and whether the platform can ever run a model that is actually yours.
- Can we bring our own API keys and pay model providers directly, on our own terms?
- Can we run open-source models on our own hardware, with no data leaving the deployment?
- Can we ever own a model — one tuned on our work, held in our deployment?
- Who decides which models are allowed to run, and can we turn one off?
How EdgeLex answers: the firm chooses across the full spectrum — frontier models via bring-your-own-key, local open-source models on firm hardware, or the firm's own fine-tuned model built in the Lex Training Center. A three-layer policy (platform, firm, user) governs which models may run at all, default-deny, and the effective allow/deny policy is enforced in the runtime path — a browser client can't override it. A routing layer also refuses to let underpowered models perform destructive actions, whichever model class the firm prefers.
3. How are answers grounded — and can you verify a citation's basis?
Generative models produce fluent text whether or not the underlying claim is true. Every vendor knows this, so every vendor says some version of "our answers are grounded in your documents." The question is what that actually means mechanically. Is grounding a retrieval step that feeds the model context and hopes for the best? Or is there a check, after the model writes and before you read, that each claim traces to something real?
- For any given sentence in an answer, can I see exactly what it rests on — a document, an email, a rule, a ledger entry?
- Are citations verified against a real corpus before I see them, or generated and hoped-for?
- Is verification deterministic where it can be, with AI judgment layered on top — or AI all the way down?
How EdgeLex answers: every claim in a Lex answer carries typed evidence — a document span, an email span, a rules-engine result, a billing ledger entry, a court order's own words, an approval artifact — and that basis is verified before the answer is shown. For caselaw, EdgeCite validates citations deterministically against a citation authority ledger built from the Free Law Project's CourtListener corpus, held inside the firm's deployment, before any AI judgment runs; unknowns escalate to live verification and judgment calls route to a human review queue. The order matters: deterministic first, AI second, human on the close calls.
4. What happens when the AI can't prove something?
This may be the most revealing question on the list, because the honest answers diverge sharply. A system that cannot ground a claim has exactly two options: say so, or say something anyway. Most systems are built — by their training and their incentives — to say something anyway, fluently. Ask the vendor to show you what a failure looks like on screen. If they can't show you one, the system doesn't have a visible failure mode, which means its failures look exactly like its successes.
Ask every vendor to show you what a failure looks like on screen. If failure looks exactly like success, you have no way to tell them apart — and neither does your associate at midnight.
How EdgeLex answers: an ungrounded legal claim is blocked, and shown as blocked — not padded over with plausible prose. The same fail-closed posture runs through the platform: Cascade refuses to invent a deadline no reviewed rule supports; signer-eligibility checks refuse to offer signing when eligibility can't be established; clause benchmarks are blocked while observations are unreviewed or source spans don't verify. We would rather show you a block than show you a guess.
5. Can the AI act — and what stops it?
The industry is moving from AI that answers to AI that acts: sending email, filing documents, creating calendar entries, completing tasks. That's genuinely useful — and it changes the risk profile completely. A wrong answer wastes your time; a wrong action goes out under your firm's name. So the moment a vendor says "agent," the interview should turn to restraint.
- Which actions require explicit human approval, and is that enforced in the architecture or requested in a prompt?
- Can the AI reach tools beyond what its current task requires, or is the tool surface bound per task?
- Are proposed actions and completed actions distinguishable in the record?
- Is there a pause, a retirement path, and a kill switch — and who holds them?
How EdgeLex answers: reads run freely; writes are governed. Anything that changes the world — send, save, update, delete — requires an approval artifact, and proposed versus completed actions are never conflated. Each request is classified into a governed task type before the model runs, producing a contract for the turn: what evidence is required, which tools are bound, whether the model may answer at all. Standing delegations run with a frozen, default-deny tool allowlist — a summarization delegation cannot send email no matter what the model decides mid-run — under iteration limits and budget ceilings, with pause, retire, and a kill switch in the Agent Control Center.
6. Where does the work land — and who pays for it?
Two practical questions, often skipped. First: when the AI produces something — a memo, a summary, a draft — where does it go? In many products the answer is a side panel or a chat transcript: unassigned, unfiled, invisible to the rest of the team, and gone when the tab closes. In a law firm, work that isn't on the record might as well not exist. Second: what does the AI cost, and can that cost be attributed the way everything else in a firm is attributed — to the matter?
- Does AI output land in our systems of record — the DMS, the task list, the matter — or in a side panel?
- Is AI work assigned, tracked, and reviewable like a person's work?
- Can AI spend be broken down by user, by model, and by client matter?
How EdgeLex answers: research memos save to the matter in the DMS; proposed tasks land in Triage; results deliver to EdgeMessage — every write approval-gated. When a delegation runs, the run is created as a real task assigned to Lex on the matter, in the same task list as the rest of the team, with status, history, and review. On cost: every request is metered at the runtime with registry rate cards — spend by provider, by model, by user, and by client matter, with projected monthly spend — so AI cost maps to the file and can be billed, absorbed, or examined deliberately. Local models run at zero marginal token cost.
7. What's on the audit trail — and how to run this checklist
When a client, a court, or a malpractice carrier asks how AI was used on a matter, "we have a policy" is not an answer. The answer is a record: which model ran, what it was allowed to see and do, what evidence its claims rested on, who approved which actions, and when. And the record itself has to be handled carefully — an audit log that contains the privileged content it's auditing is a new problem, not a solution.
How EdgeLex answers: every governance decision — the classification, the model policy check, the evidence verification, the approval, the block — is logged in one audit taxonomy, written whether the turn succeeds or fails. The logs themselves redact prompts, documents, and privileged content. Delegation definitions are immutable and versioned, so every run records exactly which version of an instruction it executed. These controls are designed to support the obligations firms actually carry — data residency, access control, audit logging, retention — which you should verify against your own compliance requirements, as you would with any vendor.
Here's the practical suggestion: send these eight questions to every vendor on your shortlist, in writing, and ask for answers in writing. A vendor whose architecture genuinely does these things will enjoy answering — the answers are specific, checkable, and demonstrable on demand. A vendor whose product can't will answer with a demo. Both responses tell you what you need to know. We wrote this guide because we're happy to be held to it.
Put the checklist to us
Book a demo and walk through a governed AI turn end to end — classification, evidence, approval, and the audit record it leaves behind. Bring the hard questions; they're the point.
See how AI Governance works →