Two years ago legal AI vendors sold chatbots. Today every one of them sells "agents." Vendors like Harvey and Clio now publicly describe their AI as doing legal work end-to-end, and the rest of the market — cloud legal AI assistants, practice-management platforms, contract tools — has adopted the same language. The word has become a tier of marketing rather than a description of architecture, which is a problem, because the difference between a chatbot and an agent is not a feature difference. It's a risk difference.
Here's a way to cut through it. Law firms already run the world's oldest delegation system: the associate. You know exactly what it takes to hand work to a junior lawyer — clear instructions, supervision proportional to consequence, work that lands on the record, and a review before anything goes out the door. Every question you should ask about an AI agent is a question you already know how to ask about an associate. So let's use that lens, and walk up the ladder one level at a time.
Level one: the chatbot — answers while you watch
A chatbot is a conversation. You ask, it answers, and nothing in the world changes except your understanding. In associate terms, this is the junior you stop in the hallway: "What's the standard for a 12(b)(6) motion in this district?" Useful, fast, and low-stakes in one specific sense — the associate isn't doing anything with the answer. You are. Whatever happens next passes through your judgment first.
That doesn't make chatbots safe by default. A hallway answer you rely on is only as good as its basis, and a fluent model will answer confidently whether or not a basis exists. So the questions that matter at level one are all about grounding: What is this answer actually resting on? Can I see the source — the document, the rule, the case — or am I taking the model's word for it? And the sharpest one: what does the system do when it can't support a claim? In EdgeLex, the answer to that last question is structural — every claim in a Lex answer carries typed evidence (a document span, an email span, a verified citation, a rule authority), and a claim that can't be grounded is blocked and shown as blocked, not padded over with plausible prose.
If a vendor's product lives entirely at this level, that's fine — a well-grounded chatbot is genuinely valuable. The trouble starts when a level-one product wears level-three language.
Level two: the assistant with tools — acts while you watch
At level two, the AI can do things: draft an email, create a task, propose a calendar entry, save a document. This is the associate at your desk, working while you supervise — "pull the agreement, draft the notice letter, and show me before it goes anywhere." The critical property of level two is that you're present. You see each step, and the work stops at your desk for review.
This is where most "agent" products actually live: a chat session with tools attached. It's a real capability jump, and it introduces the first real accountability questions, because now a wrong output isn't just a wrong answer — it's a draft that could become an action. The questions to ask: Which actions can it take at all? Is the boundary between proposed and done unmistakable — can a draft ever leave without a human decision? And is the tool surface bounded per task, or does every conversation carry every capability the platform has?
In EdgeLex, level two is governed the way you'd supervise the associate at your desk. Reads run freely; writes are governed. Anything high-impact — sending an email, saving a document, completing a court-mandated task — surfaces as an approval card and waits. Before the model even runs, the request is classified into a governed task type that binds only the tools the task needs; the model can't call what it was never given. Lex proposes, you approve — and proposed versus completed actions are never conflated in the record.
Level three: the standing agent — acts while you're gone
Level three is the real thing, and it's categorically different: the AI works when you're not there. You give an instruction once — "whenever opposing counsel files anything, verify every authority they cite and brief me" — and the work happens on the firm's events, while you're in court, at dinner, asleep. This is no longer the associate at your desk. This is the associate you've sent off with a standing assignment, and everything a firm knows about that kind of delegation suddenly applies.
What makes level three different in kind, not degree, is the trigger. A chatbot and a tool-using assistant both start when you do — they're bounded by your attention. A standing agent starts when the world does: a document lands on the matter, an email arrives, a deadline approaches, a schedule ticks. That's precisely what makes it valuable — the follow-up work of a practice mostly arrives while you're doing something else — and precisely what removes the safety net the first two levels had built in, which was you, watching.
Think about what you actually require before an associate can carry standing work. The instruction has to be written down, and if it changes, everyone needs to know which version they're operating under. The work has to land somewhere visible — assigned, dated, reviewable — not in a drawer. Anything consequential still comes back to a partner before it goes out. The work product has to show its support. And there's a limit: an associate who burns forty hours going in circles gets pulled in, not left running. None of that is bureaucracy. It's what makes delegation compatible with responsibility — and it's exactly the checklist a standing AI agent has to meet.
The associate test: if you wouldn't accept it from a junior lawyer — unwritten instructions, untracked work, unsupervised sending, unsupported claims — don't accept it from an agent.
What accountability actually requires
Here's the same checklist stated as architecture. When a vendor says "agent," these five properties are what separate a standing agent a firm can answer for from a chat session with tools and good marketing:
- Versioned instructions. The delegation's definition is written, frozen when activated, and immutable — edits create a new version, and every run records exactly which version it executed. You can always answer "what was it told to do, at the time it did this?"
- Work on the task ledger. Every run is created as a real task assigned to the agent on the matter — same task list as the humans, with status, history, and review. AI work is trackable and reportable exactly like an associate's work, not output floating in a side panel.
- Approval gates on consequential steps. Sending email, filing work product, touching deadlines — the run pauses and asks. Finished work arrives as a reviewed deliverable, never as a surprise already sent.
- Evidence on the output. The work product carries its basis — cited spans, verified authorities, rule references — and claims that can't be grounded are blocked, not smoothed over.
- Bounded runs, with a leash. A frozen, default-deny tool allowlist per delegation, iteration limits, budget ceilings, progress checks — and pause, retire, and a kill switch held by the firm.
This is how Lex Delegations are built in EdgeLex. You create one in plain language — "whenever a document is linked to this matter, summarize it and message me the key points" — and Lex compiles the sentence into a formal delegation with defined triggers, steps, and limits that you confirm before it goes live. From then on it wakes on the firm's real events, does the work inside its frozen capability surface, pauses at every consequential step, and lands its runs on the matter's task list. The Agent Control Center shows every delegation, version, run, and delivery in one place.
One more property matters at level three, because standing work is easy to lose track of: the work has to arrive somewhere you already look. In EdgeLex, a delegation's finished work arrives as a reviewed deliverable in EdgeMessage — never as a surprise already sent — and the run itself sits on the matter's task list with the rest of the team's work. When a partner asks what the AI did on this matter last month, that's a query against the ledger, not a shrug.
Notice what the frozen allowlist buys you at level three specifically. At level two, you're watching — if the assistant reaches for the wrong tool, you're there. At level three, nobody's watching, which is why the boundary can't be behavioral ("the model is instructed not to") — it has to be structural. A summarization delegation in EdgeLex cannot send email, no matter what the model decides mid-run, because the tool simply isn't bound to it. That's the difference between asking for restraint and building it.
How to place any product on the ladder
Next time a vendor says "agent," place the product with three questions. Does it act, or only answer? If it only answers, it's a chatbot — evaluate its grounding, and don't pay agent prices for it. Does it act while you watch, with your approval? Then it's an assistant with tools — evaluate the approval boundary and the tool binding. Does it act while you're gone? Then, and only then, is it a standing agent — and the full associate test applies: versioned instructions, ledgered work, approval gates, evidence, bounded runs.
There's nothing wrong with any level of this ladder. Firms need all three, and in EdgeLex all three exist under one governance layer — the grounded conversation, the approval-gated assistant, and the standing delegation are the same Lex, held to the same rules. What's wrong is selling level one or two with level-three language, because the gap between what the word promises and what the architecture delivers is precisely where a firm gets hurt. Lawyers, of all professionals, know that a title means nothing without the duties attached. "Agent" is a title. Ask about the duties.
See a delegation run end to end
Set up a standing delegation in one sentence, then watch a run go from a firm event to a reviewed, approval-gated deliverable on the task ledger — with the version, the evidence, and the kill switch all in view.
Explore Lex Delegations →