Agentic AI for Discovery Triage: A Small Firm Playbook

By Jude Lee · · Workflow

Two lawyers reviewing discovery documents and a laptop in a small firm conference room

Discovery is where small and mid-size firms lose the most unbillable, unrecoverable time. A 12,000-page production in a commercial dispute doesn’t need genius — it needs someone to open every file, decide whether it’s relevant, notice the ones with counsel in the CC line, and keep a record of the decisions. That is exactly the shape of work that AI agents handle differently from chatbots: multi-step, tool-using, repetitive, and auditable.

What “agentic” means here, without the hype

A chatbot answers a question about one document you paste in. An agent is given a goal and a set of tools, then takes multiple steps on its own: list files in a folder, open each one, extract text, classify it, write results back somewhere, and report what it did. The interesting shift is the tool access, not the intelligence.

The practical mechanism as of 2026 is the Model Context Protocol (MCP) — an open standard for giving an AI assistant governed access to specific systems and actions. Instead of a paralegal downloading files and pasting them into a chat window, the assistant connects directly to your document management system through a server that exposes exactly the operations you allow: read this matter’s folder, yes; delete anything, no. We walked through the mechanics of that in connecting Claude to Clio with MCP.

Where an agent actually fits a small-firm discovery workflow

  1. Ingest and normalize

    The agent pulls documents from the matter folder in your DMS or practice platform, extracts text (including OCR for scanned PDFs), and records file metadata: custodian, date, type, hash. Deduplication by hash is not AI at all — it’s plain code, and it should stay plain code because it is exact and free.

  2. Classify against your issue list

    You give the agent the actual issue taxonomy for the case — say, five buckets in a construction defect matter: scope changes, delay notices, payment applications, inspection records, and irrelevant. The agent tags each document, cites the passage that drove the tag, and assigns a confidence level. The citation requirement matters more than the tag: it lets a human check the reasoning in seconds.

  3. Flag privilege candidates conservatively

    Rule-based detection catches most of it — attorney domain names, known firm addresses, “privileged and confidential” strings. The agent adds the harder cases: business advice threaded with legal advice, forwarded counsel emails, in-house counsel acting commercially. Every flag is a candidate for human review, never a final call.

  4. Produce a hot-doc shortlist and chronology

    This is where partners feel the value — but the shortlist size is a parameter you set, not a capability the agent promises. You cap it at a number a lawyer can actually read in a sitting (say 40), and the agent ranks documents against your issue list to fill that cap, with a one-line reason each, plus a date-ordered chronology with document citations. Wrong ordering or a missed document is obvious to the lawyer who knows the case; a hallucinated quote is not, so require exact quoted text with a file and page reference.

  5. Write the audit log

    Every run logs which documents were touched, which model version and prompt were used, what was tagged, and who reviewed it. If your methodology is ever questioned, this log is the answer. Build it on day one, not after the first fight.

What stays human, and the rule that makes it non-negotiable

Federal Rule of Civil Procedure 26(g) requires an attorney of record to sign disclosures, discovery requests, responses and objections. Under 26(g)(1)(A), signing a disclosure certifies that it is complete and correct as of the time it is made. Under 26(g)(1)(B), signing a request, response or objection certifies — after a reasonable inquiry — that it is consistent with the rules, not interposed for an improper purpose, and not unreasonably burdensome or expensive given the needs of the case. No agent signs either one. The lawyer does, and “the model tagged it irrelevant” is not a reasonable inquiry on its own. Read the rule text itself before you write policy around it.

The ethics layer is settled enough to plan around. ABA Formal Opinion 512 (issued July 2024) addresses generative AI and lawyers’ duties of competence, confidentiality, communication, supervision, and reasonable fees; several state bars have issued their own guidance. Read your own jurisdiction’s opinion — they differ on client consent for putting confidential material into third-party tools. Also check the standing order of the judge assigned to your case; some require disclosure of AI use in filings. Confirm the specifics with the primary source or your ethics counsel before you set firm policy.

An agent can tell you what a document says. It cannot tell you that you have looked hard enough — and that is the part you sign.

Practical human checkpoints we’d argue for: 100% human review of every privilege flag; human review of every document the agent marks low-confidence; and a sampled quality check of the “irrelevant” pile, because the expensive error in discovery is the false negative nobody sees.

Buy the feature, package a skill, or build the integration

Off-the-shelf AI in your review platform
Relativity, Everlaw, DISCO and similar platforms ship AI classification, and eDiscovery vendors are actively marketing agentic features on top of it. If you already pay for one of these, use what’s in it first. You get vendor support, an established defensibility story, and no engineering burden. Downside: per-GB pricing that stings on small matters, and you work the way the platform works.
Custom agent over your own systems
A skill plus MCP access to your DMS makes sense when your documents never leave Clio/NetDocuments/SharePoint, your matters are too small to justify a hosted review platform, or your triage logic is genuinely firm-specific (a niche practice with a repeatable document set). You own the taxonomy and the audit log. Downside: you own the maintenance, the security review, and the QA.

The middle path most small firms should try first: write the triage instructions as a reusable skill — a packaged, versioned set of instructions that makes the assistant do the job the same way every time — and run it against exported documents before you connect anything to production systems. We covered how to build these in AI skills for repeatable document review, and the broader build/buy tradeoff in custom vs off-the-shelf legal software.

What firms are actually using

Adoption is real and no longer experimental, though quality varies wildly. The ABA’s annual Legal Technology Survey Report is the best neutral tracker of what US firms actually use by firm size; check the current edition rather than trusting vendor-published adoption figures.

What firms use falls into four buckets: general-purpose assistants (Claude, ChatGPT, Microsoft Copilot) for drafting and summarizing; AI features inside the practice platform they already pay for; research-grounded tools like Lexis+ AI and Westlaw’s CoCounsel; and, at larger firms, dedicated legal AI platforms. Harvey has raised large enterprise funding rounds, which tells you where the enterprise money is going but says nothing about what a six-lawyer firm should buy. For most small firms, the honest answer is that the assistant you already have plus a well-written skill outperforms a platform you can’t afford to use fully. Our take on what actually works versus hype goes deeper.

Model the payoff yourself — don’t trust a headline number

Any firm-specific savings figure you see in a vendor deck is theirs, not yours. Build your own with a simple formula:

(documents ÷ human review rate per hour) × blended hourly cost = current first-pass cost. Then estimate the agent-assisted version: (documents ÷ agent throughput) + (flagged documents ÷ human review rate) × blended cost + tooling cost.

Your numbers
Docs per matter × minutes per doc × staff cost
Illustrative model — fill in your own
100%
Share of privilege flags that should get human eyes
Our recommended policy, not a benchmark
0
Discovery certifications an agent can sign under FRCP 26(g)
FRCP 26(g) — the attorney of record signs

Run that math with pessimistic assumptions. If it still clears, pilot it. Our automation ROI walkthrough shows how to handle the part people forget: recovered hours only convert to revenue if they’re reallocated to billable or business-development work, not absorbed by Parkinson’s law.

The rate question hiding underneath all of this

Owners often frame this as a rate problem — whether they can justify a higher hourly number, or how a small practice reaches a target income. Both questions connect directly to discovery triage. Rates at the very top of the market are typical of senior partners at large firms in major metros, and far above what most small-firm rates look like; there is no single national rate, and published rate surveys vary by market and practice area, so compare against local data rather than a headline.

The more useful framing: high earnings in a small firm come from throughput and realization, not from the rate alone. If discovery triage consumes associate hours you write off, or partner hours you bill at a rate clients resist, the fix is structural — push the mechanical passes to software and rules, keep the judgment work at full rate, and stop writing off the difference. That’s the same logic behind stopping billable-hour leakage.

One staffing point to settle before the first run: name the person who owns the issue taxonomy and the person who owns the audit log. They can be the same human — usually your most experienced paralegal — but the roles are different. The taxonomy owner decides what the buckets mean and revises them as the case theory shifts. The audit log owner makes sure every run is recorded, reviewed, and retrievable if the methodology is ever challenged. Unowned, both jobs quietly stop happening around week three.

Where is your firm losing billable hours?

Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.

Get a free automation audit