Law Firm AI Automation: Vendor Agents vs Custom Build
What law firms are actually doing with AI right now
Search for law firm AI experiences online and you get two loud camps: people who say it drafted a flawless motion, and people who say it invented a citation and nearly got them sanctioned. Both are reporting real experiences with very different workloads.
The honest picture as we read it in 2026: the durable wins are clustering in repeatable, verifiable, non-adversarial work — intake triage, document review against a fixed checklist, discovery sorting, deposition summarization, billing hygiene, matter status reporting. The shaky ground remains anything where an unverified output goes straight to a court or a client.
What changed recently is that vendors stopped shipping chatbots and started shipping agents — software that takes multi-step actions inside your systems. Practice-management and billing platforms (Aderant, Clio, Smokeball, MyCase), horizontal suites (Microsoft Copilot, Google’s Gemini offerings), and general assistants like Claude are all moving this way. Announcement dates and feature availability change several times a year, so check each vendor’s own newsroom and product documentation for what actually ships today rather than trusting a date in an article — including this one.
The two build-vs-buy poles
Pick one workflow and hold the options against it. Ours below is the monthly prebill review: the partner slog of reading time entries, fixing vague narratives, flagging block billing, spotting write-offs, and catching work that never got recorded.
Strengths: already inside the system of record, so the data model, permissions, and audit trail are the vendor’s problem. No integration project. Priced as an add-on to software you already buy. Fastest path from decision to first output.
Limits: it does what the vendor built, in the vendor’s order. If your prebill rules are idiosyncratic — client billing guidelines, LEDES codes, a specific narrative style per insurer — you get generic cleanup and still re-read everything. Roadmap dependency: your request competes with every other firm’s.
Strengths: encodes your billing guidelines, your write-off thresholds, your tone. Can reach across systems your practice-management platform doesn’t own — email, shared drives, a spreadsheet the billing manager maintains. You control what the agent may read and do.
Limits: someone has to build and keep it running. It depends on your platform having a usable API — and when the vendor changes that API or its underlying schema, your agent breaks quietly. Model version upgrades can shift behavior on prompts that worked fine last quarter, so you re-test. And somebody has to own the skill document when the person who wrote it leaves the firm. It’s a budgeted project with an ongoing owner, not a checkbox. If your process is genuinely standard, this is over-engineering.
The middle path: enterprise AI suites over documents and email
Enterprise suites — Microsoft Copilot, Gemini in Google Workspace, or a Claude deployment across your tenant — deserve the same treatment as the two poles, not a footnote.
Strengths: they sit where a lot of firm knowledge actually lives — matter correspondence, closing sets, templates, chat threads. Licensing usually piggybacks on an existing Microsoft 365 or Google Workspace agreement, and the data stays inside a tenant your IT provider already administers. For “what did we agree with this client about fee caps,” they’re strong out of the box.
Limits: they’re weak on structured, transactional systems. Your time entries, WIP, and trust ledger live in a practice-management database, and a suite that only sees files and mail can’t reason about them until you connect it. That connection is increasingly MCP — an open protocol for giving an AI assistant governed access to specific tools and data. Once you’re doing that connection work, you’re partway into a build; see our walkthrough on connecting Claude to Clio with MCP.
The question isn’t “which AI is smartest.” It’s “which one can see my time entries, and who approves what it changes.”
A worked example: the prebill agent
Here’s the shape of the job, whichever path you pick.
-
Define the rules in writing first
Before any tool: what makes a narrative acceptable at your firm? Minimum length, no block billing over X hours, no internal jargon, client-specific task codes. If you can’t write it down, no agent can enforce it. This document becomes your skill — a packaged instruction set the assistant follows the same way every cycle. -
Give the agent read access, not write access
Let it produce a review memo: flagged entries, suggested rewrites, missing-time candidates inferred from calendar and email activity. Nobody’s ledger changes yet. -
Human approves in batch
The billing manager works the flagged list and applies the edits. This is the step to never automate away — narrative changes affect what a client is charged. -
Decide at day 30, on one full cycle
Run exactly one prebill cycle this way. At day 30, compare hours spent on the cycle before and after, the share of suggestions accepted without edits, and realization on the affected matters. Continue, adjust the skill, or stop — but decide with your own numbers, not a vendor case study. -
Only then let it write
Once suggestions are accepted at a high rate without edits, allow direct writes for narrow, low-risk categories (typo fixes, task-code normalization) and keep human approval for anything touching hours or amounts.
The same skeleton — read, suggest, human approve, narrow write access — applies to matter status reports, deadline audits, and vendor invoice coding. Our companion piece on stopping billable-hour leakage covers the capture side of this workflow.
What this actually costs
There’s no honest single answer, because the three paths price differently: platform add-ons are typically per-user per-month, enterprise suites are per-seat with usage tiers, and a custom build is a one-time project cost plus API usage plus maintenance. Check current pricing on the vendor’s own page — legal AI pricing has moved repeatedly and anything quoted in an article ages badly.
What you can do is model payback yourself. Fill in your own figures:
Take the hours your team currently spends on the cycle, multiply by a blended rate, subtract the hours it still takes after the agent, and compare that against the total annual cost of the tool including the internal time to run and maintain it. As an editorial rule of thumb rather than a measured benchmark, we’d judge a custom build on a 12-month payback window — if it isn’t obvious over a year with conservative assumptions, it isn’t there.
Rates, the billable hour, and the business-model question
People ask whether AI justifies a higher hourly rate. Rates vary enormously by market, practice area, and seniority; published rate surveys for your own jurisdiction are the place to check that, not an article.
The operations point is narrower and more useful: AI doesn’t raise your rate, it changes what’s inside the hour. If an agent removes six hours of prebill review, that gain only converts to revenue if those hours get redeployed to billable or business-development work. Otherwise you’ve bought comfort, not income. Ethics matter here too — ABA Formal Opinion 512 (2024) addresses generative AI, including fees and expenses: lawyers can’t bill for time they didn’t spend, and cost pass-throughs have to comply with the fee rules. Your state bar may go further; confirm with your jurisdiction’s ethics guidance, and with a qualified legal professional, before changing how you bill for AI-assisted work.
Where these agents break
The disadvantages aren’t mysterious. Agents confidently produce plausible wrong output; they don’t know what they weren’t shown; they drift when the underlying data is messy; and they create a supervision burden that partners routinely underestimate. An agent with write access and a bad instruction is faster at being wrong than any human.
There’s also a boring failure mode: buying an agent for a workflow a rules-based automation already solves. If your prebill problem is “nobody enters time on Fridays,” the fix is a reminder and a policy, not a model.
How to choose
Start with the built-in agent if your process is close to standard and you want an answer this quarter. Start with an enterprise suite if the knowledge you need is trapped in documents and email rather than a database. Build custom when your rules are genuinely yours, the workflow crosses systems, and you’ve already proven value with something simpler — the sequencing argument we make in why automation stalls after the first win.
And run the same 30-day read-only pilot regardless: one workflow, one owner, one full cycle, your own before-and-after numbers, a decision at day 30. That test is worth more than any comparison table, including this one.
Where is your firm losing billable hours?
Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.
Get a free automation audit