AI Dictation for Lawyers: Dragon vs Word vs Custom Agent
The three things people mean by “AI dictation”
Ask five lawyers about AI dictation and you’ll get three different products.
1. Dedicated legal speech recognition. Dragon Legal (from Nuance, which Microsoft acquired in 2022) is the long-standing example, usually paired with a digital dictation workflow like BigHand or Philips SpeechLive in firms that still route audio to a typist. These are marketed with legal-specific vocabularies and custom word lists. Product names, packaging, and licensing here have shifted repeatedly since the Microsoft acquisition — check current Microsoft and Nuance product documentation for pricing and availability before you budget, as of 2026.
2. Built-in dictation. Microsoft 365 Word has a dictate button. So do macOS, iOS, and Android. Quality on modern speech models is good enough that the marginal case for a separate paid dictation license is narrower than it was five years ago — my hunch, based on firms I’ve talked to rather than any survey, is that a fair number of seats go unused.
3. An agentic dictation pipeline. You speak. The audio is transcribed. An AI assistant with a defined skill reads the transcript, classifies what you just said, and produces structured outputs: a draft letter in your firm’s format, a narrative time entry, a task for the paralegal, a deadline to calendar. Then, through a connection to your practice management or document system, it puts those artifacts where they belong — as drafts, for your review.
Only the third one is AI automation in the sense this publication cares about. The first two are excellent typing replacements.
What each approach is genuinely good at
Strengths: Often already licensed as part of a suite you pay for. Real-time, in-app, no round trip. Handles proper nouns and legal terms well with a trained vocabulary. Broadly faithful to what you said — it transcribes rather than composes.
Limits: Output is a blob of text. Someone still formats it, files it, bills it, and chases the follow-ups. Accuracy degrades with background noise, accents, and cross-talk. And “deterministic” is an overstatement: modern neural transcription models produce substitutions that read fluently but are wrong, and have been documented inventing filler phrases during silence or noise. Read the transcript before you trust it.
Strengths: One dictation produces several work products. Enforces a consistent structure every time. Captures billable narrative at the moment of the work rather than at month-end reconstruction. Can trigger downstream tasks.
Limits: Generative steps can smooth over ambiguity or invent connective tissue you never said. Speaker cross-talk in hearings and depositions confuses both the transcription and the downstream summary — the agent will happily attribute opposing counsel’s position to your client. Matter misidentification is the other killer: two clients named Sanders, or three matters for the same client, and the draft lands in the wrong folder. Requires review discipline, and requires you to have decided what “good” looks like for each output.
A worked example: twenty minutes of talking, four work products
You walk out of a status conference. In the car, you open your phone and talk for four minutes: what the judge said, the new discovery cutoff, that opposing counsel will produce the supplemental medicals by the 14th, that the client needs to be told, and that you want the paralegal to pull the updated records request.
With dictation software, you now have four minutes of transcript in a note. You’ll deal with it tonight, or Thursday, or never.
With an agentic pipeline, the transcript hits a skill — a reusable, packaged set of instructions that teaches the assistant to do this exact job the same way every time. That skill is written by you and encodes your standards: how your firm’s client update letters open, what a compliant time-entry narrative looks like, that deadlines are always proposed and never auto-calendared. Output:
- A draft client update letter in your matter folder, marked DRAFT.
- A proposed time entry attached to the matter, unposted.
- A task assigned to the paralegal with the records-request instruction.
- A flagged deadline item for the calendaring rule to confirm — not written to the calendar.
You review all four in the time it takes to drink a coffee. The same skill-based approach underpins repeatable document review; dictation is just a different input.
A transcript is raw material. The value was never in the words appearing on screen — it was in someone turning them into billable, filed, followed-up work.
Where MCP fits, and why it’s the unglamorous half
The filing step is the hard part. Getting a draft into Clio, NetDocuments, or iManage under the right matter, with the right permissions, is what separates a demo from a workflow. The Model Context Protocol (MCP) is an open standard for giving an AI assistant governed access to your systems and tools — scoped, logged, and revocable, rather than pasting between windows. Practically, that means an assistant like Claude can look up the matter, confirm the client name, and write a draft document to it, without you handing over blanket credentials. We walk through the mechanics in connecting Claude to Clio with MCP.
My opinion, stated as opinion: scope the agent’s write permissions to drafts and proposals only for at least the first quarter. Read broadly, write narrowly. And require the agent to echo back the matter number it selected before it writes anything.
Which AI firms are actually running for this today
Honestly? A mix, and mostly unglamorous. Microsoft 365 dictation and Copilot because it’s bundled. Practice-management AI features (Clio Duo, Smokeball) for summarization inside the system of record. General assistants — Claude, ChatGPT, Gemini — for drafting, often on personal accounts the firm hasn’t sanctioned. Specialist platforms like CoCounsel or Harvey where research and large-document work justify the price. Purpose-built transcription (Otter, Fireflies, or Whisper-based tooling) for meetings and calls, which overlaps with but isn’t the same problem as dictation — see AI meeting notes for lawyers.
The custom layer is rarely a replacement for any of that. It’s usually thin connective glue between tools the firm already pays for.
Applying an 80/20 lens before you build anything
The 80/20 rule — the Pareto idea that a small share of inputs drives most of the output — is genuinely useful here, but only if you apply it to dictation types, not clients. Most lawyers dictate a handful of recurring things: post-hearing notes, client update letters, file memos, intake summaries, settlement authority notes. Pick the one you do most often and build a single skill for it. Skip the long tail.
A firm that builds one excellent post-hearing skill beats a firm that builds nine mediocre ones.
Modeling the payback with your own numbers
The relevant input isn’t a salary headline — it’s your realized hourly value and the hours you’d actually reallocate. Attorney compensation varies enormously by practice area, geography, and ownership structure, and no national average tells you anything useful about your own firm.
Build the model yourself. These are formulas to fill in, not benchmarks:
- Recovered admin time: hours saved per week × 48 weeks × your realized hourly rate.
- Captured billable narrative: entries you currently miss or under-describe per week × 48 × rate.
- Reallocation check: of those recovered hours, what share realistically becomes billable or business-development work? Apply that fraction — don’t assume 100%.
- Cost side: build or configuration cost + annual licenses + review time per dictation × volume.
Time your current process for one week before you touch anything. If you want a fuller structure, our automation ROI calculator walks through the cost side people usually forget.
Confidentiality, and the part that stays human
Two things to settle before rollout. First, where the audio goes: dictating privileged client information into a consumer app with unclear data-retention terms is a different risk profile than an enterprise agreement with no-training and retention controls. The ABA’s Formal Opinion 512 (2024) addresses lawyers’ duties when using generative AI tools, including confidentiality, competence, and client communication; your state bar may have its own opinion that goes further. Read both, and confirm specifics with a qualified professional in your jurisdiction before you deploy firm-wide.
Second, review. A generative step can quietly produce a fluent sentence you never said — a date that isn’t right, an inference presented as fact. So can the transcription step. That risk is highest exactly where the stakes are: deadlines, settlement authority, admissions. Keep calendaring, client-facing sends, and anything filed with a court behind a human check. Our agent oversight models cover how to structure that without making it theater.
-
Pick one dictation type
Post-hearing notes or client update letters. Not both. -
Write the skill by hand first
Draft the instructions as if training a new paralegal: structure, tone, what to never assume, what to flag instead of guess. -
Run it read-only for two weeks
Transcript in, drafts out, nothing written to your systems. Compare against what you’d have produced manually. -
Connect writes, narrowly
Add the MCP connection with permission to create drafts and proposed entries only. Log everything. -
Decide build vs. buy with real data
If the vendor feature in your practice management system now does most of this, use it. Custom is for the remainder that’s specific to your firm.
Where is your firm losing billable hours?
Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.
Get a free automation audit