AI Skills for Law Firms: Repeatable Document Review
Vendor marketing keeps pointing at the same place: document review is the center of gravity for legal AI spend. IBM, for instance, published a customer story about Fogel Law Group describing document review compressed from hours to minutes with watsonx — a vendor-published case study, so treat the specifics as marketing rather than independently verified benchmarks.
What none of that material answers is the question a five-lawyer firm actually has: what do we build ourselves, and what do we rent? The most underrated answer, in our opinion, is the least glamorous one — write the instructions down properly. A platform subscription is rented and can be repriced or discontinued; a written review procedure is an asset the firm keeps.
What a “skill” actually is
Strip the marketing away and a skill is a folder of written instructions an AI assistant loads when it’s relevant to the task at hand. Anthropic’s Agent Skills format for Claude, for example, is essentially a markdown instruction file plus optional supporting resources and scripts. The same file is portable in a very literal sense: the field list, output template and escalation rules paste directly into a custom GPT’s instructions box, into a project- or workspace-level custom-instructions field, or into the system prompt of whatever assistant your firm licenses. What changes is where you paste it and whether the platform supports attached reference files — not the content. Mechanics differ by vendor and change fast, so check current product documentation as of 2026. The concept holds regardless: you are turning tribal knowledge into a file.
That’s it. No model training, no fine-tuning, no data science. The skill for “summarize a deposition transcript” is the same thing your best paralegal would write if you asked them to document exactly how they do it: what to extract, in what order, in what format, what to flag, what never to guess.
The model is rented. The instructions are yours. That’s the part worth investing in.
Why this matters more than prompt-of-the-day tips: ad-hoc prompting produces ad-hoc output. Two associates prompting the same assistant on the same medical records get two differently structured summaries, and neither is auditable. A skill makes the output shaped — same headings, same fields, same citation format — which is what makes it reviewable, comparable across matters, and safe to drop into a template.
Where skills fit in a small firm’s document work
Four concrete jobs, in rough order of payoff:
Discovery triage. Not full-scale eDiscovery — if you’re processing hundreds of thousands of documents you need a real review platform with defensible sampling and production logs. Skill-based triage is for a mid-size production of a few hundred files, where someone has to answer “which of these touch the indemnity clause, and which are potentially privileged?” The skill defines the tag set, the responsiveness criteria from the actual RFPs, and a hard rule that anything ambiguous goes to a human queue rather than getting a confident label.
Deposition and transcript summaries. Highly structured, highly repeatable, and the format matters more than the prose. A good skill specifies: page:line cites for every assertion, a chronology, admissions against interest called out separately, and topics where the witness said “I don’t recall.”
Lease and contract abstraction. Take a PDF, output a fixed field list every time — say, 20–30 fields for a commercial lease, covering commencement date, renewal windows, assignment restrictions and notice provisions — with a page citation for each and an explicit NOT FOUND where the field is absent. That last rule is what stops the model quietly inventing a plausible renewal term.
Medical records chronology for PI and disability work: date, provider, complaint, finding, treatment, source page.
Skills, MCP and agents are three different things
This trips up most firm-level buying conversations, so be precise:
An agent is the third piece: the assistant running a multi-step task with those tools — pull the new production from the document management system, run the triage skill over each file, write tags back, open a review task for the flagged items. Skills without access are a smart intern with no keys. Access without skills is a keyholder who improvises. You want both, and you want the skill written before you wire up write permissions to anything.
Building your first review skill
-
Pick one document type you handle constantly
Not “contracts.” Commercial leases in your retail landlord practice, or IME reports in your PI practice. Narrow beats broad — a skill for one document type can encode real rules; a skill for “legal documents” degrades into generic summarization. -
Define the output before the instructions
Write the exact template you want back: field names, order, citation format. If the output has to land in a Word template or a matter field, match those names precisely. This is the same discipline that makes document automation work. -
Write the rules a new hire would need
Include the ones that feel too obvious: never infer a date not in the document; quote the operative clause verbatim; if two provisions conflict, report both rather than resolving them. -
Add two gold-standard examples
One clean document with the ideal completed output, one messy one — scanned, poorly OCR’d, missing exhibits — with the correct handling. Examples do more work than adjectives. -
Test on ten files you already know cold
Score field by field against the human-prepared version and record four counts: fields correct, fields wrong, values that appear nowhere in the source document (hallucinations), and fields the model filled in that should have returnedNOT FOUND. Divide corrected fields by total fields to get a correction rate — that’s the number the ROI section below asks you to supply. Where it diverged, patch the instructions rather than blaming the model, and repeat until the remaining divergences are genuine judgment calls. -
Version it and name an owner
Date it, keep it in the same repository or shared drive as your templates, and require sign-off to change it. An unowned skill drifts. -
Re-run the gold set after every model or platform change
Assistants get upgraded underneath you, and output format and edge-case handling can shift without any change to your file. Keep the same ten files and their known-good outputs as a standing regression test, re-run them when the vendor ships a new model version or you switch platforms, and retire or rewrite the skill if the correction rate moves materially.
Where this breaks, and what stays human
Be blunt about the failure modes. OCR quality is the hidden ceiling — a skill can’t read what the scan lost. Long documents get chunked, and cross-references between chunks are where accuracy quietly drops; test with your longest real file, not a sample. Privilege calls, responsiveness judgments at the edges, and anything going into a production log need a lawyer’s eyes. And extraction confidence is not correctness: a model will fill a field with something plausible unless you explicitly authorize it to say nothing.
Modeling the payoff without making numbers up
We have no benchmark data for your practice, and you should distrust anyone who quotes one. Model it yourself:
The honest version of this math has three assumptions you have to supply: how long the task takes now (time it), what share of the output still needs human correction (the correction rate from your ten-file pilot), and whether recovered hours actually get reallocated to revenue work or just evaporate. Our law firm automation ROI framework lays out the arithmetic. If hours don’t get reallocated, the savings are real for morale and unreal for the P&L.
When to buy the platform instead
Skills on a general-purpose assistant are cheap and fast, but they are the wrong tool if you need defensible eDiscovery workflows, privilege log generation at scale, or citation-checked legal research against a licensed database. Purpose-built legal AI platforms exist because those problems require infrastructure, not instructions.
Our rough heuristic — opinion, not a benchmark: if the job is bounded document work you already do a specific way, write a skill. If it’s research, citation validation, or high-volume review with production obligations, evaluate a platform. If it’s your systems and your workflow, and no vendor covers it, that’s the custom build conversation covered in custom vs off-the-shelf legal software. And a useful reality check on what’s earning its keep versus what’s demo-ware sits in our overview of what actually works in legal AI.
Where is your firm losing billable hours?
Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.
Get a free automation audit