AI Contract Review: Spellbook vs LegalOn vs Custom Agent

By Jude Lee · · Comparison

Two lawyers reviewing a marked-up commercial lease on a laptop in a small law firm conference room

The three real choices, not five

Strip away the marketing and there are three ways a small or mid-size firm can put AI on a contract review or abstraction job.

One: a purpose-built contract reviewer that lives in Word. Spellbook and LegalOn are the two names that come up most often in this category as of early 2026, alongside larger platforms. They sit in the drafting environment, suggest redlines, flag missing or off-market clauses, and compare a draft against a playbook.

Two: a general assistant plus a written playbook. Claude, ChatGPT or Copilot with a carefully written set of instructions — what we’d call a reusable skill for document review. No contract-specific vendor, no per-seat legal AI price, but you write and maintain the instructions yourself.

Three: a custom agent that takes action. An AI assistant connected to your document and practice management systems through MCP (the Model Context Protocol, an open standard for giving an AI governed access to specific tools and data), running a multi-step job: read the executed PDFs in a matter, extract a defined set of terms, write them into a structured record, flag the exceptions for a lawyer.

Two Word add-ins, two different starting points

Treating Spellbook and LegalOn as interchangeable is the most common mistake in this comparison. The distinction we’d draw, based on how each vendor positions and documents its product as of early 2026 — verify against current product documentation, because this category ships changes monthly:

Orientation. Spellbook leads with drafting and redlining assistance: suggest language, mark up the other side’s paper, answer questions about the draft in front of you. LegalOn leads with review against a defined standard: does this contract deviate from a playbook, and where. Both do some of both. The question is which job the product was shaped around, because that’s where the sharp edges are fewest.

Where the standard comes from. A playbook-first tool is only as good as the playbook. LegalOn’s pitch has centered on attorney-authored reference content plus your own positions; a drafting-first tool leans more on the model and your uploaded precedents. If your firm has no written positions on indemnity caps or limitation of liability, a playbook tool will make you write them — which is useful work, but it is work.

Where both stop. Neither one is a matter-file tool. Their native home is the Word document. Any push of extracted terms into Clio, NetDocuments or a deadline calendar depends entirely on what integrations the vendor happens to offer. Ask that question during the trial, not after.

Where the Word add-ins genuinely win

If your bottleneck is drafting and negotiating — marking up the other side’s MSA, papering an NDA, arguing over an indemnity cap — the add-in category is hard to beat, and building your own version of it is a poor use of money. The work already happens in Word, the reviewer knows what a redline is, and the vendor maintains the clause libraries and playbook logic. You’re buying a finished product for a job thousands of other lawyers do the same way.

Where a custom agent earns its keep

The add-in model has a ceiling: it produces a better document. It does not produce a better matter file.

Think about a firm handling commercial leases for a regional landlord, or franchise agreements, or a portfolio of vendor contracts in a diligence. The valuable output isn’t a redline — it’s a structured abstract. Commencement date, rent escalation schedule, renewal option and notice deadline, assignment restrictions, CAM cap, insurance requirements, termination triggers. Times 40 documents. Times every time the portfolio changes.

That’s abstraction, and it’s a different shape of problem. The work is repetitive, the output schema is fixed, and the value only appears when the extracted data lands in a system someone actually uses — a matter custom field in Clio, a table in your DMS, a calendar entry for the notice deadline.

An AI that hands you a beautiful summary in a chat window has moved the work, not removed it. Someone still has to retype it into the file.

This is what MCP was built for. Rather than copy-pasting between a chat window and your practice management system, you connect the assistant to your systems through a governed server exposing specific, permissioned actions — the pattern we walk through in connecting Claude to Clio via MCP.

Off-the-shelf contract AI (Spellbook, LegalOn, etc.)

Best for: drafting, redlining, playbook comparison, one-off third-party paper.

Time to value: days. Install, trial, roll out.

Maintenance: the vendor’s problem.

Ceiling: output stays in the document. Integration with your PMS is whatever the vendor built.

Cost shape: predictable per-seat subscription that scales with headcount.

Custom abstraction agent (Claude + MCP)

Best for: repeatable extraction across many documents where the output must populate a system.

Time to value: weeks, and only if the schema is genuinely stable.

Maintenance: yours — schema changes, API changes, prompt drift, evaluation.

Ceiling: high. The agent can write to your matter records, create deadlines, and answer portfolio-level questions.

Cost shape: build cost up front, then usage. Scales with volume, not headcount.

How to build the abstraction agent without kidding yourself

  1. Freeze the schema first

    Write the list of fields you want extracted, with a definition for each, before you touch any AI. If your lawyers can’t agree on what “effective date” means across your contract types, the agent will not save you.

  2. Write it as a skill, not a prompt

    Package the instructions — field definitions, extraction rules, how to handle ambiguity, when to escalate — as a reusable skill so every document gets the same treatment. “Return null and flag for review” should be an explicit instruction for anything uncertain.

  3. Run read-only for a month — here's why

    The agent extracts; a human checks and enters. The failure modes are specific and worth naming. Scanned or OCR’d documents produce garbled dates and dollar figures that read as confident, clean values. Handwritten marginal amendments and initialled changes are frequently missed entirely. The operative term often isn’t in the contract at all — it’s in a side letter, an incorporated exhibit, or an amendment filed three years later. And extraction degrades silently when a counterparty changes their template: the same field that pulled correctly for two years starts returning the wrong clause, with no error thrown. A read-only month is how you find these before they’re written into a calendar.

  4. Connect the write path through MCP

    Once accuracy holds, expose narrow, specific write actions — update these fields on this matter, create this deadline — rather than broad database access. Log every action.

  5. Keep the exception queue visible

    Everything flagged uncertain goes to a named human with a deadline. An exception queue nobody owns is where automation projects quietly die.

What firms are actually running right now

If you search for what AI most law firms use, you’ll get confident answers and little evidence. Our honest read: for small and mid-size firms, the most common deployment is the AI already bundled into software the firm bought for another reason — summarization and drafting inside Clio, Microsoft Copilot for firms on Microsoft 365, the AI layers in Lexis and Westlaw — plus a lot of unsanctioned personal ChatGPT and Claude use. The enterprise names (Harvey, CoCounsel) matter more in large-firm and in-house budgets than in a twelve-lawyer shop; if you’re weighing them, we compared them against a custom build in Harvey vs CoCounsel vs a custom AI agent.

What this does to your rate, and what it doesn’t

A fair question hiding under “is $900 an hour a lot for a lawyer”: published rate surveys routinely show senior partner rates at that level in major US markets, well above typical small-firm rates — check a current survey for your own market rather than taking a number from an article. AI does not move your hourly rate. Rate is a function of market, practice area and reputation.

What AI can move is mix and margin. If lease abstraction currently consumes associate hours you can barely bill, an agent that does the first pass converts those hours into billable work at your normal rate. It also gives you something you probably don’t have today: a per-document cost floor you can actually calculate, which is the precondition for pricing abstraction as a fixed fee instead of an apology on the invoice.

Model it as a formula, not a headline:

Hours × rate
Recovered time reallocated to billable work — measure your own baseline first
Worked example, not a benchmark
Build + run cost
Compare against the same job done by a paralegal at loaded cost
Worked example, not a benchmark
Malpractice deductible + write-off
Plug in your own figures for the error-avoidance side of the ledger
Worked example, not a benchmark

Time the job yourself across ten documents before and after. The full structure is in our automation ROI walkthrough.

The decision rule we’d use

This is opinion, stated plainly: if your contract work is varied, negotiated and drafting-heavy, buy the add-in and stop there. If it’s repetitive extraction against a stable schema, at volume, and the output belongs in a system rather than a document, build the agent. If you’re doing fewer than a handful of these documents a month, do it by hand and spend your automation budget on intake or billing instead — those are usually the bigger leaks.

The worst outcome is buying a legal AI seat for everyone, watching three people use it, and concluding AI doesn’t work for your firm. Pick one contract type, one schema, one measurable job.

Where is your firm losing billable hours?

Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.

Get a free automation audit