AI Demand Letters: EvenUp vs Supio vs a Custom Agent

By Jude Lee · · Comparison

Paralegal and attorney reviewing medical records and a drafted demand letter on laptops in a small law firm office

What actually eats the hours in a demand package

Strip a demand package down and it is five distinct jobs, not one:

  1. Collect and organize records — provider requests, follow-ups, dedupe, split a 900-page PDF into per-provider sets.
  2. Build a medical chronology — date, provider, complaint, finding, treatment, page cite.
  3. Tally specials — bills, liens, wage loss, out-of-pockets, reconcile against the ledger.
  4. Write the narrative — liability, mechanism of injury, course of treatment, pain and suffering, the ask.
  5. Assemble and send — exhibits, index, cover, adjuster, calendar the follow-up.

In my assessment, AI is genuinely strong at jobs 2 and 4, decent at 3 with tight guardrails, and mostly irrelevant to 1 and 5 — which are logistics problems that a rules-based workflow or a good paralegal checklist solves better. That split matters, because vendors sell the whole package and firms end up disappointed that the tool didn’t chase Dr. Nguyen’s office for the missing PT notes. Which brings me to the limit I’d put on every option below:

The AI can summarize every page you give it. It cannot tell you which pages you never received — and that gap is what actually sinks demands.

Option one: a purpose-built demand letter platform

EvenUp and Supio are the names most PI firms encounter first; both market AI-assisted medical chronologies and demand drafting for plaintiff-side personal injury work. I have not run a controlled comparison of either product, so treat what follows as a description of how this category positions itself rather than a verified capability claim. Check each vendor’s current documentation directly for scope, turnaround, review model, and pricing — this category changes fast and I’m not going to quote numbers I can’t stand behind.

What these vendors say you’re buying is domain specificity: products built against PI record sets that handle the messy realities — handwritten intake forms, duplicated imaging reports, provider bills embedded mid-chart — better out of the box than a general assistant you point at a folder. Several also describe human review layered into their delivery. If that matters to you, make the vendor put the review model in writing, because it changes what you’re actually purchasing: labor, not just software.

The trade-offs are the usual ones for point solutions. Your records leave your system and enter theirs, so your confidentiality and vendor-diligence analysis has to be serious, not a checkbox. You get the vendor’s chronology format and the vendor’s letter voice, with configuration rather than control. And the tool sits beside your case management system rather than inside it, so someone still moves the output back into Clio, Filevine, or Smokeball.

Option two: a general assistant plus a written skill

The cheapest experiment is to take an assistant you already pay for — Claude, ChatGPT, Copilot — and give it a skill: a reusable, packaged instruction set that teaches it to do one job the same way every time. For demands, that means a document defining your chronology columns, your citation format ([Provider – 2025-03-14 – p.412]), your rules for what counts as a gap in treatment, your letter sections in order, and an explicit instruction to output “NOT IN RECORD” rather than infer.

This is the same discipline described in our piece on building repeatable document review skills. A firm with a good skill file and a mediocre model beats a firm with a great model and no written standard, because the second firm’s output varies by whoever typed the prompt that morning.

This option deserves the same confidentiality scrutiny you’d apply to a vendor. Consumer and enterprise/business tiers of these assistants differ materially on whether inputs may be used to improve models, how long data is retained, and whether the provider will sign a business associate agreement at all. Medical records are protected health information, so tier selection is a gating decision, not a preference — confirm the current terms on each provider’s own trust or privacy documentation, and have qualified counsel review the HIPAA posture before any real chart goes in.

Where it breaks: context limits on huge record sets, no native connection to your DMS, and no automatic page-level citations unless you engineer the pipeline. It is a strong starting point for a firm doing a handful of demands a month, and a bad fit at volume.

Option three: a custom agent on your own stack

The third path is an agent that runs inside your environment: it reads records from your document system, produces the chronology, cross-checks specials against your billing ledger, drafts the letter into your template, and writes the result back to the matter. The connective tissue is typically MCP — the Model Context Protocol, an open standard for giving an AI governed, permissioned access to your data and tools. We walk through the mechanics in connecting Claude to your practice management system and building a custom MCP server over your matter data.

Point solution (EvenUp, Supio, similar)
Fast to start. Marketed as PI-trained on messy records. Vendor absorbs model upgrades. Some include human review. But: your data goes out, the format is theirs, and it lives outside your case management system.
Custom agent over your own systems
Records stay in your DMS. Your chronology schema, your letter voice, your firm’s gap rules. Writes back to the matter. But: real build and maintenance cost, and you own quality control forever.

Be honest about when custom loses. My working assumption — and you should test it against the five packages you reconstruct below — is that if you do standard soft-tissue auto cases with predictable record sets, a point solution gets you most of the available value for a fraction of the effort. Custom earns its keep when your record sources are unusual (workers’ comp panels, hospital lien portals, non-English records), when you’ve built a demand format you believe moves adjusters, or when the demand step sits inside a broader case pipeline you’ve already automated. The same build-vs-buy logic we applied to Harvey, CoCounsel, and custom agents applies here.

Model the economics before you sign anything

Don’t take anyone’s ROI slide, including mine. Pull your last five demand packages and time-reconstruct them. Then:

P × H × C
Packages per month × hours each × blended cost per hour = your current spend
× (1 − R)
Multiply by your realistic reduction, R — verification time does not go to zero

Derive R from your own pilot, not from a vendor’s slide: time the verification minutes per package before and after, on your records, and use that delta. A borrowed number is a made-up number.

The real upside usually isn’t the labor line. It’s cycle time: if a package that took three weeks goes out in five days, cases resolve sooner and the same paralegal headcount carries more files. Model that as additional matters closed per quarter rather than hours saved, and be equally explicit about the new cost — attorney or senior paralegal verification time, which is a permanent line item, not a rollout expense.

The parts that stay human

Three more that shouldn’t be delegated: the settlement number and negotiating posture; anything requiring judgment about causation or pre-existing conditions; and the decision that the record set is complete. That last one is where a simple rules-based automation beats AI outright — a checklist tracking requested-vs-received per provider needs no model at all.

On ethics, the ABA’s Formal Opinion 512 (issued July 2024) addresses lawyers’ use of generative AI, including competence, confidentiality, supervision, and fees; ABA Model Rule 1.1 comment 8 covers technology competence and Rule 5.3 covers supervision of nonlawyer assistance. Your state bar may have its own opinion that differs — check the primary source for your jurisdiction. Two practical implications: whether you can bill for time the AI performed is worth resolving before the first invoice, and your vendor agreements and business associate arrangements deserve review by qualified counsel rather than a quick skim.

If you extend this beyond demands, the same chronology-with-citations pattern is the reusable core behind discovery triage and deposition summaries — building it once, well, is worth more than buying three tools that each do it slightly differently.

Where is your firm losing billable hours?

Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.

Get a free automation audit