Local AI vs Cloud AI for Law Firms: What's Safe Enough?

By Jude Lee · · Comparison

Two lawyers reviewing AI tool settings on a laptop next to a desktop workstation in a law firm office

The real question behind “can we put client data into AI”

The ethics framework here is not new or mysterious. ABA Model Rule 1.1, Comment [8] expects lawyers to keep up with the benefits and risks of relevant technology, and Model Rule 1.6(c) requires reasonable efforts to prevent unauthorized disclosure of client information. The ABA Standing Committee on Ethics and Professional Responsibility addressed generative AI directly in Formal Opinion 512, “Generative Artificial Intelligence Tools,” issued July 29, 2024 — covering competence, confidentiality, client communication, supervision of nonlawyer assistance, and fees. It is published on the ABA’s website; read the opinion itself rather than a summary. State bars have added their own layers, including the Florida Bar’s Ethics Opinion 24-1 and the State Bar of California’s practical guidance on generative AI. Read the ones that bind you, and confirm specifics with your state bar or ethics counsel rather than with a blog.

What none of those documents say is “the model must run on your own hardware.” They ask whether the arrangement — including the vendor’s terms, retention, access controls, and your own supervision — amounts to reasonable efforts. That reframes the local-versus-cloud debate from a moral question into an engineering and contracting one.

What firms are actually running today

There is no authoritative census of AI adoption at five-to-forty-lawyer firms, so treat what follows as a sketch of a common pattern rather than measured data. The stack usually has four layers. First, AI features bolted onto software the firm already pays for: practice-management assistants, Word’s Copilot, PDF summarization in Acrobat, transcription in the phone system. Second, a general-purpose assistant — Claude, ChatGPT, or Copilot — on a business or enterprise plan. Third, sometimes one point solution for a volume task: contract review, demand letters, discovery responses, citation checking. Fourth, less commonly, something custom: an agent wired into the document management system or practice-management platform through the Model Context Protocol so it can read matter files and take scoped actions.

The work that actually gets handed to AI tends to be narrow and repetitive: first-pass document review, summarizing a records production, drafting routine correspondence, converting a transcript into a chronology, triaging inbound email, turning a messy intake call into a structured matter record. Almost none of it is “AI decides something.” Nearly all of it is “AI produces a draft a lawyer checks.”

Option one: an open-weight model on your own hardware

Local AI means downloading an open-weight model — Meta’s Llama family, Mistral’s models, and other openly released weights — and running it on a machine you control with a runner like Ollama or LM Studio. No API call leaves the building. For a firm with a client mandate prohibiting cloud processing, work under a protective order with unusual handling terms, or a practice that genuinely operates offline, that property is decisive.

Be honest about the trade. Open-weight models you can run on a single workstation are generally weaker at long-document reasoning than the frontier hosted models, and the gap shows up exactly where legal work is hardest: holding a 300-page production in context, following a multi-step instruction without drifting, resisting confident invention. You also inherit the operations — GPU purchase, model updates, patching, backups, and the security of a box that now holds a copy of everything you fed it. Get a hardware quote before you assume local is cheap; a machine that comfortably runs a large open-weight model is a capital expense, not a subscription line item.

Option two: governed cloud AI with contractual controls

The mainstream path is a business, team, or enterprise tier from a major provider — Anthropic’s Claude for Work, Microsoft 365 Copilot and Azure OpenAI, Google’s Gemini on Workspace or Vertex — where the commercial terms address training use, retention, and subprocessors, and the admin console gives you SSO, user provisioning, and logs. As of 2026 these providers generally commit, on business tiers, not to train foundation models on customer content, but terms change and differ by product and region. Read the current data processing agreement for the exact SKU you are buying, and keep a dated copy in your file.

The cloud path has its own vivid failure mode, and it is rarely the vendor’s fault. A firm buys the enterprise tier, nobody switches on audit logging in the admin console, and six months later there is no record of which user pasted which document. Or the Microsoft 365 or DMS connector is set up by someone with broad rights, so the assistant inherits visibility into every matter folder that account can reach — including the ethical-wall matters. Or default retention is longer than the matter itself, so chat transcripts containing client material sit in a workspace long after the file is closed and destroyed. Buying governance is not the same as configuring it.

The free consumer tier remains the genuinely risky one. Not because the model is worse — because the terms, retention, and account controls are built for individuals, and because nobody at the firm can see what was pasted into it. That is the practical boundary where free AI tools stop being appropriate for client matter content, and it’s why the right response to staff quietly using personal accounts is to provide a governed alternative rather than issue a ban.

Open-weight model on firm hardware
  • Data location: never leaves your network; no vendor terms to negotiate.
  • Capability: weaker on long records, multi-step instructions, and resisting invention.
  • Ops burden: you own GPU capacity, patching, backups, and physical security.
  • Integrations: limited permission and audit tooling; harder to wire into Clio, NetDocuments, or iManage.
  • Best fit: air-gapped work and client mandates that forbid cloud processing.
Governed cloud assistant
  • Data location: vendor-hosted, governed by a DPA you should read and file.
  • Capability: strongest current reasoning on long productions and complex instructions.
  • Ops burden: low infrastructure work, but you must actually configure SSO, logging, and retention.
  • Integrations: connects to DMS and practice management via MCP, with scoped permissions.
  • Best fit: the bulk of ordinary client work at a small or mid-size firm.

Where agents change the calculation

All of the above is about reading. The moment you give an assistant tools — a connection to your document system, the ability to create a matter, send an email, or file something — the dominant risk shifts from “where is the text processed” to “what is this thing permitted to do, and can I reconstruct what it did.” In our view, agent deployments are running ahead of the access-control and audit practices firms have in place for them, and questions of who bears responsibility when an agent takes a harmful action are still unsettled.

For a firm, that means a local model is no safety net if you hand it write access to the DMS with no logging. Scope matters more than geography: read-only by default, write access limited to a sandbox folder or draft status, every tool call recorded. If you are building that layer yourself, the design questions are covered in our guide to building a custom MCP server over your firm’s matter data, and the human-review question in our breakdown of oversight models for supervising AI agents.

Confidentiality is a property of your controls and contracts, not of the building your GPU sits in.

A workable decision process

  1. Classify the work, not the firm

    Sort your AI use cases into three buckets: public or firm-internal content (CLE notes, marketing drafts, research on published authority), ordinary client content, and restricted content (sealed material, trade secrets under a protective order, data covered by client-specific security addenda or sector rules such as HIPAA).
  2. Check what you already promised

    Pull your engagement letters, outside counsel guidelines, and any client security addenda. Some clients have explicit AI clauses now. These often decide the question before ethics rules do.
  3. Default to governed cloud for buckets one and two

    Business-tier accounts with SSO and logging actually enabled, a filed DPA, named users, and a retention setting you chose deliberately.
  4. Reserve local models for the restricted bucket

    And only where the task is simple enough that a smaller model performs acceptably — summarizing a short document, classification, redaction assistance. Test before you rely on it: run the same ten representative documents through both the local and the cloud path, compare each output against an attorney-prepared gold-standard version, and log how many minutes of correction each one needs. If the local path costs materially more attorney time per document, you have a number to weigh against the confidentiality benefit instead of a hunch.
  5. Write down the policy and the rationale

    One page: what tools are approved, for what data class, who may use them, what must be lawyer-reviewed before it leaves the firm. Date it. Revisit when terms change.

Modeling the cost without fooling yourself

Don’t compare a subscription fee against zero. Build both sides of the ledger with your own figures.

(hours saved/week × 52) × your effective hourly rate
Upside: recovered time reallocated to billable or business-development work
Worked formula — use your own measured numbers
seats × monthly fee × 12
Cloud path: direct software cost
Your vendor quote
hardware + setup + annual admin hours × loaded staff cost
Local path: total cost of ownership
Your IT quote and staffing estimate

The honest comparison also includes a quality term nobody likes to quantify: if the local model’s first drafts need substantially more attorney correction, the time cost can swamp the subscription you avoided. The ten-document test above gives you that number.

One last point about where the recovered time goes, because it determines whether any of this pays off. Hours saved on first drafts and record review do not convert into value by themselves — they convert only if someone decides, deliberately, what fills them. Absent that decision, saved time is absorbed by administrative drift. Firms that get a return tend to redirect it somewhere specific and trackable: faster turnaround on matters already in the door, more client contact, intake response that doesn’t wait until the afternoon, or the matter work a partner has been meaning to delegate and hasn’t. Pick the destination before you buy the tool, and you can tell afterward whether it worked.

Where is your firm losing billable hours?

Get a free automation audit: we map your intake-to-invoice workflow and show you exactly what's worth automating — before you spend a dollar.

Get a free automation audit