Despite the tech industry’s persistent optimism, a growing body of research concludes that generative AI’s tendency to fabricate facts is a structural reality, not a temporary bug. The authors and analysts below argue that entirely eliminating hallucinations remains mathematically and practically impossible.

Lawyers will never eliminate all Gen AI hallucinations using current technology. They are inherent in the way Large Lanuage Models are built.

The best we can hope for is to reduce their frequency. This is very possible, and there are many ways to go about it.

One of the best approaches is using premium, purpose-built systems like Thomson Reuters CoCounsel or Lexis+ AI. These applications layer highly customized models over proprietary, closed-universe legal databases (a specialized type of Retrieval-Augmented Generation, or RAG), forcing the AI to reason over verified primary law rather than the general internet.

Even enterprise or premium, customized legal apps aren’t perfect, but they have improved since the well-known Stanford study. Today’s best products do a better job, as explained in the new book The AI-Ready Law Firm: How Lawyers Use Generative Tools to Work Smarter and Serve Clients Better (ABA 2026).

Some lawyers report success with prompts that include language like “Act as a meticulous appellate attorney who values absolute factual precision over narrative flow.”

I have yet to find a excellent article on this topic, but vendors Machine Learning Mastery, De Community, Voiceflow, and Parasoft have some pretty good additional suggestions.

I’ll be explaining my favorite technique, reducing “answer pressure,” in a future post.

LawDroid founder and access to justice advocate Tom Martin has seen a lot of organizational organizational inertia in his day:

People keep doing the same things they’re doing. They keep buying from the same company. They keep wanting to use their tools in the same way.

Inertia is powerful indeed. I’ve found one small way to push against this. I tell every new hire:

“Your fresh eyes are valuable. You will notice problems and opportunities that may not be obvious to people who have been here five years. Make a note of them. We’ll talk in a few weeks, and then again in a few months.”

When new employees suggest better approaches, it’s often gold. Encourage them. The ability to see beyond “the way we’ve always done it” is a rare treasure.

Especially when you are willing to listen before those fresh eyes become accustomed to the scenery.

Are blogs passé? I’d argue that blogs are still one of the most effective — and underrated — marketing platforms for lawyers.

Social media is ephemeral. An Instagram post today will be old news tomorrow. A blog offers permanence. It will always be there and always under your control. They make an ideal hub for all your social media marketing. A short social media post can link to a blog post, giving potential clients more information about you and your firm.

Proprietary platforms like Medium and, more recently, Substack appeal to many, but the wise avoid getting locked into a host that can become a trap. Adding a custom domain name can help, but mapping URLs to a new host is a headache, so it’s far from a perfect solution.

Blogs are an easy and reliable way to archive articles, slide decks, and more. You don’t need to worry about a platform that may disappear or become undesirable because the site where you originally published it decides to add a paywall.

Don’t let your valuable content disappear or become lost. Keep your work at your fingertips, not at the mercy of changing platforms or disappearing websites.

David Arato (Legal Content Marketing) has some additional thoughts at Attorney At Work.

Is legal blogging still a viable client acquisition strategy in 2025?
Yes, legal blogging remains viable as a client acquisition strategy in 2025, but its effectiveness depends on execution and integration with your overall marketing approach. While blogging alone might not drive the same volume of direct client inquiries it did a decade ago, it serves critical functions in the marketing ecosystem: establishing authority, supporting SEO efforts, providing content for other channels, and educating potential clients during their decision-making process.
Why does legal blogging still matter in 2025?
Legal blogging still matters in 2025 for several reasons: 1) It demonstrates expertise in a saturated market; 2) SEO value remains significant; 3) it builds trust through consistency; and 4) it integrates with modern marketing channels. Recent data indicates that companies with regularly updated blogs generate 67% more leads than those without.

Anyone writing seriously about AI and the law learns to watch certain bylines. Michael Murray’s is one of them, and his latest book just moved to the top of my reading list.

Legal Issues of Generative and Agentic Artificial Intelligence (July 2026) is hot off the presses. It tracks the fastest-moving questions in American law, from the ChatGPT moment through the arrival of autonomous agents that research, negotiate, draft, and file on their own.

It’s the third in a series worth owning complete:

Murray is a senior law professor at the University of Kentucky, where he leads its AI and the Law Project. Blessedly, despite the professorship, his prose is highly readable — perhaps because he was a litigator and federal law clerk before he was an academic. He writes for people who practice law, not just people who study it.

Make room for all three on your bookshelf. Highly recommended.

“Human in the loop” has quickly become one of the most reassuring phrases in the modern AI vocabulary. It suggests prudence, restraint, and—above all—control. If a human must approve the system’s actions, what could go wrong?

Human in the loop, often shortened to HITL, describes any arrangement in which a person reviews or authorizes an AI system’s output before it takes effect. In high-stakes professional work, it has become the standard reassurance offered whenever someone worries aloud about AI: there will always be a human checking. The problem is that in many real-world deployments, HITL functions less as a safeguard than as a slogan. Here are some of the reasons why HITL is not everything it is often cracked up to be:

Why Having a Human In The Loop Won’t Always Save You

Automation Bias. Humans tend to trust outputs that appear polished, confident, and complete. Modern AI systems excel at producing exactly that kind of output. A well-structured answer, complete with plausible citations and a professional tone, invites acceptance. The features that make these tools useful also make them dangerous.

Mata v. Avianca, the leading case on AI hallucinations, is usually told as a story about an AI inventing cases. The real issue is that the human reviewer was the safeguard that failed.

Cognitive Overload. In practice, users of AI systems are rarely in a position to conduct careful, line-by-line verification of every output. They are busy professionals, often operating under time pressure. When AI tools are integrated into workflows that generate frequent outputs, the review process can degrade into a form of triage: approve unless something obviously looks wrong.

Scope Illusion. Users may believe they are reviewing the entirety of a decision when, in fact, they are only seeing a surface-level summary. The underlying assumptions, intermediate steps, and data sources may remain opaque. The human is “in the loop,” but only within a narrow slice of the process.

Speed Asymmetry. AI systems can generate outputs and take intermediate steps far more quickly than humans can meaningfully evaluate them. As systems scale, the human reviewer becomes a bottleneck. The natural organizational response is to streamline or reduce review, sometimes informally. Over time, scrutiny diminishes as trust increases—a paradox familiar to anyone who has studied risk management.

Why HITL Is Essential, Even Though It Is Far From Perfect

In high-stakes professional settings like the practice of law, the expectation of human judgment is not going away. These are domains where accountability, context, and ethical reasoning matter in ways that current AI systems cannot fully replicate.

HITL can be valuable if implemented thoughtfully. This includes:

  • The human reviewer must have sufficient time and incentive to conduct a real review.
  • The system must provide transparency—sources, reasoning, or at least a clear basis for its outputs.
  • The human must have both the authority and the willingness to override the system.
  • The volume of decisions must be manageable enough to permit careful scrutiny.

HITL won’t provide much help in the absence of these conditions. Remove any one and you are back to the slogan.

Conclusion

None of this is an argument against human oversight. It is essential.

Human judgment has never been infallible. Errors, biases, and rubber-stamping long predate AI, but AI introduces failure modes of its own, stacking them on top of the old ones. Oversight is fragile in both directions: the human can fail the machine, and the machine can defeat the human.

The question is not whether a human is present. It is whether the conditions that make a human’s presence meaningful are actually met—and whether anyone has checked.

What are AI agents and how do they differ from AI apps most of us have been using for the past few years?

AI Chatbots like Claude, ChatGPT, and Gemini are built on large language models. They respond to prompts and provide answers. They give us information, but don’t take any actions.

AI Agents are built on top of chatbots, but they can do more than answer questions. They act on our behalf, use tools, follow multi-step goals, interact with outside systems, and often take actions with limited human intervention.

Agentic AI refers to systems that control multiple agents and use them to accomplish more ambitious goals.

A standard AI chatbot is like sitting next to a student driver—you must constantly guide, alert, and correct them. An AI agent is like a hired driver: you hand over the keys, set the destination, and it manages the route, traffic, and step-by-step decisions. If there is a traffic slowdown ahead, the hired driver may select an alternate route.

Some sample prompts illustrate the distinction. You might prompt an AI chatbot like this: 

Summarize how the federal courts of appeals have split on applying the consumer-expectations test versus the risk-utility test in design-defect products liability claims, and note which circuits favor which approach.

The chatbot reads, reasons, and returns a summary. The exchange begins and ends with text. Whatever the answer’s quality, it has taken no action in the world—you remain the only actor, free to verify, discard, or rely on what it produced.

An agent receives something closer to an instruction than a question—a delegation of authority. We give it a goal and authorize it to pursue it across multiple steps, using tools, often without pausing for our approval. The vocabulary shift is not cosmetic. We prompt a chatbot; we task an agent, much as we delegate to an associate and then answer for the result. We might instruct an AI agent like this:

Monitor my client-matter inbox. When a new email arrives, read it, pull the relevant documents from our case files and billing records, draft a substantive reply, and send it automatically so I needn’t review routine correspondence.

AI agents pose significant new security risks, for reasons we’ll explain in a series of articles.

Most lawyers encounter artificial intelligence through cloud services like Claude, Gemini or ChatGPT. They type a question, upload a document, and receive an answer produced on computers operated by someone else.

This arrangement is convenient. It also means that the lawyer is relying on an outside company to process the information, maintain appropriate security, follow its stated retention policies, and continue offering the service on acceptable terms. However, many lawyers would rather not share information with third-party vendors, even when the terms of service imply confidentiality.

There is now another practical option. Lawyers can run increasingly capable AI models directly on their own computers.

What “Running AI Locally” Means

A cloud AI service processes a request on the provider’s computers. A local AI application processes the request on the user’s computer.

The computer may still be connected to the internet. It can receive email, open websites, and download updates. What matters is where the AI model is operating and whether the lawyer’s prompt and documents must be sent to an outside company.

This distinction is easy to overlook. Installing an application on a computer does not necessarily mean the AI itself is running on that computer. Some desktop applications are merely gateways to cloud services. Others allow the user to choose between local and cloud models.

Lawyers should therefore ask a simple question before using any AI application for sensitive work:

Is this information being processed on my computer, or is it being sent somewhere else? If you are not completely comfortable with the answer, you should be considering running your AI apps locally.

Why Local Processing Can Matter

The strongest argument for local AI is not that cloud services are unsafe. Reputable providers may offer strong security, enterprise controls, favorable contractual terms, and useful privacy protections.

The advantage of local processing is greater control. When a model runs locally, the lawyer may be able to avoid sending the contents of a document or prompt to the model provider. That reduces the number of outside parties and systems involved in handling the information.

For lawyers, that can simplify the confidentiality analysis. It does not eliminate the duty to secure the computer, supervise staff, manage backups, and understand how the software works. But it may reduce one important category of risk: the risk of transmitting client information to a remote AI service.

Enterprise AI providers spend billions of dollars annually defending their infrastructure against sophisticated cyber threats. When a firm chooses to keep its most sensitive tasks in-house, it assumes the full defensive burden. Hoarding confidential summaries on an unencrypted, locally

synchronized hard drive is not a security strategy—it is simply a bespoke vulnerability. Local AI successfully eliminates the risk of third-party data transmission, but it demands an internal IT posture robust enough to compensate for the loss of enterprise-grade armor.

More Control Over Confidential Documents

Local AI may be useful for tasks involving material that a lawyer is unwilling or unauthorized to upload to a general-purpose chatbot. Possible uses include:

  • summarizing a deposition;
  • creating a chronology;
  • extracting names, dates, and obligations;
  • comparing two contract drafts;
  • reorganizing notes;
  • rewriting correspondence;
  • classifying documents;
  • turning rough material into an outline.

A local model can perform these tasks without necessarily sending the underlying documents to an outside AI provider. That does not make its answers reliable merely because they are private. Local models can misunderstand documents, omit important details, and invent facts just as cloud models can.

Privacy and accuracy are different questions. The lawyer must still review the result.

More Control Over Retention

Cloud providers differ in how they store prompts, documents, account information, and usage records. Their practices may also vary by subscription plan and product setting. A local application may give the user more direct control over:

● whether conversations are saved;

● where files and logs are stored;

● who can access them;

● whether they are included in backups;

● how long they are retained;

● when they are deleted.

That control is useful, but it is not automatic. A supposedly private conversation may remain on an unencrypted hard drive. It may be copied into a cloud backup. It may be placed in a folder synchronized through iCloud, OneDrive, Dropbox, or another service.

The lesson is not that local AI eliminates data-management problems. It is that the law firm has more power to make its own decisions about them.

The Hardware Reality Check

The cloud, as the adage goes, is merely someone else’s computer. The corollary is that bringing AI in-house requires relying heavily on your own. Modern large language models possess a healthy appetite for unified memory and dedicated graphics processing units (GPUs).

The standard-issue law office laptop—optimized primarily for word processing and endless email chains—may find itself gasping for air when asked to synthesize a 50-page deposition. Upgrading to machines capable of running capable models, such as desktops with robust Apple M-series chips or dedicated GPUs, is often necessary to prevent document review from feeling like a dial-up experience. Local AI may not require a monthly subscription fee, but it still requires a significant capital expense to run smoothly.

The Apple Silicon Advantage

I’m far from being an Apple fanboy. However, an article like this needs to address an important related issue: Windows vs. Mac:

When outfitting an office for local AI, the traditional computing divide between Windows and macOS takes on new dimensions. While standard desktop PCs have long been the workhorses of the legal profession, Apple’s recent hardware architecture offers a distinct, albeit premium, advantage for running large language models.

The secret lies in unified memory. In a conventional PC setup, the central processor and the graphics card maintain separate pools of memory. Running a formidable 70-billion-parameter model locally typically requires stringing together multiple specialized, power-hungry graphics cards—a setup that often generates the heat and ambient noise of a commercial jet engine and looks distinctly out of place next to a mahogany credenza.

Apple’s M-series chips, by contrast, share a single, massive pool of memory across the entire system. A Mac Studio equipped with 128GB or 192GB of unified memory can comfortably load models that would choke most conventional workstations, doing so with a quiet efficiency that borders on the unsettling. For a law firm, this translates to desktop hardware capable of serious AI workloads without requiring a dedicated server room or a specialized cooling apparatus.

There are concessions, of course. The financial toll at checkout is not for the faint of heart, and the architecture is notoriously inflexible; one cannot simply pry open a Mac to add more memory later when models inevitably grow larger. Furthermore, the ecosystem remains highly controlled, lacking the chaotic customizability of open-source PC builds.

Yet, for a profession that typically values reliability and quiet competence over tinkering, these drawbacks are often secondary. If the goal is deploying robust, local AI capability with minimal friction and maximum hardware efficiency, Apple’s silicon currently provides a uniquely elegant, if expensive, path of least resistance.

More Predictable Costs

Local AI is sometimes described as free. That is misleading. The models may be available without a monthly subscription, but the computer, electricity, setup, maintenance, and support all cost money.

Still, local use can make expenses more predictable. That can be appealing when:

● the firm already owns capable computers;

● AI use is frequent;

● smaller models are adequate for the work;

● confidential material cannot conveniently be sent to cloud services;

● the firm wants to reduce dependence on subscription limits and pricing changes.

The point is not that local AI is always cheaper. Cloud services may remain the better bargain for occasional users or for work requiring the most capable models. The main benefit is better control over your information.

Less Dependence on One Vendor

Cloud AI services can change rapidly. Providers may revise prices, usage limits, features, model behavior, and contract terms. Running a model locally gives a firm more choice. It may be possible to compare models, preserve access to a preferred version, and use the same model through different applications.

That doesn’t eliminate dependence on vendors. The firm still relies on software developers, model publishers, hardware manufacturers, and technical support. However, it can reduce dependence on any single cloud service. This matters to many firms.

Internet Access Can Make Local AI More Useful

A local model does not become less local merely because the computer is online. In fact, internet access makes local use more practical. The computer can receive security updates, download new models, use approved research services, and communicate with ordinary office software.

This makes possible a hybrid workflow. A lawyer might:

1. analyze a confidential document locally;

2. conduct legal research through an approved online service;

3. use a cloud model for a difficult nonconfidential task;

4. review and combine the results personally.

That approach recognizes that local and cloud AI have different strengths. Cloud models usually offer greater capability and convenience. Local models offer greater control. A law firm does not have to choose only one.

What Local Models Do Well

The best cloud systems remain more capable than the smaller models most people can run on ordinary computers. Even so, local models may be entirely adequate for many routine tasks, including:

● rewriting;

● summarizing;

● extracting information;

● changing format or tone;

● organizing notes;

● comparing passages;

● generating preliminary outlines.

The performance gap becomes more noticeable when the work requires difficult reasoning, very long documents, current information, sophisticated research, or advanced image and audio analysis.

That suggests another useful principle: Use the smallest and simplest system that performs the task adequately. Not every assignment requires the most powerful model available.

Taking the First Step

For the practitioner inclined to test these waters, the barrier to entry is no longer a degree in computer science. Several accessible applications now serve as user-friendly gateways for running open-weight models locally.

Software such as LM Studio, GPT4All, and Ollama allows a user to download capable models—such as Meta’s Llama 3 or Mistral—and interact with them through an interface virtually indistinguishable from familiar web chatbots. These tools are free to download and operate. More importantly, they provide a practical, low-friction environment for a firm to determine whether its existing hardware, and its workflow, are truly suited for local processing before making a larger commitment.

Local AI Is Not Nirvana

Local models may be slower. They may require more setup. They may lack automatic access to current information. They may be less capable than leading cloud models. The user may also receive less technical support.

Most important, local models remain generative AI systems. They can make mistakes, misunderstand instructions, and produce confident nonsense. Hallucinations are inherent in LLMs. They are not going away any time soon. Professiona judgment still matters. Don’t just verify the accuracy of the citation; make sure the case supports the proposition that you are citing it for. 

The Apple Silicon Advantage 

When outfitting an office for local AI, the traditional computing divide between Windows and macOS takes on new dimensions. While standard desktop PCs have long been the workhorses of the legal profession, Apple’s recent hardware architecture offers a distinct, albeit premium, advantage for running large language models.

The secret lies in unified memory. In a conventional PC setup, the central processor and the graphics card maintain separate pools of memory. Running a formidable 70-billion-parameter model locally typically requires stringing together multiple specialized, power-hungry graphics cards—a setup that often generates the heat and ambient noise of a commercial jet engine and looks distinctly out of place next to a mahogany credenza.

Apple’s M-series chips, by contrast, share a single, massive pool of memory across the entire system. A Mac Studio equipped with 128GB or 192GB of unified memory can comfortably load models that would choke most conventional workstations, doing so with a quiet efficiency that borders on the unsettling. For a law firm, this translates to desktop hardware capable of serious AI workloads without requiring a dedicated server room or a specialized cooling apparatus.

There are concessions, of course. The financial toll at checkout is not for the faint of heart, and the architecture is notoriously inflexible; one cannot simply pry open a Mac to add more memory later when models inevitably grow larger. Furthermore, the ecosystem remains highly controlled, lacking the chaotic customizability of open-source PC builds.

Yet, for a profession that typically values reliability and quiet competence over tinkering, these drawbacks are often secondary. If the goal is deploying robust, local AI capability with minimal friction and maximum hardware efficiency, Apple’s silicon currently provides a uniquely elegant, if expensive, path of least resistance.

The Likely Future Is Hybrid

Cloud AI will continue to offer some advantages. It provides access to enormous computing resources, sophisticated tools, and rapidly improving models. Local AI offers something different: greater control over selected information, retention practices, model choice, and cost. Some law firms will want the best of both worlds. Cloud models can be used when their superior capabilities and convenience justify their use. Local models can be used when privacy, control, or predictable use matters more.

The key point is to understand that you have options. Don’t assume that every task belongs in the cloud. Choose the right tool for the right task.

==============

Some Plain-English Supplemental Resources

A Beginner’s Guide to Running AI Models on Your PC. This guide explains what a “local AI” is using zero jargon.

Offline AI Made Easy: A Layman’s Explainer. This article demystifies how a computer can run a chatbot completely offline. It breaks down how data stays entirely on the hard drive, making it a highly reassuring read for professionals worried about confidentiality and data security.

For a visual walkthrough designed for non-technical users, the video “Local LLM Beginner Guide: You Can Actually Run AI For Free” provides a straightforward, step-by-step demonstration that explains how anyone can download and run an AI model entirely on their own machine, without a tech background.

An AI agent can be attacked by an email you never open. No click. No download. No opened attachment. Just a “confused deputy,” an agent that has failed to distinguish a prompt from data, a problem known as prompt injection.

Three Prompt Injection Vulnerability Examples

Security researchers recently demonstrated prompte injection mechanisms with ShadowLeak: a single email containing hidden instructions. When the recipient later asked ChatGPT’s Deep Research agent to review Gmail, the agent read the hidden prompt and quietly exfiltrated inbox data from OpenAI’s cloud — where ordinary endpoint defenses would not detect it.

A later Claude.ai demonstration required even less. Researchers at Oasis Security showed how hidden instructions, triggered via a Google search-and-ad path, could be chained together to extract a user’s private conversation history. No integrations. No MCP servers. A default account.

Then came another warning sign (also explained in the Oasis Security article): researchers at PromptArmor demonstrated how a malicious document could prompt Claude Cowork to upload confidential files via the agent’s own authorized access.

The Pattern

test

Different products. Different attack paths. Same basic problem.

The agent reads untrusted content. The content contains hidden instructions. The agent has permissions. The attacker tries to make those permissions work for them.

Some of these specific vulnerabilities have been patched. The point is that prompt injection is not just another software bug waiting for next quarter’s update. It is a recurring security problem built into the basic design of agentic AI.

An agent that reads thousands of emails, PDFs, websites, and file attachments faces a constant stream of potential attacks. One success can mean the disclosure of client confidences, a waived privilege, or a missed deadline.

My article “Prompt Injection: What Lawyers Considering Agentic AI Must Know” provides a broader overview of these issues.