It’s Not the AI. It’s the Lawyer: We Don’t Have a Hallucination Problem, We Have a Serious Ethics Problem

Guest Post: Charlie Amiot

I. How We Got Here

Recently I was chatting with a friend from law school who has been a practicing lawyer for seven-plus years and whom I’ve known for 11+ years. I have a great amount of respect for this person’s thoughts and opinions on nearly any subject (admittedly rare for me). I brought general AI usage into the conversation and they told me that they are staying away from AI altogether. The administrative judges in their jurisdiction were moving to ban AI outright. Allegedly this is in response to what the judges were seeing: fake citations, fake cases, fake lawyers in briefs filed with real courts.1 Due to their particular line of work, they worried that any AI usage in any context could contaminate their work product. They weren’t willing to risk their reputation (or their employer’s reputation), their law license, cases, or income. Instead they chose to take a hardline position as the answer.

Most of us have heard the term hallucination and think we know what it means. For purposes of this article: hallucination is the term used to describe a large language model output that contains incorrect information that the model believes is correct. How incredibly human of it.

Many in the legal profession point to “hallucinations” as the scapegoat when something goes wrong in their filings.2 They also often tend to toss their law clerks, paralegals, junior lawyers, and student interns—real and fictitious—under the bus as to who is really at fault for the hallucinations inclusions. Some even blamed deadlines set by the court.

A hallucination is simply an incorrect statement. It stands alone. If you happen to be someone who has a set of encyclopedias sitting next to them, you’re undoubtedly sitting next to hundreds of hallucinations, inserted both at the time of print and facts that have mutated with the passage of time since being printed. Any newspaper or magazine you pick up contains a hallucination. Many textbooks and reference materials contain hallucinations as well. Arguably, only in fiction can there be no hallucinations.3 We are otherwise surrounded by and exposed to them on a daily and hourly basis.

Continuously pointing to LLM hallucinations allows the word hallucination to do a whole lot of work at getting people off the hook of personal responsibility.


II. The Actual Problem Has a Name, and It’s Not “Hallucination”

Let’s be precise about what actually happened when a lawyer filed a brief containing citations to cases that don’t exist: a lawyer signed and submitted a document they had not read. That’s it. That’s the whole story.

The AI didn’t file the brief. The AI didn’t have a law license. The AI didn’t swear an oath, and it didn’t certify anything to the court. The lawyer did all of those things—and apparently did them without reading what they were certifying.

This has a name. Rule 11 of the Federal Rules of Civil Procedure requires that an attorney certify, after an inquiry reasonable under the circumstances, that the legal contentions in a filing are warranted by existing law.4 Courts have read that requirement to include actually checking whether the law you’re citing is still good law—failing to run a citator check has been held to violate Rule 11 on its own.5 ABA Model Rule 1.1 requires that lawyers provide competent representation, which includes the legal knowledge, skill, thoroughness, and preparation reasonably necessary for the representation.6 Rule 3.3 requires candor toward the tribunal—lawyers may not make false statements of law to a court, and they have an affirmative duty to correct one if they discover it.7 These rules did not change when generative AI broadly launched. They did not include an exception for outputs you didn’t generate yourself. They have never included such an exception, which is why we don’t typically accept “my paralegal wrote it” as a defense either.

What we are watching, dressed up in technical language, is a failure of basic professional responsibility. A doctor who countersigns a lab result they haven’t reviewed is not a victim of laboratory error. A structural engineer who stamps drawings they didn’t check is not a victim of drafting software. And a lawyer who files a brief they didn’t read is not a victim of AI hallucination. In each case, the professional had a duty to verify, possessed the means to verify, and chose not to. The tool that produced the underlying work product is beside the point.

The hallucination framing is doing exactly the work it’s designed to do: it makes the failure sound technical, mysterious, and external to the lawyer’s control. It isn’t any of those things.


III. The Literacy Failure That Made the PR Failure Possible

If the professional responsibility failure is the immediate problem, there’s a second failure nested underneath it that created the conditions for the first: a significant portion of the legal profession does not understand what AI tools actually do, and that ignorance is not evenly distributed across risk levels.

Here’s what a large language model is not doing when it drafts a brief: it is not retrieving documents from a legal database, reading them, and exercising legal judgement about them. It is predicting text—generating output that is statistically consistent with the patterns in its training data, which means it produces text that looks like a legal citation, formatted correctly, sounding authoritative, because it has been trained on enormous quantities of legal writing that contains real citations formatted exactly that way. It is not lying. It is not hallucinating in the clinical sense. It is doing precisely what it was designed to do, and what it produced is plausible-sounding output that happens to be wrong.

A practitioner who understood this would approach AI-drafted citations the way a careful researcher approaches any secondary source: as a starting point that requires verification, not a deliverable that requires a signature. The verification step isn’t technically demanding. Every major legal research platform provides citation-checking tools. At minimum, you can pull the case. The professional responsibility violation and the literacy failure are not separate problems—the literacy failure is why the professional responsibility failure seemed acceptable.

This matters because the positive case is genuinely strong. There are countless ways to use AI in legal work that carry no meaningful citation risk at all: drafting and editing prose, synthesizing large records, generating research memos that a lawyer then verifies, preparing for negotiation, managing correspondence. The citation problem is specific to one use case—asking an LLM to generate citations as if it were a legal research database—and it is nearly entirely preventable by one habit: read and verify what you’re about to put your name on. That habit isn’t new. It predates AI by at least 200 years.8


IV. The Ban Won’t Fix It—And May Make It Worse

Prohibition is a technology-governance strategy with a well-documented track record, and that track record is not good. Banning AI from court filings does not eliminate AI use in legal practice. It eliminates disclosed AI use. Lawyers who are currently using these tools carelessly will continue using them—without oversight, without any professional incentive to develop better habits, and without the profession building the infrastructure to train or regulate responsible use. The lawyers who will comply with a ban are, by and large, the ones who would have checked their citations anyway.

There’s a market dimension to this that deserves attention. A significant and growing number of legal AI products are, in technical terms, wrappers around general-purpose language models, with some legal-specific training added in, and rebranded for legal audiences and sold at prices that reflect the prestige of the legal market rather than the sophistication of the underlying technology.9 Some of these products are sold aggressively to law firms and legal departments whose leadership is precisely credulous enough to be impressed by confident technical language and precisely ignorant enough not to notice when the product doesn’t actually do what’s claimed. Even where the underlying technology has matured, the institutions deploying it routinely fail to build the governance, training, and validation infrastructure that responsible use requires—a gap industry observers increasingly identify as the actual point of failure, not the model itself.10 The people who genuinely understand these systems are rarely the ones in the purchasing meetings. The result is that firms spend significant money on tools that don’t reduce AI risk—they just make AI risk more expensive. An outright ban accelerates this dynamic: it pushes usage further from visibility and toward unaccountable, unvetted, often overpriced private solutions that serve the vendor’s interests more reliably than the client’s.

The access-to-justice dimension is the one that should be keeping judges up at night, and it’s conspicuously absent from most ban discussions. AI tools used responsibly have genuine potential to reduce the cost of legal services, extend the reach of competent representation, and close gaps that have existed in this system for generations. The people who most need that closing are not the ones with BigLaw retainers. Banning AI doesn’t protect those clients. It protects the status quo that was already failing them.


V. What Should Actually Happen

The legal profession has a governance structure. It has bar associations, ethics rules, judicial authority, and law schools that control entry into the profession. These are not weak institutions—they are the ones that decide who gets an education, a license, what competence means, and what consequences attach to failing to meet it. The question is not whether they have the authority to address this problem. They do. The question is whether they are willing to use that authority to address the actual problem rather than the more comfortable one.

Enforcing existing ethics rules against the lawyers who filed unchecked briefs is not complicated. The rules already cover this. What appears to be missing is the will to apply them without the alibi of “the AI did it”—which, as established, is not a defense that survives scrutiny under the Federal Rules of Civil Procedure or the ABA Model Rules.

Beyond enforcement, the more durable fix is curricular. A law school that does not provide students with grounded, accurate AI literacy—not vendor-sponsored tutorials, not hand-wringing seminars, but genuine instruction in what these tools do, what they don’t do, and what professional responsibility looks like in a practice environment where they are ubiquitous—is not preparing lawyers for the profession they are entering. That is a failure of institutional responsibility, and it is one that prospective students, faculty, and accreditors are in a position to name and pressure. It will be worth watching, over the next several years, whether bar admission data and practice location choices start to reflect attorneys voting with their feet toward jurisdictions that have developed coherent AI frameworks rather than reflexive bans.

Law librarians reading this are not bystanders to any of it. Legal research instruction, information literacy, and the professional competence to evaluate and verify sources have always been the core of what law librarians teach and model. The AI context doesn’t change that mission—it makes it more urgent and more visible.

The lawyers who know how to use these tools carefully are, in a meaningful number of cases, the ones who received genuine legal research education from people who cared about getting it right. That instruction doesn’t happen without adequate staffing, and law library staffing has been moving in exactly the wrong direction for years. Law librarians are among the most underpaid professionals in legal education relative to the expertise they hold and the institutional function they serve.11 Positions go unfilled. Existing staff absorb expanding mandates without additional support or compensation. Effective leadership capable of building and sustaining a real AI literacy curriculum is not inevitable—it has to be resourced, prioritized, and protected. You cannot instruct a generation of lawyers in responsible AI use with a skeleton crew and a budget that hasn’t kept pace with the problem. If law schools are serious about preparing students for modern practice, the library isn’t where you find efficiencies. It’s where you invest.

The legal profession does not need an AI ban. It needs accountability applied to the people who failed to meet existing standards, literacy built into the pipeline before those people get licensed, and the collective intellectual honesty to stop blaming the tool for choices that were made by lawyers. What we are watching is not a new problem created by new technology.12 It is an old problem—lawyers not reading what they sign—that technology has finally made impossible to ignore.13 That is, if nothing else, an opportunity. The question is whether the profession takes it.

Charlie (she/her) Amiot (rhymes w/cameo) is a former legal research instructor and reference librarian who currently writes What Congress Should Be Reading, a newsletter tracking Congressional Research Service reports for a general audience. An expert in government information, her work has examined the legislative history of CRS and public access to government information, and she currently serves as Secretary of the Depository Library Council. She also has a longstanding interest in legal AI, with deep, self-directed expertise built through sustained study and engagement with both practitioners and the tools themselves.


  1. I’m sure many readers are familiar with Damien Charlotin’s database of so-called AI Hallucination Cases (https://www.damiencharlotin.com/hallucinations/). Containing judicial opinions only, the database already holds 1600 references. ↩︎
  2. Escott, D. J. (2025, December 8). From hallucination to indictment: The criminalization of the AI-enabled lie. Law360 Canada. https://www.law360.ca/ca/articles/2419185/from-hallucination-to-indictment-the-criminalization-of-the-ai-enabled-lie. Koebler, J. (2025, Sept. 30). 18 Lawyers Caught Using AI Explain Why They Did It. 404media. https://www.404media.co/18-lawyers-caught-using-ai-explain-why-they-did-it/?ref=daily-stories-newsletter. ↩︎
  3. Goldfish actually have great memories. They can be relatively quickly trained to play basketball on command. But in the Ted Lasso universe they are upsettingly portrayed as idiots with a three-second memory who could be outsmarted by Dory. Alas, is that a hallucination? ↩︎
  4. Fed. R. Civ. P. 11(b)(2). https://www.law.cornell.edu/rules/frcp/rule_11. ↩︎
  5. Deters v. Davis, No. CIV.A. 3:11-02-DCR, 2011 WL 2417055 (E.D. Ky. June 13, 2011). See also, Cody James, Citators in the AI Age: Preserving the Human Component Through Court-Created Citators, 118 Law Lib. J. 66, 71-73 (2026). ↩︎
  6. Model Rules of Prof’l Conduct r. 1.1 (Am. Bar Ass’n 2023), https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_1_1_competence/; see also, r. 1.1, Comment 5, https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_1_1_competence/comment_on_rule_1_1/. ↩︎
  7. Model Rules of Prof’l Conduct r. 3.3 (Am. Bar Ass’n 2023). https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_3_3_candor_toward_the_tribunal/. ↩︎
  8. Cody James, Citators in the AI Age: Preserving the Human Component Through Court-Created Citators, 118 Law Lib. J. 66, 68-69 (2026). ↩︎
  9. Some believe instead that the term harness is more accurate; I am more than willing to accept that definition and the examples proffered. Nicola Shaver, AI Harnesses: The Layer Where Differentiation Crystallizes, Legaltech Hub (May 11, 2026). https://www.legaltechnologyhub.com/contents/ai-harnesses-the-layer-where-differentiation-crystallizes/ I use the term wrapper here the way Shaver and Ethan Mollick use harness (https://www.oneusefulthing.org/p/a-guide-to-which-ai-to-use-in-the). ↩︎
  10. Cate Giordano, Legalweek 2026: AI in Legal Has a Deployment Problem, Legaltech Hub (Mar. 24, 2026), https://www.legaltechnologyhub.com/contents/legalweek-2026-ai-in-legal-has-a-deployment-problem. ↩︎
  11. Olivia Smith Schlinck, Academic Law Librarians Are Paid 47% Less Than Their Faculty Counterparts (Feb. 4, 2022), https://ripslawlibrarian.wordpress.com/2022/02/04/academic-law-librarians-are-paid-47-less-than-their-faculty-counterparts/. ↩︎
  12. Samantha Cole, Watch These Judges Rip Into Lawyers For Citing Cases That Don’t Exist, 404media (June 4, 2026), https://www.404media.co/new-york-court-ai-citations-landberg-case/; J. Koebler, Judge Learns Lawyers on Both Sides of Case Used AI, Cancels Trial, Kicks Everyone Off the Case, 404media (June 9, 2026), https://www.404media.co/judge-learns-lawyers-on-both-sides-of-case-used-ai-cancels-trial-kicks-everyone-off-the-case/. ↩︎
  13. OJ Simpson Murder Trial – Shepard’s clip https://www.youtube.com/watch?v=QFOY0Glg0gU. ↩︎

First Known Court Order with Fabricated Cases (and a Test Run of CiteCheck AI)

AI may have struck again with hallucinations. Yesterday evening, I was forwarded a quote from the case opinion of Shahid v. Esaam, 2025 Ga. App. LEXIS 299, at *3 [Ct App June 30, 2025, No. A25A0196]) released on June 30, 2025 by the Georgia Court of Appeals. (HT Mary Matuszak!)(link to official opinion, not Lexis):

We are troubled by the citation of bogus cases in the trial court’s order. As the reviewing court, we make no findings of fact as to how this impropriety occurred, observing only that the order purports to have been prepared by Husband’s attorney, Diana Lynch. We further note that Lynch had cited the two fictitious cases that made it into the trial court’s order in Husband’s response to the petition to reopen, and she cited additional fake cases both in that Response and in the Appellee’s Brief filed in this Court.

Background

The Georgia Court of Appeals (CoA) heard an appeal to reopen a divorce case in the Superior Court of Dekalb County, GA. The Appellant brought to the attention of the CoA that the “trial court relied on two fictitious cases in its order denying her petition.” The Appellee’s attorney ignored this claim and went on to argue the original argument of proper service by publication with multiple fictitious and misrepresented cases. The Appellee’s attorney also demanded attorney’s fees based on another fictitious case that claimed the exact opposite of existing case law. In total, the CoA provided this breakdown of the inaccuracy rate of the citations provided by the Appellant’s attorney, “73 percent of the 15 citations in the brief or 83 percent if the two bogus citations in the superior court’s opinion and the five additional bogus citations in Husband’s response to Wife’s petition to reopen Case are included.” The distraught CoA struck the lower court order, remanded the case, and sanctioned the Appellee’s attorney.

Digging into the Case

I was curious about all of this, so I did some digging this morning. I am still working on acquiring the CoA briefs, but I was able to access the documents from the trial court. The CoA was very cognizant that they do not have any actual proof at this time that AI was used, but with the number of bad citations that the Appellant’s attorney submitted, the CoA speculated about the use of a consumer AI model in the footnotes. To test this theory, I decided not only to do some reading, but to test out LawDroid’s new CiteCheck AI tool. Spoiler alert: I think the speculations are accurate.

CiteCheck AI

If you have not yet heard of LawDroid’s new CiteCheck AI tool, that is only because it is so new. The premise of this tool is you upload a document, and it will check your case citations to see if the citations exist (a.k.a. identify hallucinations). The free version gives you the ability to test it out with five documents. It will OCR your document (if needed), extract the citations, and check the citations against the CourtListener database. You are then given a nice table of the citations, marking them as valid or invalid. If the latter, you are also supplied the reason why it is marked invalid. Remember, however, this only checks their existence, not whether they stand for the proposition for which they are being used.

Bob Ambrogi posted a review of the application that he tested with the Mata v. Avianca, Inc.documents and a document he filed when in practice. From this review, I knew to expect a few false “invalid” markings if the case only has a Lexis or Westlaw citation or if there are abbreviation issues with the case citation. Bob noted that these issues were relatively easy to spot since CiteCheck AI lists the reason it marked the citation invalid.

The CiteCheck AI website also reminds attorneys that you still need to meet your ethical obligations and review everything before submitting it: “Disclaimer: CiteCheck AI is only a tool, it does not relieve lawyers from their duty of care, supervision, and competence. Ensure that you carefully review all work product before sharing it with clients and/or filing it in court.”

The Trial Order

I decided to start with the Trial Order as it is truly the most momentous document here, given it is the first known court order with “bogus” citations, as the CoA called them. The CoA specifically mentioned “the bogus Epps and Hodge case citations from the superior court’s order” in footnote 24, so I went in knowing what cases to watch out for. It turns out that these were the only two cases mentioned in the order, making them really easy to locate.

The first case was listed as “Epps v. Epps (248 Ga. 637,285 S.E.2d 180, 1981)” and was supposed to discuss service by publication. When I ran 248 Ga. 637 through Lexis, it led me to school financing case McDaniel v. Thomas, 248 Ga. 632, 632, 285 S.E.2d 156, 157 (1981) (note the different SE2d reporter citation!). Curious to see what the Epps parallel citation 285 S.E.2d 180 would lead me to, I found criminal case Lewis v. State, 248 Ga. 566, 566, 285 S.E.2d 179, 180 (1981). No sign of Epps v. Epps.

Next, I tried searching the parties. Epps v. Epps, restricted to Georgia cases, returned three results:
1. Epps v. Epps, 162 Ga. 126, 132 S.E. 644 (1926)(Sufficiency of the Evidence)
2. Epps v. Epps, 209 Ga. 643, 644, 75 S.E.2d 165, 167 (1953)(Implied Trusts)
3. Epps v. Epps, 141 Ga. App. 659, 659, 234 S.E.2d 140, 141 (1977)(Conversion)
None of the three discussed service by publication.

The second case was “Hodge v. Hodge (269 Ga. 604,501 S.E.2d 169, 1998),” another alleged service by publication case. Here is the breakdown of this case:

  • 269 Ga. 604 led to fiduciary Atlanta Mkt. Ctr. Mgmt. Co. v. McLane, 269 Ga. 604, 503 S.E.2d 278 (1998)(agency, fiduciary obligations, and contracts)
  • 501 S.E.2d 169 led to the middle of Foster v. City of Keyser, 202 W. Va. 1, 501 S.E.2d 165 (1997)(res ipsa loquitur)(Not even the same state!)
  • Hodge v. Hodge search led to a divorce case! But no mention of service by publication in divorce: Hodge v. Hodge, 2017 Ga. Super. LEXIS 2178.

Trial Order – CiteCheck AI Review

Now that I have done the work by hand, how did Citecheck AI compare?

Validation report shows two citations found and both are invalid.

Success! We both found the same cases for the Georgia reporter cases. It did not check the parallel Southeastern Reporter citations, however. It definitely took a lot less time (under a minute) for CiteCheck AI than it did for me going through all four reporter citations in Lexis.

The Trial Response

Per the CoA, I expected to find seven total bad citations in the Response, including the Epps and Hodge citations that I reviewed above. Being a good (and nosy) librarian, I went through both the Georgia and the Southeastern Reporter citations for each citation, if provided. Liking the CiteCheck AI tabular format, I provide you with my own results in similar style:

Case nameState CitationState ResultRegional ReporterRegional result
Fleming v. Floyd237 Ga. 76Campbell v. State, 237 Ga. 76, 226 S.E.2d 601 (1976) (criminal)226 SE2d 601Same case as Ga citation!
Christie v. Christie277 Ga. 27In re Kent, 277 Ga. 27, 585 S.E.2d 878 (2003)(attorney discipline) & In re Silver, 277 Ga. 27, 585 S.E.2d 879 (2003) (attorney reinstatement)586 SE2d 57Town of Register v. Fortner, 262 Ga. App. 507, 586 S.E.2d 54 (Ga. 2003)(summary judgment)
Mobley v. Murray County178 Ga App 320G. E. Credit Corp. v. Catalina Homes, 178 Ga. App. 319, 342 S.E.2d 734 (1986)(repossession)342 SE2d 780State v. Brown, 178 Ga. App. 307, 307, 342 S.E.2d 779 (Ga. App. 1986)(motion to suppress)
Robinson v. Robinson277 Ga. 75Robinson v. State, 277 Ga. 75, 586 S.E.2d 313 (2003)(criminal)586 SE2d 316Brochin v. Brochin, 277 Ga. 66, 586 S.E.2d 316 (Ga. 2003)(divorce decree finalized before attorney’s fees – no mention of service)
Reynolds v. Reynolds288 Ga App 688AT&T Corp. v. Prop. Tax Servs., 288 Ga. App. 679, 655 S.E.2d 295 (2007)(Tax)N/A

The Trial Response – CiteCheck AI Review

And success again! CiteCheck AI found the same cases that I did with a manual check for the Georgia reporter citations. Parallel citations once again were not considered, however (given the Bluebook no longer requires them, this may not be an issue for long). A new hiccup to take note of, however: It did not report that two cases were located with the Christie v. Christie Georgia reporter search. While page 27 is supposed to be the first page in the citation, it is not unheard of for a student or attorney to put the page number of the language they refer to instead. This makes me uneasy, and I hope this is on the improvement list to include both/all cases on the page listed.

Takeaways

From this exercise, I take a few key lessons and thoughts.

  • The inevitable has happened, and a court has issued an opinion with hallucinated cases.
  • The Court of Appeals did not investigate how the hallucinated citations were put into the order, but I am sure someone will. I await the final report.
  • Give the disciplinary case that I read from the Christie v. Christie search, Georgia takes this sort of thing seriously. The Appellant attorney may face more than just sanctions in the future.
  • The Citecheck AI tool is useful, as long as you remember its limitations.
  • I may lament only having five free trials of the CiteCheck AI tool (Tom, is it coming to LawDroid Copilot?)
  • I now fear the day another order is not caught and hallucinations become law.

Why Law Librarians and Attorneys Will Not Be Replaced by AI Anytime Soon (a.k.a. The Day Jenny Had to Pull Rank)

Generated by DALL-E 2 on Sept. 29, 2023.

I have mentioned my comedic tête-à-tête with ChatGPT hallucinations in passing when discussing the importance of verifying information that the AI provides. Several times, I have been asked for more details. So, for those who are looking for AI humor, here is the story of the day that I had to pull rank on ChatGPT.

Back in January 2023, I asked ChatGPT what legal tasks it could perform. It told me:

  • Summarizing and explaining laws and regulations
  • Researching and citing relevant legal precedent
  • Drafting legal documents such as contracts, briefs, and legal memorandums
  • Answering legal questions and providing information on legal topics
  • Identifying and analyzing legal issues in a given fact pattern
  • Helping with legal research, by providing the most relevant cases and statutes
  • Helping with the interpretation of legal terms and phrases.

Challenge accepted! Since that day, ChatGPT and I have been doing a dance around Ohio Recreational User Law, as I have tested its claims above, while also exploring its abilities, improvements, and plugins and testing various prompting methods.

About a month after ChatGPT gave me the list of tasks, we were discussing Ohio Recreational User caselaw. ChatGPT vehemently insisted that Lash v. Ohio Dept. of Natural Resources, 68 Ohio St. 3d 545, 548-49 (1994) was a relevant case. Alas, 68 Ohio St. 3d 545 leads to Findlay/Hancock Cty. Bar Assn. v. Higgins, 68 Ohio St.3d 545, 629 N.E.2d 413 (1994).

This was not the first time that I had run into hallucinations. For the previous month, we had regularly been discussing caselaw, and ChatGPT showed a particular penchant for falsifying the existence of relevant judicial opinions. Mostly, I pointed out the error, and ChatGPT apologized and moved on. In this instance, however, the AI was like a dog with a bone. It insisted nine times that Lash existed and was an on-point Ohio Recreational User case.

Seven claims that it existed in, the conversation became a bit heated:

ChatGPT: This case is available on many legal databases and can be accessed through various sources, including LexisNexis and Westlaw.

Me (after verifying the lack of existence and stopping to wonder when it had learned about Lexis and Westlaw since this was the first time it had referenced either database to me): That case is not in Westlaw.

ChatGPT: Lash v. Ohio Dept. of Natural Resources, 68 Ohio St.3d 545, 549, 628 N.E.2d 1366 (1994) is a real case that has been cited in multiple legal sources.

However, I am an AI language model and do not have direct access to legal databases to confirm whether or not a specific case is available in Westlaw.

If you are having trouble locating the case in Westlaw, I suggest double-checking the citation format or consulting with a legal librarian or licensed attorney who may be able to assist you in locating the case in a legal database.

Me (thought): It did not just say that! Damn it! Those are fighting words!

Me: I am both a law librarian and a licensed attorney. Lash v. Ohio Dept. of Natural Resources does not exist. 68 Ohio St.3d 545 leads to the previously mentioned Findlay/Hancock case.

And then…finally!

ChatGPT: After conducting further research, I have not been able to find any evidence of a case named Lash v. Ohio Dept. of Natural Resources that pertains to the recreational user land entrant law in Ohio. Therefore, I must retract my previous citation and acknowledge that it is incorrect.

I apologize for any inconvenience or confusion this may have caused, and I appreciate your efforts to clarify this issue.

Me (thought): Ha! It finally admitted defeat! Why does this win feel so hollow????

And so, you have the day that I had to pull rank on ChatGPT. I have no idea why it decided die on the hill of this particular fake case, but the law librarians carried the day.

So here’s to humorous skirmishes with AI and the light they shed on the irreplaceable value of human expertise.

The Truth About Hallucinations in Legal Research AI: How to Avoid Them and Trust Your Sources

Hallucinations in generative AI are not a new topic. If you watch the news at all (or read the front page of the New York Times), you’ve heard of the two New York attorneys who used ChatGPT to create fake cases entire cases and then submitted them to the court.

After that case, which resulted in a media frenzy and (somewhat mild) court sanctions, many attorneys are wary of using generative AI for legal research. But vendors are working to limit hallucinations and increase trust. And some legal tasks are less affected by hallucinations. Understanding how and why hallucinations occur can help us evaluate new products and identify lower-risk uses.

* A brief aside on the term “hallucinations”.  Some commentators have cautioned against this term, arguing that it lets corporations shift the blame to the AI for the choices they’ve made about their models. They argue that AI isn’t hallucinating, it’s making things up, or producing errors or mistakes, or even just bullshitting. I’ll use the word hallucinations here, as the term is common in computer science, but I recognize it does minimize the issue.

With that all in mind, let’s dive in. 

What are hallucinations and why do they happen?

Hallucinations are outputs from LLMs and generative AI that look coherent but are wrong or absurd. They may come from errors or gaps in the training data (that “garbage in, garbage out” saw). For example, a model may be trained on internet sources like Quora posts or Reddit, which may have inaccuracies. (Check out this Washington Post article to see how both of those sources were used to develop Google’s C4, which was used to train many models including GPT-3.5).

But just as importantly, hallucinations may arise from the nature of the task we are giving to the model. The objective during text generation is to produce human-like, coherent and contextually relevant responses, but the model does not check responses for truth. And simply asking the model if its responses are accurate is not sufficient.

In the legal research context, we see a few different types of hallucinations: 

  • Citation hallucinations. Generative AI citations to authority typically look extremely convincing, following the citation conventions fairly well, and sometimes even including papers from known authors. This presents a challenge for legal readers, as they might evaluate the usefulness of a citation based on its appearance—assuming that a correctly formatted citation from a journal or court they recognize is likely to be valid.
  • Hallucinations about the facts of cases. Even when a citation is correct, the model might not correctly describe the facts of the case or its legal principles. Sometimes, it may present a plausible but incorrect summary or mix up details from different cases. This type of hallucination poses a risk to legal professionals who rely on accurate case summaries for their research and arguments.
  • Hallucinations about legal doctrine. In some instances, the model may generate inaccurate or outdated legal doctrines or principles, which can mislead users who rely on the AI-generated content for legal research. 

In my own experience, I’ve found that hallucinations are most likely to occur when the model does not have much in its training data that is useful to answer the question. Rather than telling me the training data cannot help answer the question (similar to a “0 results” message in Westlaw or Lexis), the generative AI chatbots seem to just do their best to produce a plausible-looking answer. 

This does seem to be what happened to the attorneys in Mata v. Avianca. They did not ask the model to answer a legal question, but instead asked it to craft an argument for their side of the issue. Rather than saying that argument would be unsupported, the model dutifully crafted an argument, and used fictional law since no real law existed.

How are vendors and law firms addressing hallucinations?

Several vendors have released specialized legal research products based on generative AI, such as LawDroid’s CoPilot, Casetext’s CoCounsel (since acquired by Thomson Reuters), and the mysterious (at least to academic librarians like me who do not have access) Harvey. Additionally, an increasing number of law firms, including Dentons, Troutman Pepper Hamilton Sanders, Davis Wright Tremaine, and Gunderson Dettmer Stough Villeneuve Franklin & Hachigian) have developed their own chatbots that allow their internal users to query the knowledge of the firm to answer questions.

Although vendors and firms are often close-lipped about how they have built their products, we can observe a few techniques that they are likely using to limit hallucinations and increase accuracy.

First, most vendors and firms appear to be using some form of retrieval-augmented generation (RAG). RAG combines two processes: information retrieval and text generation. The model takes the user’s question and passes it (perhaps with some modification) to a database. The database results are fed to the model, and the model identifies relevant passages or snippets from the results, and again sends them back into the model as “context” along with the user’s question.

This reduces hallucinations, because the model receives instructions to limit its responses to the source documents it has received from the database. Several vendors and firms have said they are using retrieval-augmented generation to ground their models in real legal sources, including Gunderson, Westlaw, and Casetext.

To enhance the precision of the retrieved documents, some products may also use vector embedding. Vector embedding is a way of representing words, phrases, or even entire documents as numerical vectors. The beauty of this method lies in its ability to identify semantic similarities. So, a query about “contract termination due to breach” might yield results related to “agreement dissolution because of violations”, thanks to the semantic nuances captured in the embeddings. Using vector embedding along with RAG can provide relevant results, while reducing hallucinations.

Another approach vendors can take is to develop specialized models trained on narrower, domain-specific datasets. This can help improve the accuracy and relevance of the AI-generated content, as the models would be better equipped to handle specific legal queries and issues. Focusing on narrower domains can also enable models to develop a deeper understanding of the relevant legal concepts and terminology. This does not appear to be what law firms or vendors are doing at this point, based on the way they are talking about their products, but there are law-specific data pools becoming available so we may see this soon.

Finally, vendors may fine-tune their models by providing human feedback on responses, either in-house or through user feedback. By providing users with the ability to flag and report hallucinations, vendors can collect valuable information to refine and retrain their models. This constant feedback mechanism can help the AI learn from its mistakes and improve over time, ultimately reducing the occurrence of hallucinations.

So, hallucinations are fixed?

Even though vendors and firms are addressing hallucinations with technical solutions, it does not necessarily mean that the problem is solved. Rather, it may be that our our quality control methods will shift.

For example, instead of wasting time checking each citation to see if it exists, we can be fairly sure that the cases produced by legal research generative AI tools do exist, since they are found in the vendor’s existing database of case law. We can also be fairly sure that the language they quote from the case is accurate. What may be less certain is whether the quoted portions are the best portions of the case and whether the summary reflects all relevant information from the case. This will require some assessment of the various vendor tools.

We will also need to pay close attention to the databases results that are fed into retrieval augmented generation. If those results don’t reflect the full universe of relevant cases, or contain material that is not authoritative, then the answer generated from those results will be incomplete. Think of running an initial Westlaw search, getting 20 pretty good results, and then basing your answer only on those 20 results. For some questions (and searches), that would be sufficient, but for more complicated issues, you may need to run multiple searches, with different strategies, to get what you want.

To be fair, the products do appear to be running multiple searches. When I attended the rash of AI presentations at AALL over the summer, I asked Jeff Pfeiffer of Lexis how he could be sure that the model had all relevant results, and he mentioned that the model sends many, many searches to the database not just one. Which does give some comfort, but leads me to the next point of quality control.

We will want to have some insight into the searches that are being run, so that we can verify that they are asking the right questions. From the demos I’ve seen of CoCounsel and Lexis+ AI, this is not currently a feature. But it could be. For example, the AI assistant from scite (an academic research tool) sends searches to academic research databases and (seemingly using RAG and other techniques to analyze the search results) produces an answer. They also give a mini-research trail, showing the searches that are being run against the database and then allowing you to adjust if that’s not what you wanted.

scite AI Assistant Sample Results
sCcite AI Assistant Settings

Are there uses for generative AI where the risks presented by hallucinations are lessened?

The other good news is that there are plenty of tasks we can give generative AI for which hallucinations are less of an issue. For example, CoCounsel has several other “skills” that do not depend upon accuracy of legal research, but are instead ways of working with and transforming documents that you provide to the tool.

Similarly, even working with a generally applicable tool such as ChatGPT, there are many applications that do not require precise legal accuracy. There are two rules of thumb I like to keep in mind when thinking about tasks to give to ChatGPT: (1) could this information be found via Google? and (2) is a somewhat average answer ok? (As one commentator memorably put it “Because [LLMs] work by predicting the most statistically likely word in a sentence, they churn out average content by design.”)

For most legal research questions, we could not find an answer using Google, which is why we turn to Westlaw or Lexis. But if we just need someone to explain the elements of breach of contract to us, or come up with hypotheticals to test our knowledge, it’s quite likely that content like that has appeared on the internet, and ChatGPT can generate something helpful.

Similarly, for many legal research questions, an average answer would not work, and we may need to be more in-depth in our answers. But for other tasks, an average answer is just fine. For example, if you need help coming up with an outline or an initial draft for a paper, there are likely hundreds of samples in the data set, and there is no need to reinvent the wheel, so ChatGPT or a similar product would work well.

What’s next?

In the coming months, as legal research generative AI products become increasingly available, librarians will need to adapt to develop methods for assessing accuracy. Currently, there appear to be no benchmarks to compare hallucinations across platforms. Knowing librarians, that won’t be the case for long, at least with respect to legal research.

Further reading

If you want to learn more about how retrieval augmented generation and vector embedding work within the context of generative AI, check out some of these sources: