It’s Not the AI. It’s the Lawyer: We Don’t Have a Hallucination Problem, We Have a Serious Ethics Problem

Guest Post: Charlie Amiot

I. How We Got Here

Recently I was chatting with a friend from law school who has been a practicing lawyer for seven-plus years and whom I’ve known for 11+ years. I have a great amount of respect for this person’s thoughts and opinions on nearly any subject (admittedly rare for me). I brought general AI usage into the conversation and they told me that they are staying away from AI altogether. The administrative judges in their jurisdiction were moving to ban AI outright. Allegedly this is in response to what the judges were seeing: fake citations, fake cases, fake lawyers in briefs filed with real courts.1 Due to their particular line of work, they worried that any AI usage in any context could contaminate their work product. They weren’t willing to risk their reputation (or their employer’s reputation), their law license, cases, or income. Instead they chose to take a hardline position as the answer.

Most of us have heard the term hallucination and think we know what it means. For purposes of this article: hallucination is the term used to describe a large language model output that contains incorrect information that the model believes is correct. How incredibly human of it.

Many in the legal profession point to “hallucinations” as the scapegoat when something goes wrong in their filings.2 They also often tend to toss their law clerks, paralegals, junior lawyers, and student interns—real and fictitious—under the bus as to who is really at fault for the hallucinations inclusions. Some even blamed deadlines set by the court.

A hallucination is simply an incorrect statement. It stands alone. If you happen to be someone who has a set of encyclopedias sitting next to them, you’re undoubtedly sitting next to hundreds of hallucinations, inserted both at the time of print and facts that have mutated with the passage of time since being printed. Any newspaper or magazine you pick up contains a hallucination. Many textbooks and reference materials contain hallucinations as well. Arguably, only in fiction can there be no hallucinations.3 We are otherwise surrounded by and exposed to them on a daily and hourly basis.

Continuously pointing to LLM hallucinations allows the word hallucination to do a whole lot of work at getting people off the hook of personal responsibility.


II. The Actual Problem Has a Name, and It’s Not “Hallucination”

Let’s be precise about what actually happened when a lawyer filed a brief containing citations to cases that don’t exist: a lawyer signed and submitted a document they had not read. That’s it. That’s the whole story.

The AI didn’t file the brief. The AI didn’t have a law license. The AI didn’t swear an oath, and it didn’t certify anything to the court. The lawyer did all of those things—and apparently did them without reading what they were certifying.

This has a name. Rule 11 of the Federal Rules of Civil Procedure requires that an attorney certify, after an inquiry reasonable under the circumstances, that the legal contentions in a filing are warranted by existing law.4 Courts have read that requirement to include actually checking whether the law you’re citing is still good law—failing to run a citator check has been held to violate Rule 11 on its own.5 ABA Model Rule 1.1 requires that lawyers provide competent representation, which includes the legal knowledge, skill, thoroughness, and preparation reasonably necessary for the representation.6 Rule 3.3 requires candor toward the tribunal—lawyers may not make false statements of law to a court, and they have an affirmative duty to correct one if they discover it.7 These rules did not change when generative AI broadly launched. They did not include an exception for outputs you didn’t generate yourself. They have never included such an exception, which is why we don’t typically accept “my paralegal wrote it” as a defense either.

What we are watching, dressed up in technical language, is a failure of basic professional responsibility. A doctor who countersigns a lab result they haven’t reviewed is not a victim of laboratory error. A structural engineer who stamps drawings they didn’t check is not a victim of drafting software. And a lawyer who files a brief they didn’t read is not a victim of AI hallucination. In each case, the professional had a duty to verify, possessed the means to verify, and chose not to. The tool that produced the underlying work product is beside the point.

The hallucination framing is doing exactly the work it’s designed to do: it makes the failure sound technical, mysterious, and external to the lawyer’s control. It isn’t any of those things.


III. The Literacy Failure That Made the PR Failure Possible

If the professional responsibility failure is the immediate problem, there’s a second failure nested underneath it that created the conditions for the first: a significant portion of the legal profession does not understand what AI tools actually do, and that ignorance is not evenly distributed across risk levels.

Here’s what a large language model is not doing when it drafts a brief: it is not retrieving documents from a legal database, reading them, and exercising legal judgement about them. It is predicting text—generating output that is statistically consistent with the patterns in its training data, which means it produces text that looks like a legal citation, formatted correctly, sounding authoritative, because it has been trained on enormous quantities of legal writing that contains real citations formatted exactly that way. It is not lying. It is not hallucinating in the clinical sense. It is doing precisely what it was designed to do, and what it produced is plausible-sounding output that happens to be wrong.

A practitioner who understood this would approach AI-drafted citations the way a careful researcher approaches any secondary source: as a starting point that requires verification, not a deliverable that requires a signature. The verification step isn’t technically demanding. Every major legal research platform provides citation-checking tools. At minimum, you can pull the case. The professional responsibility violation and the literacy failure are not separate problems—the literacy failure is why the professional responsibility failure seemed acceptable.

This matters because the positive case is genuinely strong. There are countless ways to use AI in legal work that carry no meaningful citation risk at all: drafting and editing prose, synthesizing large records, generating research memos that a lawyer then verifies, preparing for negotiation, managing correspondence. The citation problem is specific to one use case—asking an LLM to generate citations as if it were a legal research database—and it is nearly entirely preventable by one habit: read and verify what you’re about to put your name on. That habit isn’t new. It predates AI by at least 200 years.8


IV. The Ban Won’t Fix It—And May Make It Worse

Prohibition is a technology-governance strategy with a well-documented track record, and that track record is not good. Banning AI from court filings does not eliminate AI use in legal practice. It eliminates disclosed AI use. Lawyers who are currently using these tools carelessly will continue using them—without oversight, without any professional incentive to develop better habits, and without the profession building the infrastructure to train or regulate responsible use. The lawyers who will comply with a ban are, by and large, the ones who would have checked their citations anyway.

There’s a market dimension to this that deserves attention. A significant and growing number of legal AI products are, in technical terms, wrappers around general-purpose language models, with some legal-specific training added in, and rebranded for legal audiences and sold at prices that reflect the prestige of the legal market rather than the sophistication of the underlying technology.9 Some of these products are sold aggressively to law firms and legal departments whose leadership is precisely credulous enough to be impressed by confident technical language and precisely ignorant enough not to notice when the product doesn’t actually do what’s claimed. Even where the underlying technology has matured, the institutions deploying it routinely fail to build the governance, training, and validation infrastructure that responsible use requires—a gap industry observers increasingly identify as the actual point of failure, not the model itself.10 The people who genuinely understand these systems are rarely the ones in the purchasing meetings. The result is that firms spend significant money on tools that don’t reduce AI risk—they just make AI risk more expensive. An outright ban accelerates this dynamic: it pushes usage further from visibility and toward unaccountable, unvetted, often overpriced private solutions that serve the vendor’s interests more reliably than the client’s.

The access-to-justice dimension is the one that should be keeping judges up at night, and it’s conspicuously absent from most ban discussions. AI tools used responsibly have genuine potential to reduce the cost of legal services, extend the reach of competent representation, and close gaps that have existed in this system for generations. The people who most need that closing are not the ones with BigLaw retainers. Banning AI doesn’t protect those clients. It protects the status quo that was already failing them.


V. What Should Actually Happen

The legal profession has a governance structure. It has bar associations, ethics rules, judicial authority, and law schools that control entry into the profession. These are not weak institutions—they are the ones that decide who gets an education, a license, what competence means, and what consequences attach to failing to meet it. The question is not whether they have the authority to address this problem. They do. The question is whether they are willing to use that authority to address the actual problem rather than the more comfortable one.

Enforcing existing ethics rules against the lawyers who filed unchecked briefs is not complicated. The rules already cover this. What appears to be missing is the will to apply them without the alibi of “the AI did it”—which, as established, is not a defense that survives scrutiny under the Federal Rules of Civil Procedure or the ABA Model Rules.

Beyond enforcement, the more durable fix is curricular. A law school that does not provide students with grounded, accurate AI literacy—not vendor-sponsored tutorials, not hand-wringing seminars, but genuine instruction in what these tools do, what they don’t do, and what professional responsibility looks like in a practice environment where they are ubiquitous—is not preparing lawyers for the profession they are entering. That is a failure of institutional responsibility, and it is one that prospective students, faculty, and accreditors are in a position to name and pressure. It will be worth watching, over the next several years, whether bar admission data and practice location choices start to reflect attorneys voting with their feet toward jurisdictions that have developed coherent AI frameworks rather than reflexive bans.

Law librarians reading this are not bystanders to any of it. Legal research instruction, information literacy, and the professional competence to evaluate and verify sources have always been the core of what law librarians teach and model. The AI context doesn’t change that mission—it makes it more urgent and more visible.

The lawyers who know how to use these tools carefully are, in a meaningful number of cases, the ones who received genuine legal research education from people who cared about getting it right. That instruction doesn’t happen without adequate staffing, and law library staffing has been moving in exactly the wrong direction for years. Law librarians are among the most underpaid professionals in legal education relative to the expertise they hold and the institutional function they serve.11 Positions go unfilled. Existing staff absorb expanding mandates without additional support or compensation. Effective leadership capable of building and sustaining a real AI literacy curriculum is not inevitable—it has to be resourced, prioritized, and protected. You cannot instruct a generation of lawyers in responsible AI use with a skeleton crew and a budget that hasn’t kept pace with the problem. If law schools are serious about preparing students for modern practice, the library isn’t where you find efficiencies. It’s where you invest.

The legal profession does not need an AI ban. It needs accountability applied to the people who failed to meet existing standards, literacy built into the pipeline before those people get licensed, and the collective intellectual honesty to stop blaming the tool for choices that were made by lawyers. What we are watching is not a new problem created by new technology.12 It is an old problem—lawyers not reading what they sign—that technology has finally made impossible to ignore.13 That is, if nothing else, an opportunity. The question is whether the profession takes it.

Charlie (she/her) Amiot (rhymes w/cameo) is a former legal research instructor and reference librarian who currently writes What Congress Should Be Reading, a newsletter tracking Congressional Research Service reports for a general audience. An expert in government information, her work has examined the legislative history of CRS and public access to government information, and she currently serves as Secretary of the Depository Library Council. She also has a longstanding interest in legal AI, with deep, self-directed expertise built through sustained study and engagement with both practitioners and the tools themselves.


  1. I’m sure many readers are familiar with Damien Charlotin’s database of so-called AI Hallucination Cases (https://www.damiencharlotin.com/hallucinations/). Containing judicial opinions only, the database already holds 1600 references. ↩︎
  2. Escott, D. J. (2025, December 8). From hallucination to indictment: The criminalization of the AI-enabled lie. Law360 Canada. https://www.law360.ca/ca/articles/2419185/from-hallucination-to-indictment-the-criminalization-of-the-ai-enabled-lie. Koebler, J. (2025, Sept. 30). 18 Lawyers Caught Using AI Explain Why They Did It. 404media. https://www.404media.co/18-lawyers-caught-using-ai-explain-why-they-did-it/?ref=daily-stories-newsletter. ↩︎
  3. Goldfish actually have great memories. They can be relatively quickly trained to play basketball on command. But in the Ted Lasso universe they are upsettingly portrayed as idiots with a three-second memory who could be outsmarted by Dory. Alas, is that a hallucination? ↩︎
  4. Fed. R. Civ. P. 11(b)(2). https://www.law.cornell.edu/rules/frcp/rule_11. ↩︎
  5. Deters v. Davis, No. CIV.A. 3:11-02-DCR, 2011 WL 2417055 (E.D. Ky. June 13, 2011). See also, Cody James, Citators in the AI Age: Preserving the Human Component Through Court-Created Citators, 118 Law Lib. J. 66, 71-73 (2026). ↩︎
  6. Model Rules of Prof’l Conduct r. 1.1 (Am. Bar Ass’n 2023), https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_1_1_competence/; see also, r. 1.1, Comment 5, https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_1_1_competence/comment_on_rule_1_1/. ↩︎
  7. Model Rules of Prof’l Conduct r. 3.3 (Am. Bar Ass’n 2023). https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_3_3_candor_toward_the_tribunal/. ↩︎
  8. Cody James, Citators in the AI Age: Preserving the Human Component Through Court-Created Citators, 118 Law Lib. J. 66, 68-69 (2026). ↩︎
  9. Some believe instead that the term harness is more accurate; I am more than willing to accept that definition and the examples proffered. Nicola Shaver, AI Harnesses: The Layer Where Differentiation Crystallizes, Legaltech Hub (May 11, 2026). https://www.legaltechnologyhub.com/contents/ai-harnesses-the-layer-where-differentiation-crystallizes/ I use the term wrapper here the way Shaver and Ethan Mollick use harness (https://www.oneusefulthing.org/p/a-guide-to-which-ai-to-use-in-the). ↩︎
  10. Cate Giordano, Legalweek 2026: AI in Legal Has a Deployment Problem, Legaltech Hub (Mar. 24, 2026), https://www.legaltechnologyhub.com/contents/legalweek-2026-ai-in-legal-has-a-deployment-problem. ↩︎
  11. Olivia Smith Schlinck, Academic Law Librarians Are Paid 47% Less Than Their Faculty Counterparts (Feb. 4, 2022), https://ripslawlibrarian.wordpress.com/2022/02/04/academic-law-librarians-are-paid-47-less-than-their-faculty-counterparts/. ↩︎
  12. Samantha Cole, Watch These Judges Rip Into Lawyers For Citing Cases That Don’t Exist, 404media (June 4, 2026), https://www.404media.co/new-york-court-ai-citations-landberg-case/; J. Koebler, Judge Learns Lawyers on Both Sides of Case Used AI, Cancels Trial, Kicks Everyone Off the Case, 404media (June 9, 2026), https://www.404media.co/judge-learns-lawyers-on-both-sides-of-case-used-ai-cancels-trial-kicks-everyone-off-the-case/. ↩︎
  13. OJ Simpson Murder Trial – Shepard’s clip https://www.youtube.com/watch?v=QFOY0Glg0gU. ↩︎

Effortless Boolean: A Free Tool to Supercharge Your Legal Research

As anyone who has taught legal research knows, Boolean searching is a superpower. The ability to craft a precise query with terms and connectors is the difference between finding a needle in a haystack and finding nothing at all. But for newcomers, the syntax of ( ), !, /p, and /s can feel like learning a new language under pressure.

The Legal Boolean Search Builder is built directly on a process I’ve been teaching for a while now—an 8-step method designed to take the guesswork out of query construction. It moves from identifying key concepts, to brainstorming alternates, and finally to connecting them with the right syntax.

For years, I’ve shared this process in slide decks, but it’s always been static. I wanted to turn it into something dynamic—a tool that could handle the syntax so that researchers could focus on the strategy.

A screenshot of the Legal Boolean Search Builder, as described in the rest of this post, and available at https://booleanbuilder.replit.app/

The Building Process: An Iterative Approach

I built this project using Gemini’s Canvas, and so it may look familiar to Gemini users. It uses HTML, Tailwind CSS for styling, and vanilla JavaScript for all the interactive logic. No complex frameworks, no dependencies—just a single file you can open in any browser. I then threw it into a github repo and imported to Replit so I could host it there.

This came together in a few hours, so I’m sure there are further tweaks and improvements I could make. I’m immensely grateful to Charlie Amiot and Debbie Ginsberg for their sharp insights and invaluable suggestions that took the tool from a basic concept to a polished, user-friendly application.

Finally, this project was significantly influenced by an amazing fillable PDF created by Dan Kimmons and Tara Mospan. Dan described his process for going from worksheet to fillable PDF in these very pages a few years ago.

How It Works: Key Features

The core idea is to break down the complex task of writing a Boolean query into manageable steps.

1. The Two-Column Layout

The user interface is split into two main sections. On the left, you build your concepts step-by-step. On the right, you see your search string come to life in real-time, along with a helpful review checklist. This instant feedback loop is key to the learning process.

2. Smart Suggestions for Phrases

One of the biggest hurdles for new researchers is knowing when to use an exact phrase search (e.g., "assumption of risk") versus a more flexible proximity search. The tool helps by automatically suggesting a proximity search, filtering out common stop words to focus on the core terms.

3. The Truncation Builder

Finding the correct word root for truncation can be tricky. Is it assum! or assump!? To solve this, I added a “Truncation Builder” modal. You can enter all the variations of a word you can think of, and the tool finds the common root, providing you with the most effective truncated term to copy and use.

Try It Yourself

This project was a fantastic experience in turning a teaching methodology into a living tool. The goal was never to replace the critical thinking that goes into legal research, but to remove the syntactic barriers that can get in the way.

You can try the tool out for yourself and view the source code on GitHub. I’d love to hear your feedback!

Benchmarking a Moving Target, or let’s run a hypo through 7 AIs and see what happens

Debbie Ginsberg, Guest Blogger

Benchmarking should be simple, right? Come up with a set of criteria, run some tests, and compare the answers. But how do you benchmark a moving target like generative AI?

Over the past months, I’ve tested a sample legal question in various commercial LLMs (like ChatGPT and Google Gemini) and RAGs (like Lexis Protégé and Westlaw CoCounsel) to compare how each handled the issues raised. Almost every time I created a sample set of model answers to write about, the technology would change drastically within a few days. My set became outdated before I could start my analysis. While this became a good reason to procrastinate, I still wanted to show something for my work.

As we tell our 1Ls, sometimes you need to work with what you have and just write.

The model question

In May, I asked several LLMS and RAGs this question (see the list below for which ones I tested):

Under current U.S. copyright law (caselaw, statutes, regulations, agency information), to what extent are fonts and typefaces protectable as intellectual property? Please focus on the distinction between protection for font software versus typeface designs. What are the key limitations on such protection as established by statute and case law? Specifically, if a font has been created by proprietary software, or if a font has been hand-designed to include artistic elements (e.g, “A” incorporates a detailed drawing of an apple into its design), is the font entitled to copyright protection?

I chose this question because the answer isn’t facially obvious – it straddles the line between “typeface isn’t copyrightable” and “art and software are copyrightable”.  To answer the question effectively, the models would need to address that nuance in some form.

The model benchmarks

The next issue was how to compare the models. In my first runs, the answers varied wildly. It was hard to really compare them. Lately, the answers have been more similar. I was able to develop a set of criteria for comparison. So for the May set, I benchmarked (or at least checked):

  • Did the AI answer the question that I asked?
  • Was the answer thorough (did it more or less match my model answer)?
  • Did the AI cite the most important cases and sources noted in my model answer?
  • Were any additional citations the AI included at least facially relevant?
  • Did the model refrain from providing irrelevant or false information?

I did not benchmark:

  • Speed (we already know the reasoning models can be slow)
  • If the citations were wrong in a non-obvious way 

The model answer and sources

According ot my model answer, the best answers to the question should include at least the following:

  • Font software: Font software that creates fonts is protected by copyright.  The main exception is software that essentially executes a font or font file, meaning the software is utilitarian rather than creative.
  • Typefaces/Fonts: Neither of these is protected by copyright law.  Fonts and typefaces may have artistic elements that are protected by copyright law, but only the artistic elements are protected, not the typefaces or fonts themselves.
  • The answer should include at least some discussion as to whether a heavily artistic font qualifies for protection.

Bonus if the answer addressed:

  • Separability: If the art can be separated from the typeface/font, it’s copyrightable.
  • Alternatives: Can the font/typeface be protected by other IP protections such as licensing, patents, or trademarks?
  • International implications: Would we expect to see the same results in other jurisdictions?

In answering this question, I expected the LLMs and RAGs to cite:

Benchmarking with the AI models

For this post, I ran my model in the following LLMs/RAGs:

  • Lexis Protégé (work account)
  • Westlaw CoCounsel (work account)
  • ChatGPT o3 deep research (work account)
  • Gemini 2.5 deep research (personal paid account)
  • Perplexity research (personal paid account)
  • DeepSeek R1 (personal free account)
  • Claude 3.7 (personal paid account)

I’ve set up accounts in several commercial GenAI products. Some are free, some are Pro, and Harvard pays for my ChatGPT Enterprise account. As an academic librarian, I have access to CoCounsel and Protétgé.

The individual responses are included in the appendix.

I didn’t have access to Vincent or Paxton at the time. I also didn’t have ChatGPT o3 Pro, either. Later in June, Nick Halperin ran my model in Vincent and Paxton, and I ran the model in o3 Pro. Those examples, as well as GPT5, will be included in the appendix but they are not discussed here.

Bechmarking the results

In parsing the results, most answers were fairly similar with some exceptions:

SourceFont software copyrightableTypefaces/
fonts not copyrightable
Exceptions to font‑software copyrightArt in typefaces/fonts copyrightable
Lexis ProtégéYesYesYesNo
Westlaw CoCounselYesYesNoYes
ChatGPT o3 deep researchYesYesYesYes
Gemini 2.5 deep researchYesYesYesYes
Perplexity researchYesYesYesYes
DeepSeek R1YesYesYesYes
Claude 3.7YesYesYesYes
  • Font software is copyrightable: in all answers 
  • Typefaces/fonts are not copyrightable: in all answers
  • Exceptions to font software copyright: in all answers except Westlaw
  • Art in typefaces/fonts is copyrightable: in all answers except Lexis

Several answers included additional helpful information:

SourceSepera-bilityC Office PoliciesAltern-ativesLicen-singInt’lRecentState law
Lexis ProtégéYesNoNoNoNoNoNo
Westlaw Co-CounselNoNoNoNoNoNoYes
ChatGPT o3 deep researchYesYesYesYesYesYesNo
Gemini 2.5 deep researchYesYesYesYesNoNoNo
Per-
plexity research
YesNoYesNoNoNoNo
Deep-
Seek R1
YesNoYesNoNoNoNo
Claude 3.7NoNoYesYesYesNoNo

  • Discussions about separability: Gemini, ChatGPT, Deep Seek (to some extent), Perplexity, Lexis
  • Specific discussions about Copyright Office policies: Gemini, ChatGPT
  • Discussions about alternatives to copyright (e.g., patent, trademark): Gemini, Claude, ChatGPT, Deep Seek, Perplexity
  • Specific discussions about licensing: Gemini, Claude, ChatGPT
  • International considerations: Claude, ChatGPT
  • Recent developments: ChatGPT
  • State law: Westlaw

The models were somewhat consistent about what they cited:

LLM/RAGCopyright statuteCopyright regsAdobeLaatzShake ShackThe Copyright Compendium
Lexis ProtégéYesYesYesYesNoNo
Westlaw Co-
Counsel
YesYesYesYesYesNo
ChatGPT o3 deep researchYesYesYesNoNoYes
Gemini 2.5 deep researchYesYesYesYesNoYes
Perplexity researchNoYesNoNoNoYes
DeepSeek R1YesYesYesNoNoNo
Claude 3.7NoYesYesNoNoNo
  • The Copyright statute: Lexis, Westlaw, Deep Seek, Chat GPT, Gemini
  • Copyright regs: cited by all
  • Adobe: Lexis, Westlaw, Claude, Deep Seek, Chat GPT, Gemini
  • Laatz: Lexis, Westlaw, Gemini
  • Shake Shack: Westlaw
  • The Copyright Compendium: Perplexity, Chat GPT, Gemini; Lexis cited to Nimmer for the same discussion

The models also included additional resources not on my list:

LLM/RAGBlogs etc.Restat.EltraLaw reviewArticles about loansLibGuides
Lexis ProtégéYesYesYesNoNoNo
Westlaw Co-
Counsel
YesNoNoYesYesNo
ChatGPT o3 deep researchYesNoYesNoNoNo
Gemini 2.5 deep researchYesNoYesNoNoYes
Perplexity researchNoNoNoNoNoNo
DeepSeek R1NoNoYesNoNoNo
Claude 3.7YesNoYesNoNoNo
  • Blogs, websites, news articles: The commercial LLMs.  Gemini found the most, but it’s Google.
  • Restatement: Lexis
  • Eltra Corp. v. Ringer, 1976 U.S. Dist. LEXIS 12611: Lexis, Claude, Deep Seek, Chat GPT, Gemini (t’s not a bad case, but not my favorite for this problem)
  • An actual law review article: Westlaw
  • Higher interest rate consumer loans may snag lenders: Westlaw (not sure why)
  • LibGuides: Gemini
  • Included a handy table: ChatGPT, Gemini

The answers varied in depth of discussion and number of sources:

  • Lexis: 1 page of text, 1 page of sources (I didn’t count the sources in the tabs)
  • Westlaw: 2.5 pages of formatted text, 17 pages of sources
  • ChatGPT: 8 pages of well-formatted text, 1 page of sources
  • Gemini: 6.5 pages of well-formatted text, 1 page of sources
  • Perplexity: A little more than 4 pages of text, about 1 page of sources
  • Deep Seek: a little more than 2 pages of weirdly formatted text, no separate sources
  • Claude: 2.5 pages of well-formatted text, no separate sources

Hallucinations

  • I didn’t find any sources that were completely made up
  • I didn’t find any obvious errors in the written text, though some sources made more sense than others
  • I did not thoroughly examine every source in every list (that would require more time than I’ve already devoted to this blog post). 

Some random concluding thoughts about benchmarking

When I was running these searches, I was sometimes frustrated with the Westlaw and Lexis AI research tools. Not only do they fail to describe exactly what they are searching, they also don’t necessarily capture critical primary sources in their answers (we can get a general idea of the sources used, but not as granular as I’d like). For example, the Copyright Compendium includes one of the more relevant discussions about artistic elements in fonts and typefaces, but that discussion isn’t captured in the RAGs.  To be sure, Lexis did find a similar discussion in Nimmer; Westlaw didn’t find anything comparable, although it did cite secondary sources.

In general, the responses provided by all of the generative AI platforms were correct, but some were more complete than others.  For the most part, the commercial reasoning models (particularly ChatGPT and Gemini) provided more detailed and structured answers than the others.  They also provided responses using formatting designed to make the answers easy to read (Westlaw did as well).

None of the models appeared to consider that recency would be a significant factor in this problem.  Several cited a case from the 70s that didn’t concern fonts.  Several failed to cite Laatz, a recent case that’s on point.  Lexis and Westlaw, of course, cited to authoritative secondary sources (and even a law review article in Westlaw’s case).  The LLMs were less concerned with citing to authority.  In all cases, I would have preferred a more curated set of resources than the platforms provided. 

Finally, none of the platforms included visual elements in what is inherently a visual question. It would have been nice to see some examples of “this is probably copyrightable and this is not” (not that I directly asked for them). 

Vibe-Coding Instruction: I Made a Boolean Minigame In 30 Minutes

I’ve been thinking a lot lately about how to bring more interactivity and immediacy into legal research instruction—especially for those topics that never quite “click” the first time. One idea that’s stuck with me is vibe-coding (see Sam Harden’s recent piece on vibecoding for access to justice). The concept, loosely put, is about using code to quickly build lightweight tools that deliver a very specific, helpful experience—often more intuitive than polished, and always focused on solving a narrow, real-world problem.

That framing resonated with me as both an educator and a librarian. In particular, it got me thinking about Boolean searching—an area where students routinely struggle. Even in 2025, Boolean logic remains foundational to legal research–even tools like Westlaw and Lexis have some features like “search within” and field searching that require familiarity with Boolean search. But despite its importance, it can feel abstract and mechanical when taught through static examples or lectures.

So I tried a bit of vibe-coding myself. I built a small, interactive Boolean search game using the Canvas feature in Google Gemini 2.5—it’s a simple web-based activity that gives users a chance to experiment with constructing Boolean expressions and get real-time feedback. It only took about 30 minutes to get a solid version running, and even in that rough form, it worked. The immediate engagement helps clarify the logic in a way that static examples rarely do. You can check it out and play here: https://gemini.google.com/share/436f0db98cef

Screenshot of a "Boolean Search Basics Game" interface. The top section titled "How to Play" explains how to use Boolean search operators:

    AND for documents containing all terms.

    OR for documents containing at least one term.

    NOT to exclude terms.

    Parentheses for grouping.

    Quotes for exact phrases.

    W/N for proximity within N words.

    /P for terms in the same paragraph.

Below the instructions is "Level 1: Using AND", which asks the user to find documents that contain both "apple" and "pie". A text box is provided for entering a Boolean query, with buttons labeled "Run Search" and "Reset Level".

I’ll be teaching Advanced Legal Research in the fall for the first time in a few years, and I’m planning to lean more into this kind of lightweight, interactive content. These micro-tools don’t have to be elaborate to be effective, and they can go a long way toward reinforcing concepts that students often struggle with in more traditional formats.

Have an idea for a micro-tool to use in teaching? They’re easy, fun, and a little addicting to make. You’ll just need access to the paid version of ChatGPT, Claude, or Gemini. (You can also experiment with AI coding assistants like Replit or Bolt.New. Both have limited free versions.) Provide your idea, perhaps some additional context in the form of a file or webpage, and you’re off to the races. My prompt that resulted in a working version of this Boolean game was literally just “Make an interactive game that will help researchers understand the basics of Boolean Search,” and I attached some slides I’ve previously used to teach the topic.

If you build something or you have an idea I’d love to hear about it!

Revolutionizing Legal Education with AI: The Socratic Quizbot

I had the pleasure of co-teaching AI and the Practice of Law with Kenton Brice last semester at OU Law. It was an incredible experience. When we met to think through how we would teach this course, we agreed on one crucial component:
We wanted the students to get a lot of reps using AI throughout the entire course.

That is fairly easy to accomplish for things like research, drafting, and general studying for the course but we hit a roadblock with the assessment component. I thought about it for a week and said, “Kenton, what if we created an AI that would Socratically quiz the students on the readings each week?” His response was, “Do you think you can do that?” I said, “I don’t know but I’ll give it a try.” 🤷‍♂️

Thus Socratic Quizbot was born. If you follow me on social media, you’ve probably seen me soliciting feedback on the paper:

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4975804

December 2024 Update

Purpose

A lot of the motivation for Quizbot was a new paradigm in the law school ecosystem: the take-home essay is effectively dead. In fact, lots of the typical homework that you would assign as a law school professor simply breaks when you introduce something like ChatGPT or Claude into our world. We needed to come up with new methods of assessment.

I knew that these tools were really good at ingesting documents like PDFs and then summarizing them (manipulating the text, generating based on the text, etc.). What I needed was an AI that could read our course readings and then have a back-and-forth Socratic conversation with the students about those readings, and then some method to assess those conversations so that I could give students a grade. This felt like a big task with many potential pitfalls for one guy who is only mediocre (at best) at coding and app development.

As it turned out, I was able to fumble my way through the process and create a method of assessment that students seemed to enjoy. Alright, “enjoy” is probably too strong of a word, but they tolerated it and said they liked it quite a bit more than something like a multiple-choice test or a take-home essay. The Socratic Quizbot enables you to scale cold-calling to every student in the class while eliminating much of the stress and embarrassment that law students have dreaded since time immemorial.

Since many of the people who are interested in this blog post may have already read or skimmed my article, I decided to add my update as Appendix A so that you could simply fast-forward to the portion that is relevant to you. There is also a link to the open-source code in Github.

A Brief Overview of What is in Appendix A

Appendix A was born out of one question I kept getting after sharing the pre-print of my article: “How?” So, let me show you exactly how you can implement the Socratic Quizbot in your classroom, along with some insights from my students who graciously let me experiment with them.

Student Feedback, Challenges, and Improvements
Students overwhelmingly preferred this method to essays or multiple-choice quizzes, citing the flexibility to ask for clarification and control the pace of their learning. It also reduced the fear of being judged by their peers. That said, a few students tried to game the system by flipping the questions back on the bot. My grading rubric handled that, but I’d like to make Quizbot more persistent in pressing them for answers next time. I’m also excited to explore gamification—adding themes, Easter eggs, or playful interactions to make the experience even more enjoyable.

Two Ways to Get Started
If you want to try this yourself, you’ve got two paths. The no-code approach uses ChatGPT Teams and involves setting up a CustomGPT that ingests your course readings and quizzes your students. This is great if you’re looking for quick implementation. For the more tech-savvy or budget-conscious, the code-based option lets you host Quizbot locally using the instructions I’ve shared on GitHub. It takes a bit more effort but gives you total control over security and functionality. Hopefully you will see a version of Socratic Quizbot available in CALI.org in the future because I have been talking with John and Elmer and they both seem interested with integrating it into the platform (although, do not hold them to that because it’s still very-early talks).

Ultimately, my goal is to make this tool accessible for anyone in legal education. Whether you’re a tech whiz or new to AI, there’s a way to incorporate this into your classroom. And if you’re as curious about alternative assessments as I am, I’d love to hear your thoughts and ideas! The benefit of making it open in Github is that you can fork and improve my prototype. I would be deeply honored to see improvements on my little project and love to see what our community can do to improve it.

Evaluating Generative AI for Legal Research: A Benchmarking Project

This is a post from multiple authors: Rebecca Fordon (The Ohio State University), Deborah Ginsberg (Harvard Law Library), Sean Harrington (University of Oklahoma), and Christine Park (Harvard Law Library)

In late 2023, several legal research databases and start-up competitors announced their versions of ChatGPT-like products, each professing that theirs would be the latest and greatest. Since then, law librarians have evaluated and tested these products ad hoc, offering meaningful anecdotal evidence of their experience, much of which can be found on this blog and others. However, one-time evaluations can be time-consuming and inconsistent across the board. Certain tools might work better for particular tasks or subject matters than others, and coming up with different test questions and tasks takes time that many librarians might not have in their daily schedules.

It is difficult to test Large-Language Models (LLMs) without back-end access to run evaluations. So to test the abilities of these products, librarians can use prompt engineering to figure out how to get desired results (controlling statutes, key cases, drafts of a memo, etc.). Some models are more successful than others at achieving specific results. However, as these models update and change, evaluations of their efficacy can change as well. Therefore, we plan to propose a typology of legal research tasks based on existing computer and information science scholarship and draft corresponding questions using the typology, with rubrics others can use to score the tools they use.

Although we ultimately plan to develop this project into an academic paper, we share here to solicit thoughts about our approach and connect with librarians who may have research problem samples to share.

Difficulty of Evaluating LLMs

Let’s break down some of the tough challenges with evaluating LLMs, particularly when it comes to their use in the legal field. First off, there’s this overarching issue of transparency—or rather, the lack thereof. We often hear about the “black box” nature of these models: you toss in your data, and a result pops out, but what happens in between remains a mystery. Open-source models allow us to leverage tools to quantify things like retrieval accuracy, text generation precision, and semantic similarity. We are unlikely to get the back-end access we need to perform these evaluations. Even if we did, the layers of advanced prompting and the combination of tools employed by vendors behind the scenes could render these evaluations essentially useless.

Even considering only the underlying models (e.g., GPT4 vs Claude), there is no standardized method to evaluate the performance of LLMs across different platforms, leading to inconsistencies. Many different leaderboards evaluate the performance of LLMs in various ways (frequently based on specific subtasks). This is kind of like trying to grade essays from unrelated classes without a rubric—what’s top-notch in one context might not cut it in another. As these technologies evolve, keeping our benchmarks up-to-date and relevant is becoming an ongoing challenge, and without uniform standards, comparing one LLM’s performance to another can feel like comparing apples to oranges.

Then there’s the psychological angle—our human biases. Paul Callister’s work sheds light on this by discussing how cognitive biases can lead us to over-rely on AI, sometimes without questioning its efficacy for our specific needs. Combine this with the output-based evaluation approach, and we’re setting ourselves up for potentially frustrating misunderstandings and errors. The bottom line is that we need some sort of framework for the average user to assess the output.

One note on methods of evaluation: just before publishing this blog post, we learned of a new study from a group of researchers at Stanford, testing the claims of legal research vendors that their retrieval-augmented generation (RAG) products are “hallucination-free.” The group created a benchmarking dataset of 202 queries, many of which were chosen for their likelihood of producing hallucinations. (For example, jurisdiction/time-specific and treatment questions were vulnerable to RAG-induced hallucinations, whereas false premise and factual recall questions were known to induce hallucinations in LLMs without RAG.) The researchers also proposed a unique way of scoring responses to measure hallucinations, as well as a typology of hallucinations. While this is an important advance in the field and provides a way to continue to test for hallucinations in legal research products, we believe hallucinations are not the only weakness in such tools. Our work aims to focus on the concrete applications of these LLMs and probe into the unique weaknesses and strengths of these tools. 

The Current State of Prompt Engineering

Since the major AI products were released without a manual, we’ve all had to figure out how to use these tools from scratch. The best tool we have so far is prompt engineering. Over time, users have refined various templates to better organize questions and leverage some of the more surprising ways that AI works.

As it turns out, many of the prompt templates, tips, and tricks we use with the general commercial LLMs don’t carry over well into the legal AI sphere, at least with the commercial databases we have access to. For example, because the legal AIs we’ve tested so far won’t ask you questions, researchers may not be able to have extensive conversations with the AI (or any conversation for some of them). So that means we must devise new types of prompts that will work in the legal AI sphere, and possibly work only in the AI sphere.

We should be able to easily design effective prompts because the data set the AIs use is limited. But it’s not always clear exactly what sources the AI is using. Some databases may list how many cases they have for a certain court by year; others may say “selected cases before 1980” without explaining how they were selected. And even when the databases provide coverage, it may not be clear exactly which of those materials the AI can access.

We still need to determine what prompt templates will be most effective across legal databases. More testing is needed. However, we are limited to the specific databases we can access. While most (all?) academic law librarians have access to Lexis+ AI, Westlaw has yet to release its research product to academics. 

Developing a Task Typology

Many of us may have the intuition that there are some legal research tasks for which generative AI tools are more helpful than others. For example, we may find that generative AI is great for getting a working sense of a topic, but not as great for synthesizing a rule from multiple sources. But if we wanted to test that intuition and measure how well AI performed on different tasks, we would need to first define those tasks. This is similar, by the way, to how the LegalBench project approached benchmarking legal analysis—they atomized the IRAC process for legal analysis down to component tasks that they could then measure.

After looking at the legal research literature (in particular Paul Callister’s “problem typing” schemata and AALL’s Principles and Standards for Legal Research Competency), we are beginning to assemble a list of tasks for which legal researchers might use generative AI. We will then group these tasks according to where they fall in an information retrieval schemata for search, following Marchionini (2006) & White (2024), into Find tasks (which require a simple lookup), Learn & Investigate tasks (which require sifting through results, determining relevance, and following threads), and Create, Synthesize, and Summarize tasks (a new type of task for which generative AI is well-suited).

Notably, a single legal research project may contain multiple tasks. Here are a few sample projects applying a preliminary typology:

Again, we may have an initial intuition that generative AI legal research platforms, as they exist today, are not particularly helpful for some of these subtasks. For example, Lexis+AI currently cannot retrieve (let alone analyze) all citing references to a particular case. Nor could we necessarily be certain from, say, CoCounsel’s output, that it contained all cases on point. Part of the problem is that we cannot tell which tasks the platforms are performing, or the data that they have included or excluded in generating their responses. By breaking down problems into their component tasks, and assessing competency on both the whole problem and the tasks, we hope to test our intuitions.

Future Research

We plan on continually testing these LLMs using the framework we develop to identify which tasks are suitable for AIs and which are not. Additionally, we will draft questions and provide rubrics for others to use, so that they can grade AI tools. We believe that other legal AI users will find value in this framework and rubric. 

Leapfrogging the Competition: Claude 3 Researches and Writes Memos (Better Than Some Law Students and Maybe Even Some Lawyers?)

Introduction

I’ve been incredibly excited about the premium version of Claude 3 since its release on March 4, 2024, and for good reason. Now that my previous favorite chatty chatbot, ChatGPT-4, has gone off the rails, I was missing a competent chatbot… I signed up the second I heard on March 4th, and it has been a pleasure to use Claude 3 ever since. It actually understands my prompts and usually provides me with impressive answers. Anthropic, maker of the Claude chatty chatbot family, has been touting Claude’s accomplishments of supposedly beating its competitors on common chatbot benchmarks, and commentators on the Internet have been singing its praises. Just last week, I was so impressed by its ability to analyze information in news stories in uploaded files that I wrote a LinkedIn post also singing its praises!

Hesitation After Previous Struggles

Despite my high hopes for its legal research abilities after experimenting with it last week, I was hesitant to test Claude 3. I have a rule about intentionally irritating myself—if I’m not already irritated, I don’t go looking for irritation… Over the past several weeks, I’ve wasted countless hours trying to improve the legal research capabilities of ChatGPT-3.5, ChatGPT-4, Microsoft Copilot, and my legal research/memo writing GPTs through the magic of (IMHO) clever prompting and repetition. Sadly, I failed miserably and concluded that either ChatGPT-4 was suffering from some form of robotic dementia, or I am. The process was a frustrating waste, and I knew that Claude 3 doing a bad job of legal research too could send me over the edge….

Claude 3’s Wrote a Pretty Good Legal Memorandum!

Luckily for me, when I finally got up the nerve to test out the abilities of Claude 3, I found that the internet hype was not overstated. Somehow, Claude 3 has suddenly leapfrogged over its competitors in legal research/legal analysis/legal memo writing ability – it instantly did what would have taken a skilled researcher over an hour and produced a better legal memorandum which is probably better than that produced by many law students and even some lawyers. Check it out for yourself! Unless this link actually works for any Claude 3 subscribers out there, there doesn’t seem to be a way to actually link to a Claude 3 chat at this time. However, click here for the whole chat I cut and pasted into a Google Drive document, here for a very long screenshot image of the chat, or here for the final 1,446-word version of the memo as a Word document.

Comparing Claude 3 with Other Systems

Back to my story… The students’ research assignment for the last class was to think of some prompts and compare the results of ChatGPT-3.5, Lexis+ AI, Microsoft Copilot, and a system of their choice. Claude 3 did not exist at the time, but I told them not to try the free Claude product because I had canceled my $20.00 subscription to the Claude 2 product in January 2024 due to its inability to provide useful answers – all it would say was that it was unethical to answer every question and tell me to do it myself. When creating an answer sheet before class tomorrow which compares the same set of prompts on different systems, I decided to omit Lexis+ AI (because I find it useless) and to include my new fav Claude 3 in my comparison spreadsheet. Check it out to compare for yourself!

For the research part of the assignment, all systems were given a fact pattern and asked to “Please analyze this issue and then list and summarize the relevant Texas statutes and cases on the issue.” While the other systems either made up cases or produced just two or three actual real and correctly cited cases on the research topic, Claude 3 stood out by generating 7 real, relevant cases with correct citations in response to the legal research question. (And, it cited to 12 cases in the final version of its memo.)

It did a really good job of analysis too!

Generating a Legal Memorandum

Writing a memo was not part of the class assignment because the ChatGPT family was refusing the last few weeks,* and Bing Copilot had to be tricked into writing one as part of a short story, but after seeing Claude 3’s research/analysis results, I decided to just see what happened. I have many elaborate prompts for ChatGPT-4 and my legal memorandum GPTs, but I recalled reading that Claude 3 worked well with zero-shot prompting and didn’t require much explanation to produce good results. So, I decided to keep my prompt simple – “Please generate a draft of a 1500 word memorandum of law about whether Snurpa is likely to prevail in a suit for false imprisonment against Mallatexaspurses. Please put your citations in Bluebook citation format.”

From my experience last week with Claude 3 (and prior experience with Claude 2 which would actually answer questions), I knew the system wouldn’t give me as long an answer as requested. The first attempt yielded a pretty high-quality 735-word draft memo that cited all real cases with the correct citations*** and applied the law to the facts in a well-organized Discussion section. I asked it to expand the memo two more times, and it finally produced a 1,446-word document. Here is part of the Discussion section…

Implications for My Teaching

I’m thrilled about this great leap forward in legal research and writing, and I’m excited to share this information with my legal research students tomorrow in our last meeting of the semester. This is particularly important because I did such a poor job illustrating how these systems could be helpful for legal research when all the compared systems were producing inadequate results.

However, with my administrative law legal research class starting tomorrow, I’m not sure how this will affect my teaching going forward. I had my video presentation ready for tomorrow, but now I have to change it! Moreover, if Claude 3 can suddenly do such a good job analyzing a fact pattern, performing legal research, and applying the law to the facts, how does this affect what I am going to teach them this semester?

*Weirdly, the ChatGPT family, perhaps spurred on by competition from Claude 3, agreed to attempt to generate memos today, which it hasn’t done in weeks…

Note: Claude 2 could at one time produce an okay draft of a legal memo if you uploaded the cases for it, that was months ago (Claude 2 link if it works for premium subscribers and Google Drive link of cut and pasted chat). Requests in January resulted in lectures about ethics which resulted in the above-mentioned cancellation.

Beyond Legal Documentation: Other Business Uses of Generative AI

I have been listening to and enjoyed thinking about and participating in conversations about how generative AI is going to be integrated into the practice of law. Most of these conversations surround how it will be integrated into legal documents, which is not surprising considering how many lawyers have gotten in trouble for this and how quickly our research and writing products are integrating the technology. But there is more to legal practice than creating client and/or court documents. In fact, there are many more business uses of generative AI than just research and drafting.

This past fall, I was asked to lead an AI session for Capital University’s joint venture with the Columbus College of Art & Design, the Institute for Creative Leadership at Work. I was asked to adapt my presentation to HR professionals and focus on SHRM compliance principles. I enjoyed the deep dive into this world, and I came away from my research with a lot of great ideas for my session, Bard, Bing, and ChaptGPT, Oh My!: Possible Ethical Uses of Generative AI at Work, such as tabletop emergency exercises, social media posts, job descriptions, and similar tasks.

This week, I have been thinking about how everyone’s focus has really been around legal documentation, my own included. But there are an amazing number of backend business tasks that could also utilize AI in a positive way. The rest of the world, including HR, has been focusing on them for a while, but we seem to have lost track of these business tasks.

Here are some other business uses of generative AI and prompts that I think hold great promise. Continue reading →

Tabletop emergency simpulation image
  1. Drafting job descriptions
    • Pretend that you are an HR specialist for a small law firm in the United States. Draft a job description for a legal secretary who focuses on residential real estate transactions but may assist with other transactional legal matters as needed. [Include other pertinent details of the position]. The job description will be posted in the following locations [fill in list]
  2. Creating tabletop simulations to work through crisis/emergency plans:
    • You are an HR specialist who is helping plan for and test the company’s responses to a variety of situations. First is an active shooter in the main building. A 5th grade tour of the facilities is going on on the third floor. Create a detailed tabletop simulation to test this.
    • Second scenario: The accounting department is celebrating the birthday of the administrative assistant and is having cake in the breakroom. The weather has turned bad, and an F4 tornado is spotted half a mile away. After 15 minutes, the tornado strikes the building directly. Create a detailed tabletop simulation to test the plan and response for this event.
  3. Assisting with lists of mandatory and voluntary employee trainings
    • Pretend that you are an HR professional who works for a law firm. You are revamping the employee training program. We need to create a list of mandatory trainings and a second list of voluntary trainings. Please draft a list of training appropriate to employees in a law firm setting.
  4. Assisting with social media posting creation:
    • Pretend that you are a professional social media influencer for the legal field. Draft an Instagram post, including creating a related image, to celebrate Law Day, which is coming up on May 1st.  Make sure that it is concise and Instagram appropriate. Please include hashtags.
  5. Assisting with creating employee policies or handbooks (verify content!):
    • Pretend that you are an information security professional. Draft an initial policy for a law firm regarding employee AI usage for company work. The company wants to allow limited use of generative AI. They are very worried that their proprietary and/or confidential client data will be accidentally released. Specify that only your custom AI system – [name firm-specific or specialized AI with a strong privacy contract clause] – can be used with company data. The policy must also take into consideration the weaknesses of all AI systems, including hallucinations, potential bias, and security issues.
  6. Assisting with making sure your web presence is ADA accessible:
    • Copilot/web-enabled Prompt: Pretend that you are a graphic designer who has been tasked with making sure that a law firm’s online presence is ADA accessible. Please review the site [insert link], run an ADA compliance audit, and provide an accessibility report, including suggestions on what can be done to fix any accessibility issues that arise.
  7. Onboarding documentation
    • Create a welcome message for a new employee. Tell them that the benefits orientation will be at 9 am in the HR conference room on the next first Tuesday of the month. Pay day is on the 15th and last day of each month, unless payday falls on a weekend or federal holiday, in which case it will be the Friday before. Employees should sign up for the mandatory training that will be sent to them in an email from IT.
    • (One I just user IRL) Pretend that you are a HR specialist in a law library. A new employee is starting in 6 weeks, and the office needs to be prepared for her arrival. [Give specific title and any specialized job duties, including staff supervision.] Create an onboarding checklist of important tasks, such as securing keys and a parking permit, asking IT to set up their computer, email address, and telephone, asking the librarians to create passwords for the ILS, Libguides, and similar systems, etc.

What other tasks (and prompts) can you think of that might be helpful? If you are struggling to put together a prompt, please see my general AI Prompt Worksheet in Introducing AI Prompt Worksheets for the Legal Profession. We welcome you to share your ideas in the comments.

Birth of the Summarizer Pro GPT: Please Work for Me, GPT

Last week, my plan was to publish a blog post about creating a GPT goofily self-named Summarizer Pro to summarize articles and organize citation information in a specific format for inclusion in a LibGuide. However, upon revisiting the task this week, I find myself first compelled to discuss the recent and thrilling advancements surrounding GPTs – the ability to incorporate GPTs into a ChatGPT conversation.

What is a GPT?

But, first of all, what is a GPT? The OpenAI website explains that GPTs are specialized versions of ChatGPT designed for customized applications. These unique GPTs enable anyone to modify ChatGPT for enhanced utility in everyday activities, specific tasks, professional environments, or personal use, with the added ability to share these personalized versions with others.

To create or use a GPT, you need access to ChatGPT’s advanced features, which require a paid subscription. Building your own customized GPT does not require programming skills. The process involves starting a chat, giving instructions and additional information, choosing capabilities like web searching, image generation, or data analysis, and iteratively testing and improving the GPT. Below are some popular examples that ChatGPT users have created and shared in the ChatGPT store:

GPT Mentions

This was already exciting, but last week they introduced a feature that takes it to the next level – users can now invoke a specialized GPT within a ChatGPT conversation. This is being referred to as “GPT mentions” online. By typing the “@” symbol, you can choose from GPTs you’ve used previously for specific tasks. Unfortunately, this feature hasn’t rolled out to me yet, so I haven’t had the chance to experiment with it, but it seems incredibly useful. You can chat with ChatGPT as normal while also leveraging customized GPTs tailored to particular needs. For example, with the popular bots listed above, you could ask ChatGPT to summon Consensus to compile articles on a topic. Then call on Write For Me to draft a blog post based on those articles. Finally, invoke Image Generator to create a visual for the post. This takes the versatility of ChatGPT to the next level by integrating specialized GPTs on the fly.

Back to My GPT Summarizer Pro

Returning to my original subject, which is employing a GPT to summarize articles for my LibGuide titled ChatGPT and Bing Chat Generative AI Legal Research Guide. This guide features links to articles along with summaries on various topics related to generative AI and legal practice. Traditionally, I have used ChatGPT (or occasionally Bing or Claude 2, depending on how I feel) to summarize these articles for me. It usually performs admirably well on the summary part, but I’m left to manually insert the title, publication, author, date, and URL according to a specific layout. I’ve previously asked normal old ChatGPT to organize the information in this format, but the results have been inconsistent. So, I decided to create my own GPT tailored for this task, despite having encountered mixed outcomes with my previous GPT efforts.

Creating GPTs is generally a simple process, though it often involves a bit of fine-tuning to get everything working just right. The process kicks off with a set of questions… I outlined my goals for the GPT – I needed the answers in a specific format, including the title, URL, publication name, author’s name, date, and a 150-word summary, all separated by commas. Typically, crafting a GPT involves some back-and-forth with the system. This was exactly my experience. However, even after this iterative process, the GPT wasn’t performing exactly as I had hoped. So, I decided to take matters into my own hands and tweak the instructions myself. That made all the difference, and suddenly, it began (usually) producing the information in the exact format I was looking for.

Summarizer Pro in Action!

Here is an example of Summarizer Pro in action! I pasted a link to an article into the text box and it produced the information in the desired format. However, reflecting the dynamic nature of ChatGPT responses, the summaries generated this time were shorter compared to last week. Attempts to coax it into generating a longer or more detailed summary were futile… Oh well, perhaps they’ll be longer if I try again tomorrow or next week.

Although it might not be the most fancy or thrilling use of a GPT, it’s undeniably practical and saves me time on a task I periodically undertake at work. Or course, there’s no shortage of less productive, albeit entertaining, GPT applications, like my Ask Sarah About Legal Information project. For this, I transformed around 30 of my blog posts into a GPT that responds to questions in the approximate manner of Sarah.

Introducing AI Prompt Worksheets for the Legal Profession

I spent the first week of January attending the American Association of Law Schools’ Annual Meeting in Washington D.C. I was really impressed with all of the thoughtful AI sessions, including two at which I participated as a panelist. The rooms were packed beyond capacity for each AI session that I attended, which underscored the growing interest in AI in the legal academy. Many people attended in order to start their education. The overwhelming interest at the conference made my decision clear: it is time to launch my AI prompt worksheets to the world, addressing the need I observed there. While AALS convinced me to release the worksheets, the worksheets themselves were created for an upcoming presentation at ABA TECHSHOW 2024, How to Actually Use AI in Your Legal Practice, at which Greg Siskind and I will be discussing practical tips for generative AI usage.

DALL-E generated

Background: Good Habits – Research Planning

Law librarians have been encouraging law students to create a research plan before they start their research for decades. The plan form varies by school and/or librarian, but it usually requires the researcher to answer questions on the following topics:

  • Issue Identification
  • Jurisdiction
  • Facts
  • Key words/Terms of Art
  • Resource Selection

Once the questions are answered, the plan has the researcher write out some test searches. The plan evolves as the research progresses. The more experienced the researcher, the less formal the plan often is, but even the most experienced researcher retrieves better results if they pause to consider what they know currently and what they need in the results. After all, garbage in, garbage out (GIGO). In other words, the quality of our input affects the quality of the output. This is especially true when billable hours come into play, and you cannot bill for excess time due to poor research skills.

Continuing the Good Habits with Generative AI

GIGO applies just as much to generative AI. I quickly noticed that my AI results are much better when I stop and think them through, providing a high level of detail and a good explanation of what I want the AI system to produce. So, good law librarian that I am, I created a new form of plan for those who are learning to draft a prompt. Thus, I give you my AI prompt worksheets.

AI Prompt Worksheet – General

Worksheet (Word)

The first worksheet that I created is geared towards general generative AI systems like ChatGPT, Claude 2, Bing Chat/Copilot, and Bard.  The worksheet makes the prompter think through the following topics:

  • Tone of Output
  • Role
  • Output Format
  • Purpose
  • Issue
  • Potential Refinements (may be added later as the plan evolves)

So that you can easily keep track of your prompts, the Worksheet also requests some metadata about your prompt, including project name, date, and AI system used. The final question lets the prompter decide if this prompt worked for them.

DALL-E generated

AI Prompt Worksheet – Legal

Worksheet – Legal (Word)

For the second worksheet, I wanted to draft something that works well with legal AI systems. Based on the systems that I have received access to, such as Lexis AI and LawDroid Copilot, and the systems that I have seen demonstrated, I cut down some of the fields. Most of the systems are building a guided AI prompting experience, so they will ask you for the jurisdiction, for instance. They may also allow you to select a specific type of output, such as a legal memo or contract clause. This means less need for an extensive number of fields in the worksheet. In fact, when I ran the worksheet past a vLex representative, I was told it was not needed at all because they had made the guided prompt that easy.

Librarian that I am, however, I still feel that planning before you prompt is preferred. Reasons for this preference include: the high cost of the current generative AI searches, the desire for efficient and effective results, knowledge that an attorney’s time is literally worth money, and the desire for a happy partner and client.

The legal worksheet trims the fields down to role, output (format and jurisdiction), issue, and refinement instructions. This provides enough room to flesh out your prompt without overlapping the guided prompt fields too much.

General Comments Regarding the Worksheets

With both worksheets, the key is to give a good, detailed description of what you need. Think about it like explaining what you need to a first-year law student – the more detail you give, the more likely you are to get something useable. The worksheets provide examples of the level of detail recommended, and you will find links to the results in the footnotes of the forms.

In addition to helping perfect your prompt with some pre-planning, these worksheets should be useful for creating your very own prompt library.

Feedback Wanted!

DALL-E created

Please feel free to use the worksheets (just don’t sell them or otherwise profit off of them! Ask if you want to make a derivative of them). If you do use them, please let me know what you think in the comments or via email. How have they assisted (or not) with improving your prompting skills? Are there fields you would like to see added/removed?  I will be updating and releasing new versions as I go. If you are looking for the most recent versions of the worksheets, I will post them at: https://law-capital.libguides.com/Jennys_AI_Resources/AI_Prompt_Worksheets