AI in Legal Review

Prompt Engineering for Legal Contract Analysis

Structured prompts prevent AI hallucinations in legal analysis.

Cover illustration for “Prompt Engineering for Legal Contract Analysis”
Cover illustration for “Prompt Engineering for Legal Contract Analysis”

A Northern District of Alabama court disqualified attorneys in July 2025 after five ChatGPT-hallucinated citations appeared across two filings. The firm had a written AI governance policy on file, and the sanctions came anyway, because the policy addressed conduct after the fact rather than the structure of the prompts that produced the problem. Bar regulators in each state where the sanctioned attorneys were licensed were notified. The case illustrates the argument this piece makes: accuracy and auditability in legal contract analysis come from how a prompt is built.

The failure that produced the Alabama sanctions is common and follows a predictable pattern. A lawyer treats a prompt as a question put to a knowledgeable colleague, phrased the way one might ask a question out loud, and the model responds the way a generalist would: fluently, confidently, and without any particular attachment to the jurisdiction, the client's position, or the document actually in front of the person asking. General-purpose AI finds clauses reliably. It fails on precise language, numeric thresholds, multi-part requirements, cross-references, and absence checks, which happen to be exactly the features of a contract that generate legal liability when they're misread or missed. Spellbook's 2026 practitioner guide puts the distinction plainly: "the difference between a generic prompt and a well-structured one can mean the difference between a rough starting point and usable output."

The comparison between tools comes later in this piece, since model selection affects the outcome far less than the structure of the prompt does. For now, the point is narrower: no model, however capable, compensates for a prompt that fails to specify who is asking, under what law, with what document in hand, in what format, and with what rules for handling uncertainty. The Alabama court found no fault in how the firm's AI policy was written. It found that nothing in the firm's process caught five fabricated citations before they reached a filing, because no part of the architecture was built to catch them.

The Legal Prompts 2026 guide names the fix as a structural anatomy rather than a style tip: the five-component RCTFC framework, standing for Role, Context, Task, Format, and Constraints, describes what a reliable legal prompt actually contains. This piece organizes its argument around five structural decisions drawn from that same logic: assigning a role, specifying jurisdiction, supplying full document context, defining an output schema, and building in anti-hallucination rules. Each is a decision a lawyer makes before the model generates a single word, and each closes off a specific way the output can go wrong.

Diagram: The Five-Component Legal Prompt: RCTFC. Visualizes: Visualize the five structural decisions that make a legal prompt reliable, drawn directly from the RCTFC framework named in the article: Role, Context, Task, Format, and Constraints.

The first structural decision in any legal prompt is deciding who the model is being asked to be. Large language models predict the most likely next word given everything that came before, so when you give it a role, that prediction narrows toward specialist vocabulary, specialist depth, and specialist caution before the model even sees the task. HAQQ's 2026 legal prompting guide describes the mechanism directly: "Telling the AI to act as a specific type of legal professional narrows the scope of its response and improves relevance, instead of a generic answer, you get analysis from the perspective of a specialist."

The Legal Prompts guide shows how much daylight separates a role-free prompt from a role-anchored one. "Write a contract clause about non-compete" produces whatever the model's training data suggests a non-compete clause typically looks like, with no attachment to any particular legal standard. "You are a California-licensed transactional attorney. Draft a non-compete clause that complies with California Business and Professions Code Section 16600" produces something else entirely, because the role instruction grounds the task in a named statute and a specific professional standard before the drafting even starts.

Role assignment carries a second layer that matters just as much in contract work: whose side the model is reasoning from. Spellbook's guide treats stating who the lawyer represents as its own step, apart from naming the professional role. Adding a line like "I represent [Party]" shifts the entire analysis from neutral description to advocacy, so the model flags risk and leverage from the client's side of the table rather than summarizing the agreement as a disinterested observer would. A prompt without that layer tends to produce a balanced summary of the contract's terms, which has its uses in academic review but does little for a lawyer trying to negotiate better terms or allocate risk away from a client. The role, in other words, is not a flourish added for tone. It sets the analytical frame the model applies to everything that follows, and everything that follows, including jurisdiction, depends on getting that frame right first.

Specifying jurisdiction and governing law as a constraint, not background color

Jurisdiction is the second structural decision, and it functions as an active constraint on how the model reasons. If a prompt leaves out governing law, you don't get a neutral or universal answer. It produces an answer that sounds legally fluent while applying rules that may not govern the contract at all, which is a more dangerous outcome than producing no answer whatsoever. The Legal Prompts 2026 guide states the stakes without hedging: "An overlooked jurisdiction-specific rule can constitute malpractice. Generic AI output that ignores the governing law of a specific state is worse than useless, it is dangerous."

No legal AI vendor invented that standard as a marketing claim. ABA Model Rules 1.1, on competence, and 8.4, on misconduct, now cover AI use by attorneys through ABA Formal Opinion 512, issued in 2024, which maps the existing rule text onto AI use without requiring any amendment to the rules themselves. Jurisdiction-awareness is now part of competent representation, and you can enforce it under rules that predate large language models by decades.

Jurisdiction also doesn't operate as an isolated fact bolted onto a prompt. It constrains how the role, the task, and the output format all get interpreted. A force majeure clause redrafted under English law drafting conventions needs to come out differently than the same clause redrafted under New York law, because the drafting conventions, the interpretive doctrines, and the default rules that fill gaps in the clause all differ. Juro's 2026 prompt guide shows jurisdiction embedded directly inside the task instruction rather than appended afterward: "Redraft this force majeure clause... compliant with English law drafting conventions." Cross-border work raises the stakes further. HAQQ's guide offers a pattern built for exactly this situation: "You are reviewing a cross-border supply agreement between a US manufacturer and an EU distributor. The agreement is governed by German law." Both the factual setting and the governing law appear before any task instruction, because a model asked to review that agreement without knowing which law governs it has no basis for judging whether any given clause is standard, enforceable, or even legal.

The August 2026 New York City Bar Association policy paper lands on the same conclusion from the regulatory side. The Association's Emerging Companies & Venture Capital Committee found that AI tools can help with legal work, but they can't substitute for professional legal judgment under existing rules of professional conduct. Jurisdiction-specific judgment is precisely the thing an unprompted model has no way to supply on its own, which makes the lawyer's prompt the only place that judgment can enter the process before the model starts reasoning.

Providing full document context to prevent the model from reasoning in a vacuum

Once a model knows who it is and what law governs the matter, the next structural question is what it can actually see. A model asked to analyze a single clause in isolation has no way to detect cross-references to other sections, no way to catch internal inconsistencies between provisions, and no way to flag obligations that are implicit in the structure of the agreement. These are not edge cases that only arise in unusually complex deals. Cross-references and implicit obligations are the structural features of a contract that litigators exploit most often and that courts spend the most time interpreting, which makes them exactly the features a thin-context prompt is most likely to miss.

General-purpose AI tends to catch missing standard clauses, non-standard language deviations, and surface-level inconsistencies without much trouble. It struggles with implicit obligations, with requirements that depend on cross-referenced provisions elsewhere in the document, and with novel legal constructions that fall outside the patterns in its training data. HAQQ's guide frames the fix as ambiguity elimination: "Include the type of case, the jurisdiction, the parties involved, the relevant legal framework, and any specific constraints. The more context you provide, the less the AI has to guess." Juro's guide recommends you build a standing contract playbook or contract management policy, so you don't have to re-type the baseline context into every prompt, and the model stays grounded in the same reference points across an entire matter.

Context works best layered. Juro recommends an iterative approach for legal work: start broad, narrow the request with continuous feedback, and build context cumulatively rather than trying to hand the model every fact at once. HAQQ's guide reaches the same conclusion from a different angle, noting that breaking a complex task into sequential steps produces meaningfully better results than asking for everything in a single pass.

Retrieval-Augmented Generation is the production-scale version of this same principle. Rather than a lawyer manually assembling context for each prompt, this kind of retrieval system pulls the relevant documents at the moment of the query and grounds the model's response in current, real data rather than whatever pattern its training happened to absorb. In practice, this means every output should point to a specific paragraph in a specific document, with verified citations linking back to primary law or to the firm's own clause library. RAG doesn't replace the discipline of providing context. It automates what a well-built context section in a prompt already does by hand. None of it matters, though, if the output that comes back has no defined shape, because even a model with perfect context and perfect jurisdiction-awareness produces something unusable if there's no schema for what the answer should look like.

The comparison between tools comes later in this piece, since model selection affects the outcome far less than the structure of the prompt does.

Defining a structured output schema so the analysis is auditable, not just readable

The fourth structural decision is what shape the output takes, and it is the decision that turns AI-generated analysis into something that can actually be checked. If you don't define a schema, even accurate output resists systematic review, and the human review gate that legal ethics rules require becomes impractical once volume rises past a handful of documents. A narrative paragraph describing a contract's risks might read well, but it offers no fixed location for any given finding, no way to compare it against a playbook standard, and no clean path back to the specific clause it describes.

A three-column table, listing clause reference, issue identified, and recommended fix, solves all three problems at once: every finding has an address in the document, every finding can be checked against a standard, and every finding traces back to source text. Juro's guide identifies exactly this structure as the mechanism that makes AI output "instantly actionable," specifying "a three-column table: clause reference, issue identified, recommended fix" as part of the prompt before any instruction about the content of the review itself. Spellbook's guide lists other schema choices built for different users: rank findings by severity, build a side-by-side comparison between two drafts, or produce a bulleted executive summary for a client who needs the headline risks without the clause-by-clause detail. Which schema to choose depends on who reads the output, not on which format happens to look cleanest.

The Legal Prompts guide adds a smaller but practical technique: placeholder patterns, such as [mm/dd/yyyy]: [description], that show the model what format is expected. That single technique can turn output that needs heavy rewriting into something close to a finished work product. Schema and scope reinforce each other, too. Juro notes that narrowing the scope of a request, telling the model to focus on specific issues rather than review everything, keeps the output from becoming generic or overwhelming, and a prompt that defines both the schema and the scope together forces precision at every level of the response.

There's a further benefit to schema discipline that extends past the individual prompt: once output follows a fixed, predictable structure, a mechanical validator can check it before a human ever sees it, confirming that the structure is intact, that citations are present, and that nothing is missing, in the same pass that screens for signs of hallucination. At that point the schema becomes an interface contract that downstream systems can enforce automatically. That connects directly to the fifth and final structural layer, because a schema only protects a lawyer from bad output if the prompt also tells the model what to do when it isn't sure of an answer.

Building anti-hallucination rules into the prompt as a mandatory engineering layer, not an afterthought

The fifth structural decision addresses the failure mode that can undo everything built by the first four: a model stating something false with the same confidence it uses for something true. Asked direct, verifiable questions about federal court cases, four leading general models returned hallucination rates ranging from 58% to 88%, according to research by Dahl and colleagues at a Stanford legal informatics lab. The gap between how these models perform on general legal reasoning benchmarks and how reliably they handle precise contract work remains one of the field's most significant unresolved problems. Model capability alone, in other words, can't do the job of quality control. The prompt has to do that work.

Anti-hallucination rules inside a prompt take a few concrete forms. Every factual claim should point to a specific paragraph in the document under review or to a named primary source, not to whatever the model absorbed during training. The prompt should instruct the model to flag low-confidence conclusions rather than present them with the same authority as high-confidence ones. And the prompt needs an explicit abstention rule: when a clause can't be found, the model should report its absence. HAQQ's guide frames this as a positive instruction rather than a prohibition: "Set guardrails, define what the AI must do, not just what it should avoid." None of this removes the need for a mandatory human review gate before output reaches a client-facing system. That gate is now an engineering requirement built into the process, not a workflow suggestion left to individual discretion.

One finding complicates the intuitive advice to make models "show their work." If you ask a model to reason step by step before stating a conclusion, chain-of-thought prompting generally produces more accurate, better-reasoned output on complex legal analysis. But the ContractEval benchmark, presented at the NLLP Workshop in November 2025, found that reasoning-based prompting can increase the conciseness of output while lowering correctness specifically on clause identification tasks. More elaborate reasoning instructions, applied to the wrong kind of task, can actively hurt accuracy on the exact work most contract review depends on. Chain-of-thought instructions should be matched to the task rather than applied as a blanket default, helping complex multi-step legal reasoning while potentially hurting straightforward clause-identification work.

A newer threat has entered this same layer of prompt engineering. Contract review AI now faces injected clauses embedded in documents that attempt to override the system's own instructions, a threat vector that didn't feature in earlier practitioner guidance on legal prompting. This adds input validation to what was previously treated as a pure output-quality problem. The anti-hallucination layer of a legal prompt now has to account not only for what the model might invent, but for what an adversarial document might try to make the model do.

Prompt architecture as a durable professional skill

The most serious objection to everything argued above is that purpose-built legal AI tools already embed legal logic at the model level, which would seem to reduce the need for a lawyer to engineer prompts by hand. That objection confirms the argument. Purpose-built tools work as well as they do because their developers applied the same five structural principles, role, jurisdiction, context, output schema, and anti-hallucination rules, at the level of the product. Spellbook's guide says as much directly: "Prompts for general AI tools require more structure and context, but legal-specific tools like Spellbook have built-in legal training, reducing the need for extensive prompt engineering." The need is reduced. It is not eliminated. LegalOn's contract review resource describes tools that break a contract down into structured, provision-level checks evaluated against precise legal standards, which is the output-schema principle and the role principle built into the architecture of the product itself, not a replacement for the structural thinking behind it.

What a purpose-built tool cannot embed is the lawyer's specific client, the governing law of the specific deal on the table today, the party being represented, and the particular playbook standards the firm applies. Those inputs have to come from somewhere, and a well-structured prompt, or an equivalent set of configuration choices inside a legal AI product, is the only place they can enter the process. The August 2026 New York City Bar Association policy paper makes the governance stakes explicit: AI tools may assist legal work, but they cannot substitute for professional legal judgment under existing rules of professional conduct. Professional legal judgment now includes recognizing when a prompt's architecture, or a product's built-in defaults, isn't sufficient to trust the output it produces.

Role, jurisdiction, context, output schema, and anti-hallucination rules aren't tricks specific to any one model or product generation. They are the structural requirements that any reliable legal AI system has to satisfy, whether a lawyer sets them explicitly in a prompt or a vendor has already built them into the product. A lawyer who understands these five principles gains something more durable than a set of prompt templates: the ability to audit any AI tool's output and judge, with some precision, where its embedded logic is thin and where a human still needs to supply the missing structure.

Sources

  1. Answering Questions in Stages: Prompt Chaining for Contract QA

    Provided the finding that reasoning-based prompting can increase conciseness while lowering correctness on clause identification tasks, cited as the ContractEval benchmark from the NLLP Workshop.

  2. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools

    Supplied the hallucination rate figures (58% to 88%) for leading general models on federal court case questions, attributed to Dahl and colleagues at a Stanford legal informatics lab.

Priya Subramaniam

Senior Contributing Editor

Priya spent nine years as a commercial contracts attorney at a Fortune 500 technology company before moving into legal operations consulting, where she advises in-house teams on clause standardization and negotiation strategy. Her writing focuses on the practical mechanics of contract redlining, fallback frameworks, and how legal teams translate playbooks into repeatable workflows.

More in AI in Legal Review

← Front page