An analysis of the Northern District of California’s July 2026 stipulated ESI protocol and companion protective order — the first I’ve seen to govern generative AI as its own category of document review — and what it means for how you negotiate discovery in the age of AI. Reading time approximately 12 minutes.
Case: James v. Cerebras Systems, Inc., No. 4:25-cv-09361 · Court: United States District Court, Northern District of California · Orders: ECF No. 74 (ESI Protocol) and ECF No. 75 (Stipulated Protective Order) · Entered: July 7, 2026 · Judge: Magistrate Judge Robert M. Illman
By Kelly Twigger
Podcast | Transcript
A first: parties wrote the AI rulebook for discovery, and simply agreed to it
For the past year on Case of the Week, I have tracked how courts are working through artificial intelligence in discovery — the privilege fights over a litigant’s prompts, the judges beginning to write AI terms into protective orders, the slow realization that the rules we have were not built for any of this. James v. Cerebras Systems is the milestone in that arc.
What the parties did here is, as far as I have seen, a first: a stipulated ESI protocol that governs the use of generative AI to review and produce documents as its own category — with its own disclosure rules, its own validation math, and a requirement that a party using AI to decide what gets produced disclose the prompts it used to do it. And they agreed to all of it. Nobody fought. That, as much as any single provision, is the story. My first reaction was that I would never agree to any of this — which is exactly why it is worth understanding before some version of it lands on your desk. Orders like this become templates, and the lawyers who wrote this one are writing the AI-copyright playbook for the rest of the country.
There are two orders, both entered by Magistrate Judge Robert M. Illman on July 7, 2026, and you have to keep them straight. Document 74 is the ESI Protocol — how you use AI to review and produce. Document 75 is a companion Stipulated Protective Order — how you protect sensitive material and use AI tools on it. They are designed to interlock, and if you read one without the other you will misunderstand both.
Case of the Week
Nothing in the Federal Rules required any of this
Hold onto one frame throughout. Nothing in the Federal Rules required what these parties agreed to. Rule 26 obligates a party to provide relevant, proportional discovery. Rule 34 governs the form of production. Neither says a word about disclosing your prompts, your validation methodology, or your elusion rate. Those obligations exist here for one reason: the parties agreed to take them on.
This was a stipulated order — negotiated and consented to — not one a court imposed after a fight, and not the scorched-earth battle we watched in the OpenAI copyright cases. We have debated for years on this show how much parties are really obligated to share on TAR and search terms; the answer has always been left to what parties agree to or what a court orders on a motion. We have not seen sophisticated parties on both sides simply agree, up front, to a disclosure regime this far past what Rule 26 requires.
So keep asking: why would they? Much of what this protocol demands brushes against material that is arguably privileged or work product — your prompts, how you trained your model, your validation results, your error rates. The most plausible read is that counsel who lived through the OpenAI discovery fights decided a detailed, agreed protocol up front is cheaper and more predictable than litigating every AI question by motion, one expensive fight at a time — and that OpenAI lost most of those fights and had to disclose anyway. This level of detail buys peace. But you buy that peace by expanding your own obligations and trading away protections. That is a real strategic decision, not an obvious one, and it goes to a theme we return to constantly: sit down early, understand the issues and who is across the table, and draft to the case in front of you. Sometimes buying peace is right — but do it on purpose, not by pulling a form off a shelf.
One protocol, three ways to cull: search terms, TAR, and AI
The AI piece only makes sense once you see where it sits. Section VII — Search and Review — authorizes and governs three ways to cull documents, and it addresses combining them. Search terms, capped at twenty per custodian per party, with hit reports and a null-set test on the documents that do not hit. TAR — the predictive-coding process where you train a classifier to rank documents by likely responsiveness — which gets its own rulebook in Appendix 3. And artificial intelligence, which gets its own separate rulebook in Appendix 4. Section VII.6 governs “layering,” using more than one method on the same set. You choose your method, disclose that choice, and if it is TAR or AI you hand over the matching appendix. What is new is that AI is not folded under the TAR umbrella. It is broken out as its own regime.
AI review becomes its own disclosed category
Section VII.8 creates a regime the order calls “AI Responsiveness Review,” defined broadly as workflows using “large-language models (LLMs), deep-learning classifiers, embedding-based similarity analysis, semantic clustering, predictive redaction systems, or generative summarization models” to decide whether a document is produced, withheld, or redacted. Six technologies, not one — a far wider net than TAR. If you use any of them to make responsiveness or privilege calls, you must disclose that election, and under Section VII.9 provide an AI Review Protocol containing everything in Appendix 4.
The workflow reality of Appendix 4 is sweeping: if you use AI to decide what gets produced, how the AI reached those decisions must be documented, your privilege calls must be human-confirmed, and the entire process is disclosed to your opponent. Appendix 4 requires you to disclose, at a minimum:
- The system — the identity, version, and hosting environment of each model (on-premise, private cloud, or vendor-secured), plus a statement that it is “generally accepted within the eDiscovery industry as a reasonably reliable tool.” If you use these tools, you have seen the “relaunch” prompt that appears almost daily — meaning you may be on a different model version day to day. Does that mean documenting every model change? It reads like yes.
- The universe — the criteria used to select the reviewed documents, including any culling before the AI ran (search terms, dates, custodians, deduplication, threading, near-duplicate suppression), and which document types were excluded and how they were reviewed instead — down to the reviewers’ credentials.
- The training and prompts — the methodology used to train or instruct the system; the sources of training and validation data; who designed, trained, and validated it; documents excluded from training; and disclosure of all prompts, templates, instruction sets, and parameter configurations, with any prompt change served in redline within three business days. Think about how often you iterate a prompt — that is a feat on a three-day clock, and it demands a well-defined process up front or it becomes a nightmare.
- The scoring — how the system generates scores, classifications, or decision boundaries and the thresholds used; the score distribution from the initial run; and how many documents landed in each category.
- The oversight — who conducted oversight, privilege review, and validation; the review workflow; the number manually reviewed; and quality-control procedures that must include human confirmation of privilege calls, checks for hallucination, over-summarization, and misclassification, and testing of prompt performance.
- The iteration — how many times the model or prompts were adjusted and how that changed results.
- The validation — a sample of the documents the AI excluded as non-responsive, sized at a 95% confidence level with a plus-or-minus 2% margin of error; the resulting richness, recall, precision, and elusion rates; and a duty to meet and confer and take corrective measures if elusion exceeds 3%.
- The audit trail and security — all AI processing inside the secure review environment, with no content exported to an outside model absent written agreement; retained audit materials including prompt-iteration logs and configuration records; and any genuinely proprietary prompt withheld but logged and made available for attorneys’-eyes-only inspection under the Protective Order.
That is a remarkable amount of transparency to sign up for. Here are the pieces that matter most.
Prompts are the provision people will fixate on. You disclose all of them, and every change, in redline on a three-day clock. We have covered whether a litigant’s own prompts are protected: Warner v. Gilbarco said yes, they are work product, and gave us the line that an AI platform is a “tool, not a person”; Morgan v. V2X agreed for a pro se litigant. So there is law that prompts can be protected — and here the parties simply agreed to hand them over. That is not a court overriding a work-product objection; it is both sides trading that protection away by stipulation. There is one escape hatch, and it is the first place the two orders touch: a genuinely proprietary prompt can be withheld, but it must be logged and shown to opposing counsel’s eyes only under the Protective Order. And note a distinction that will matter as this law develops. In Warner and Morgan, the prompts were the litigant’s own analytical prompts — case strategy, mental impressions. Here, prompts are being used to scope the review population and the production set — a process that has traditionally been privileged too. This protocol changes that dynamic by stipulation, and whether it becomes a norm is one of the most important open questions in discovery right now.
The validation numbers are set at the demanding end. A 95% confidence level is standard. But a plus-or-minus 2% margin of error is tight — a lot of TAR validation runs at a 5% margin, and tightening from 5% to 2% multiplies your validation sample by roughly six times, because sample size grows with the square of the precision you demand. And an elusion expectation of 3% or lower is aggressive: the case law that blessed TAR — Da Silva Moore, Rio Tinto — focused on whether the process was reasonable and cooperative, not on a hard numeric target, and recall of 70 to 80 percent has generally been treated as defensible. What these parties agreed to sits at the very demanding end of anything the TAR world has required, now applied to AI. If you copy these numbers without understanding what they cost to meet, you will regret it.
The AI cannot absorb your Rule 26(g) duty. Appendix 4 requires checks for hallucination, over-summarization, and misclassification — and remember Rule 37(b) authorizes sanctions for violating a court order. It also states that “legal counsel remains responsible for ensuring the accuracy and completeness of all productions.” You cannot point at the model when a production is wrong. These tools do not lower the standard of care; they raise it. On security, all processing must stay inside the secure review environment, with no content piped out to a public model. If that is not how you want your case to run, do not copy this language — adjust it for the way the technology works in your matter.
And watch that “generally accepted” phrase. It will not be a fight here because the parties agreed. But who decides what is generally accepted for a tool that may be six months old? We have decades of literature validating TAR and nothing like that consensus for LLMs used to make responsiveness calls, and the tools change faster than anyone can evaluate them. That undefined standard is going to get tested somewhere.
Layering — Section VII.6 — is the quiet provision I think is genius. Layering stacks two culling methods on the same set, say search terms and then TAR. Each method independently misses some responsive documents, so stacking them compounds the loss: the keyword pass strips out responsive documents that lack your terms, and the TAR or AI system never sees them. So the protocol makes you disclose the intent to layer, meet and confer, compare hit counts with and without layering, and let the requesting party sample the excluded set. The risk is not the technology; it is the undisclosed, unvalidated combination of technologies. We have not seen anything this comprehensive on layering before.
How the two orders interlock — and why over-designation is the risk
The scope difference between the orders drives a strategic point. The ESI Protocol governs your entire review and production — every document, whether or not it is ever designated. The Protective Order reaches only material a party actually designates as protected. The protocol applies to everything; the protective order applies to a subset.
They meet most importantly at Section 13.4 of the Protective Order, which lets a party run Protected Material through generative AI tools only where the tool does not train on it and meets industry-standard security — and then names names: “the parties’ enterprise ChatGPT and Harvey accounts satisfy the foregoing requirements.” Read that carefully. It does not mean those are the only tools you can use, or that they are blessed for everyone. It means these parties reviewed the data processing agreements — the DPAs — behind their own enterprise ChatGPT and Harvey accounts, confirmed those contracts prohibit training on the data and meet the security bar, and put that finding in the order. This is a contract point, not a technology point. Plenty of tools meet the same bar. The difference between a consumer and an enterprise account is entirely the data terms — the enterprise DPA promises no training and delivers the controls; a consumer account does not. So if you adopt this language, the takeaway is not “use ChatGPT or Harvey.” It is: review your own tool’s DPA, confirm it meets the requirements, and get that verification into the record.
This descends directly from Morgan v. V2X, where Magistrate Judge Braswell wrote AI-specific language into a protective order requiring any AI platform to carry a contractual training prohibition, third-party flow-down, a right to delete, and written documentation — functionally, the terms of a DPA, imposed from the bench. These parties stipulated to that same logic and went further by verifying and naming their tools. (Note the contrast with Warner, where the Court held the parties did not have to disclose their tools at all.)
And here is the over-designation problem I flagged when we covered Morgan. Because Section 13.4 attaches only to Protected Material, its weight depends entirely on how much gets designated — creating a new incentive to over-designate. If designating material controls what tools the other side can run it through, a producing party has a reason to designate aggressively, to keep its data out of the opponent’s unvetted tools. That is a different axis from Morgan, where the over-breadth was in the tool definition (“any modern artificial intelligence platform” sweeping in Westlaw, Relativity, Copilot); here it is in the scope of designation. Same disease. The guardrails are in the order — Section 5.1 bars “mass, indiscriminate, or routinized designations” and threatens sanctions, and Section 6 gives a challenge procedure — but they only work if the court enforces them and counsel actually challenges. Designation now carries AI-usage consequences it did not carry a few years ago.
Two more provisions show how carefully this was tailored. Section 2.19 makes Training Data attorneys’ eyes only, so the datasets at the heart of the merits — Books3 and The Pile — are viewable by opposing counsel and experts under AEO rather than walled off in the top tier. Section 2.18 defines Source Code for this dispute specifically — “training pipeline code,” “data ingestion or filtering scripts,” “tokenization or preprocessing scripts,” “model architecture files” — while carving Training Data out. And the ESI Protocol says source code is not covered and gets its own separate protocol, which tells you the real technical fight, over how Cerebras-GPT was built, is still coming.
The access-to-justice gap is widening
A protocol like this — hosting disclosures, 95% validation, audit logs, prompt redlines on a three-day clock — is only achievable by a well-resourced party with sophisticated vendors. At the other end are the pro se litigants we saw in Warner and Morgan, a growing wave using consumer AI to run their own cases and creating real problems for courts and opposing counsel who must sift through filings and productions nobody validated or governed. The same technology is producing two opposite problems at once: at the top, elaborate voluntary governance; at the bottom, none. This order is a template for the first group and does nothing for the second, and the gap is widening. Magistrate Judge Braswell raised it in Morgan, and it is only getting sharper.
Put this on the map we have built all year. In Warner, Heppner, and Morgan, the AI user was a litigant using a chatbot to think through a case, and the question was privilege. Here the AI user is counsel and their vendors, using AI as the engine that culls and produces, and the question is governance. Same technology, completely different discovery problem. What is striking is that these parties did not wait for anyone to rethink the rules. They wrote their own, by agreement, and a magistrate judge signed them.
What to do this quarter
Start with the obligation that reaches every matter: your clients are already using AI, whether their businesses sanction it or not. You now have an affirmative obligation to understand how AI artifacts and evidence are created and to plan to preserve them, produce them, and ask for them in discovery. This is where key evidence will live — and as the OpenAI cases have shown, prompts can be evidence of intent; what a person asked the model, and how, can go directly to state of mind. If you are in-house, understand how AI is used inside each business unit. If you are outside counsel, sit down with clients now, before your duty to preserve arises.
Before you adopt a protocol like this, decide whether you want its obligations. Nothing in Rules 26 or 34 requires prompt disclosure, validation reporting, or elusion testing. This is a stipulation, and stipulations bind you. Make the trade on purpose, not by copying a form.
If you will use AI to review, build your AI Review Protocol before you deploy the tool — settle your hosting environment, thresholds, validation plan, and prompt-disclosure position in advance, because you may be handing all of it to the other side. If you are requesting, this order legitimizes real leverage: demand the AI election, the validation metrics, the elusion testing, and the prompts, and ask whether they checked for hallucination and over-summarization.
Know your data terms. Other tools qualify beyond the two named here, but you have to review your own tool’s DPA and confirm it meets the no-training and security requirements — a consumer account will not clear the bar. Get the verification into the record. And watch your designations: because the AI-tool restriction attaches only to Protected Material, over-designation now carries an added cost. Designate with discipline if you produce; challenge over-designation under Section 6 if you receive.
Do not overlook the unglamorous detail. This protocol handles Teams and Slack messages in units of no less than 24 hours, preserving emojis and threading; lets a requesting party demand current versions of up to 100 hyperlinked documents within 14 days; requires version history from collaboration tools; and sets the privilege log in Excel within 45 days. Sophisticated parties sweat these details for a reason.
If you take nothing else: generative AI in discovery is no longer riding under the TAR umbrella, and the parties writing these protocols now are setting the template you will be handed next year. A stipulation is a choice. Before you sign one that looks like this, be sure you want everything in it — because nothing in the rules made you agree to it, and everything in it will bind you.
Listen to the full episode
This week’s Case of the Week segment of the Meet and Confer podcast breaks down both orders in James v. Cerebras Systems — the AI Responsiveness Review regime, the prompt-disclosure requirement, the validation math, and how the ESI protocol and protective order work hand in hand. Listen on Meet and Confer →
See Minerva26 in action
Minerva26 is the discovery intelligence platform that connects case law, rules, and real-world workflows. We curate and tag close to 50,000 discovery decisions by issue, so when you are preparing for a meet-and-confer, drafting a Rule 29 stipulation, or building an opposition to a motion to compel expert AI use, the rulings you need are already organized for you. Book a 30-minute demo
Related on Minerva26: Warner v. Gilbarco · U.S. v. Heppner · Morgan v. V2X · the OpenAI copyright cases · Da Silva Moore v. Publicis · Rio Tinto v. Vale
This decision is available on the Minerva26 platform with full issue tagging. If you’re a litigator navigating discovery strategy and want to stay ahead of decisions like this one, visit Minerva26.com to learn more or schedule a demo. Every decision covered on Case of the Week is searchable by issue, jurisdiction, and judge.
If it’s about the discovery of ESI, it’s covered in Minerva26, your discovery strategy platform.
- 📩 Get the newsletter
- 🎧 Listen to Meet and Confer
- 🔁 Share this post on LinkedIn


