Two stipulated orders in the biggest generative AI copyright case in the country set out how to produce subscriber prompts and the images and video they generated: a sample of 385 per character, drawn by a published method anyone can rerun, backed by a signed certification and a list of every record in the eligible population so the other side can check the draw itself. That last piece is something I have not seen in a sampling protocol before, and it is the difference between a sample you have to trust and a sample you can audit. When the AI output is the evidence, you need both a plan to sample it and a way to prove the sample. Reading time approximately 13 minutes.
Case: Disney Enter., Inc. v. Midjourney, Inc., Lead Case No. 2:25-cv-05275-JAK (AJRx), consolidated with No. 2:25-cv-08376-JAK (Ex) | Court: United States District Court, Central District of California | Decision: Two stipulated orders governing production of subscriber prompts and outputs. Dkt. No. 152 sets the sample size, randomization method, field list, certification and Verification List requirements, and extrapolation boundary for private (stealth mode) material. Dkt. No. 166 applies that machinery to the public material and adds a phased metadata report, a systems-not-URLs production obligation, a reasonable-request backstop, and contingent scope for a disputed set of characters. | Judge: United States Magistrate Judge A. Joel Richlin | Decided: Dkt. No. 152 entered August 10, 2026 · Dkt. No. 166 stipulated August 10, 2026, approved August 12, 2026, filed August 14, 2026
Read the sampling protocol order (Dkt. 152) on Minerva26 →
Read the public prompts order (Dkt. 166) on Minerva26 →
By Kelly Twigger
Podcast | Transcript
When the Gen AI Output Is the Evidence
Suppose the evidence in your matter is not documents in a file share, or email, or text messages. Suppose it is everything a piece of software ever produced for its users — every request somebody typed in, every result the system handed back, sitting in a database that was never built with discovery in mind. There is no natural stopping point in a data set like that. You cannot collect all of it, you cannot review all of it, and the other side is unlikely to accept whatever slice you decide to hand over.
This is a new branch on a tree we have been climbing all year. Every AI-in-discovery ruling we have covered so far, from Warner v. Gilbarco and Morgan v. V2X through Heppner, Tate, Conservation Law Foundation v. Shell, James v. Cerebras Systems, and Schulte v. LinkedIn, has asked the same question: when someone uses an AI tool, what does the other side get to see about it? In all of them, the AI was a tool somebody used to prepare or process the case.
Here the AI output is not a tool anyone used. It is the evidence itself. The studios are suing over what Midjourney’s model generated, which means that the prompts (inputs) subscribers typed and the images and video that came back (outputs) are the documents at the center of the case. When the Gen AI output is the evidence, and there is an effectively unlimited amount of it, how does anybody produce it?
The answer at that scale is sampling, and courts have been comfortable with sampling for years. But sampling raises a question most protocols never answer: how does the other side know your sample is honest? You picked it, you ran the search, you decided what came out. What these two orders do, and the reason I picked them, is answer that question in writing.
Case of the Week
What the Case Is About
In June of 2025, Disney and Universal sued Midjourney in the Central District of California, and Warner Bros. filed its own case that was later consolidated. The plaintiff group is essentially the entire American character library, from Marvel and Lucasfilm to DC Comics, Hanna-Barbera, and the Cartoon Network. Midjourney is a text-to-image and text-to-video service: a subscriber types a prompt, the model returns image or video files, and the whole exchange is stored as what the platform calls a job. If the studios want to prove subscribers were generating Darth Vader, or my personal favorite, Han Solo, or Scooby-Doo or the Minions, the proof is in those individual jobs.
Two features of the platform shape the entire production problem. The first is volume, since this is a consumer service with a very large subscriber base generating continuously. The second is stealth mode, which lets a subscriber generate privately and splits the discovery population in half. That is why there are two orders: Docket 152 handles the private material, and Docket 166 handles the public material and borrows nearly all of its machinery from 152.
How the Protocol Works, in Short
Everything is organized around a Character Bucket, a group of keywords tied to a specific character or group of characters. That is a design choice and a smart one. They did not sample the corpus as a whole, they sampled per character, because each character is a separate asserted work and a separate infringement question. Your buckets have to line up with the questions the case turns on, or a sample cannot answer anything. Consider too what building those buckets took across that entire range of studios and characters, before a single search was run.
From there the machinery is straightforward to describe. For each bucket, Midjourney produces a sample of 385 jobs, a number Section 8 ties to a margin of error of approximately plus or minus five percentage points at the 95% confidence level. The sample is drawn by applying a SHA-256 hash to each job’s unique identifier, sorting the values, and taking the 385 lowest, which shuffles the deck in a way the other side can reproduce exactly. Deduplication happens only by unique identifier, with two deliberate carve-outs: identical prompt text does not collapse two jobs into one, and a job in more than one bucket stays in all of them. Section 6 then lists more than twenty fields that travel with each job, well past prompt and output, including the parent job ID and status flags showing what Midjourney’s own systems concluded about the content.
The public order adds four things: a complete metadata report of the entire public population within 30 days, before any output file moves; an obligation to produce outputs from Midjourney’s own systems “without regard to whether such outputs remain accessible at any public URL”; a reasonable-request backstop for specific jobs the sample missed; and contingent scope for characters still caught up in a live scope dispute.
That is the summary. The provision-by-provision walkthrough is in this week’s episode, and if you are going to borrow this language for your own protocol, you can read through the Transcript, listen to the Podcast or watch the video embedded here.
The Provision I Have Not Seen Before
Section 7 is why this case is worth your time, and it does two things past the production itself.
First, a written certification signed by responsible personnel confirming that the hash function and sampling method were actually followed, that the hash was applied to the real prompt identifiers in the production database, that nothing was excluded except through the agreed criteria, and that the production contains the 385 lowest hash jobs.
Second, a Verification List. Contemporaneously with each sample, Midjourney produces a machine-readable CSV listing every prompt ID in the eligible universe — every ID the keyword search returned, before any hashing — with each ID’s hash value. No prompt text, no outputs, just identifiers and hashes, and the order makes that obligation independent of the obligation to produce the sample.
Pay attention to what that gives the requesting party. It gives them the true size of the population, the one number a producing party almost never has to disclose and the one you cannot do without if you ever want to turn a percentage into a count. It lets them sort the hashes themselves and confirm the 385 they received really are the lowest. And if a job they know about is missing from the list, they can prove the search missed it. Those two provisions together are what turn sampling from something you have to trust into something you can audit.
Where Sampling Still Breaks Down
Before you copy 385 into your own protocol, understand two limits on what that number can do.
The first is the difference between a percentage and a count. A sample gives you a share, not a total, and five points of error on a percentage becomes 5% of the population once you multiply it out. Seventy percent of a 50,000-job bucket is 35,000 give or take 2,500; seventy percent of a 50 million-job bucket is 35 million give or take 2.5 million. Identical sample, identical quality of estimate on the rate, and a range on the actual number that is a thousand times wider. If the count of infringing outputs ends up driving damages, I think that is where the fight will be.
The second limit has nothing to do with population size, and I think that is where the real exposure sits. It is about how rare the thing you are counting is. Lilith Bat-Leah at Epiq wrote this up for ACEDS in the technology assisted review context, and her example is the one to hold onto: if a sample of 383 documents turns up only eight of the thing you are measuring, the range around your estimate runs from 35% to 97%. That is not an estimate, that is a shrug. Buckets here are defined by keyword hits, and keyword hits carry noise, so if the jobs that actually depict a character are a thin slice of what the keywords pulled in, the estimate will be loose and then multiplied by a population in the millions. Nothing in Section 8 addresses how common or rare the target is inside a bucket, and the parties gave up the right to argue the sample is unrepresentative.
What OpenAI Could Not Win by Motion
Put this next to the other place courts have made an AI company produce prompts and outputs at scale. In the consolidated OpenAI copyright litigation, the news plaintiffs went after ChatGPT output logs. OpenAI had retained tens of billions in the ordinary course. Plaintiffs moved to compel 120 million, then accepted OpenAI’s own counterproposal of a 20 million de-identified sample. In October, OpenAI changed course and asked to run search terms across that sample and produce only the hits. Magistrate Judge Wang granted the motion to compel the full sample on November 7, denied reconsideration on December 2, and extended the order to the class plaintiffs on December 5. The District Court affirmed on January 5, 2026. Logs that did not contain the plaintiffs’ works were still relevant because they bore on OpenAI’s fair use defense, privacy was adequately protected by the reduction plus de-identification plus the protective order, and no case law requires a court to order the least burdensome discovery available.
Two differences are worth noting for your matters. The first is when each party asked to narrow the sample. Both companies wanted the same thing, a produced set narrowed by keyword rather than handed over whole. OpenAI asked after it had already agreed to the size of the sample, and did not get it. Midjourney has keyword-defined populations written into the definition of the sample itself, agreed before anything was collected. To be fair about the distinction underneath that, OpenAI opened its own door by raising fair use, which is what let Magistrate Judge Wang hold it could not confine production to logs touching the plaintiffs’ works. But the timing point stands: one was a stipulation written at the start, the other a motion filed after the fact.
The second difference is what each sample can be checked against. Twenty million is a negotiating number, and nothing in that process tells the requesting party how the 20 million were chosen from the tens of billions or lets them verify the selection afterward. Three hundred eighty-five is a statistical number, drawn by a published method, certified by a signed statement, and accompanied by a list of the entire eligible universe. The OpenAI plaintiffs got vastly more documents. The Midjourney plaintiffs got the ability to prove the sample is what it claims to be. A small verified sample can be worth more than a large unverified one, and a producing party that offers verification is in a much stronger position to argue for a small sample.
One caution: the Midjourney case had something the OpenAI case did not, which is defined characters to search for, and the OpenAI cases were the first of their kind. Comparing discovery decisions is almost like comparing apples to oranges, but we draw the parallels where we can. I spend real time on that comparison in the episode, including how Magistrate Judge Richlin cited the OpenAI litigation in this same case back in June while still protecting the studios’ own exploratory prompts as work product.
The Case Law Is Turning Into Protocol Language
Here is the pattern I think matters more than any single provision. Parties are reading these decisions and rewriting their protocols in light of how courts are holding, and that is not speculation. Tate told the parties to go amend their protective order to spell out whether and how confidential material may be put into an AI tool. Conservation Law Foundation held that counsel’s general language of “notes” in a Rule 29 stipulation did not reach the expert’s AI prompts, a direct instruction to name those artifacts in the document. The Cerebras protocol was negotiated by OpenAI’s lead trial counsel in the copyright cases, in the months after OpenAI lost the discovery fights that mattered in the Southern District of New York. And these two orders answer, in advance and by agreement, nearly every objection OpenAI raised and lost. That is case law turning into protocol language in under a year. The protocol you are handed this year should not look like the one you were handed last year, and if it does, somebody is not staying on top of the case law.
What to do this week
Make your sampling protocol verifiable, not just agreed. In Schulte v. LinkedIn, plaintiffs went to court for the elusion estimates, error rate, and reviewer counts behind LinkedIn’s use of Relativity aiR, and Magistrate Judge Beeler said no, that is discovery on discovery and it is disfavored. Here the producing party handed over the equivalent machinery voluntarily. What Schulte tells you that you cannot win on a motion, Midjourney and Cerebras show you that you can get by negotiation.
Use a defensible number and put the statistics in the order. Twenty million in the OpenAI litigation was a negotiating position. Three hundred eighty-five here is a statistical conclusion, with the confidence level and margin of error stated in the order itself. When the number is a conclusion rather than a compromise it is much harder to attack later.
Name the AI artifacts explicitly in every document you draft. Generic categories will not get you AI artifacts, whether you are trying to protect them or obtain them. Use the actual words: prompts, queries, outputs, job identifiers, moderation status. If you are not sure what they are, look them up.
Define what counts as a duplicate, in writing. Identical prompt text does not make two jobs one job, and a record falling in two categories stays in both. Remember the open question the magistrate judge in Concord Music Group v. Anthropic found relevant: how many prompts did it take to generate each example. You cannot answer that if your deduplication rules collapsed the attempts.
Get an inventory of the population before you get the documents. Section 3 of the public order is a report, not a production, and it lands in 30 days. Compare that to the OpenAI litigation, which opened with a motion to compel 120 million logs and took four orders and six months to work down from there, with both sides arguing volume blind the whole way. Ask for the inventory first and you negotiate with numbers instead of adjectives.
Never agree to sampling without a path to verify it. This is the provision OpenAI wanted and could not get. Without a written mechanism to request specific items outside the sample, you have agreed to a ceiling and called it a methodology.
Build the amendment clause in when you draft. These protocols can be modified by written stipulation without leave of Court, which is a deliberate choice worth noticing, since we have covered protocols that required a showing of good cause instead. Either way, the technology in your case will change during the case, and Schulte is what happens when nobody goes back to update the order.
Know the shape of the prompt rules and go back to the episodes for the details. Your client’s own prompts used in the litigation are generally work product, though it depends on the jurisdiction, and Tate was in Texas state court with a specific rule protecting prompts prepared by or for a party. Your expert’s prompts are generally discoverable. The other side’s product prompts are producible on a protocol like this one. The details live in the cases, all of them in Minerva26 under the Generative AI tag.
Do not just make a plan. Plan for the plan. Somebody had to know that Midjourney stores an exchange as a job, and that the job is the unit you sample. Somebody had to know stealth mode existed and split the population in half. Somebody had to think about how people actually use the product, running the same prompt over and over toward a result, and realize that deduplicating on prompt text would erase the evidence of that repetition, which is the entire volume story. None of that is legal drafting. All of it is understanding the technology first and writing the protocol second.
My Take and What to Watch For
The most interesting line in either order is the boundary the parties drew in Section 8. Experts may extrapolate from the sample to estimate how many prompts generated outputs depicting a character, and the parties gave up the right to say the sample is unrepresentative. They expressly kept the right to challenge any conclusion that particular outputs are infringing. So the count can be extrapolated and the infringement cannot.
That boundary holds cleanly in discovery. I am not at all sure it holds at the damages stage, and this is my read rather than anything the Court has decided. Statutory damages under the Copyright Act run per work infringed, and an estimate of how many prompts produced a character, plus or minus a few points that may translate into millions of jobs, is not the same thing as a finding of how many works were infringed. The parties settled the process. What they did not settle, and what no court has answered, is what a number produced that way is good enough to prove. My guess is the first real fight comes at the expert stage.
The other thing I am watching is how far outside AI this travels, because none of this sampling process is really about artificial intelligence. The 385, the hash, the certification, and the Verification List work for any data source too large to produce whole: consumer platform content, chat and messaging across large custodian sets, call recordings, telematics and sensor data, clickstream and server logs, body cam video, transaction data.
Listen to the full episode
This post is the summary. The episode is where the good stuff lives. I go through both orders provision by provision: why the sample size is 385 and what the number actually buys you, how the SHA-256 draw makes the selection reproducible and why that matters more than it sounds, the two deduplication carve-outs and the damages theory underneath them, the full Section 6 field list and what those moderation flags tell the studios, the certification and Verification List in detail, and the phased report and URL clause in the public order. If you are drafting or negotiating a sampling protocol, that is the version to listen to. Listen on Meet and Confer →
See Minerva26 in action
Minerva26 is the discovery intelligence platform that connects case law, rules, and real-world workflows. We include the stipulated orders and protocols themselves, not just the opinions, and you can now search directly for stipulated orders or filter them out, so when you are drafting a sampling protocol the language courts have actually entered is already organized for you. Book a 30-minute demo
Related on Minerva26: James v. Cerebras Systems, Inc. · Schulte v. LinkedIn Corp. · Warner v. Gilbarco, Inc. · Morgan v. V2X, Inc. · United States v. Heppner · Tate Group Automotive, LLC v. Legacy Automotive Capital, LLC · Conservation Law Foundation, Inc. v. Shell Oil Co. · Concord Music Group, Inc. v. Anthropic PBC · In re OpenAI, Inc., Copyright Infringement Litigation
This decision is available on the Minerva26 platform with full issue tagging. If you’re a litigator navigating discovery strategy and want to stay ahead of decisions like this one, visit Minerva26.com to learn more or schedule a demo. Every decision covered on Case of the Week is searchable by issue, jurisdiction, and judge.
Kelly Twigger is CEO and founder of Minerva26 and Principal at ESI Attorneys. She has been a discovery strategist and practicing attorney for nearly 30 years. Case of the Week is a segment of the Meet and Confer podcast, breaking down one recent ESI discovery decision each week into practical strategy you can use.
- 📩 Get the newsletter
- 🎧 Listen to Meet and Confer
- 🔁 Share this post on LinkedIn


