1 2 3 4 UNITED STATES DISTRICT COURT 5 NORTHERN DISTRICT OF CALIFORNIA 6 EUREKA DIVISION 7 8 PAUL TREMBLAY, et al., Case No. 23-cv-03223-AMO (RMI)
9 Plaintiffs, ORDER RE: FOURTH DISCOVERY 10 v. DISPUTE
11 OPENAI, INC., et al., Re: Dkt. No. 153 12 Defendants.
13 14 Now pending before the court is a jointly-filed letter brief setting forth a discovery dispute 15 through which Defendants seek to compel certain discovery over which Plaintiffs have asserted a 16 work product privilege. See Ltr. Br. (dkt. 153) at 1-5. Pursuant to Federal Rule of Civil Procedure 17 78(b) and Civil Local Rule 7-1(b), the court finds the matter suitable for disposition without oral 18 argument. For the reasons stated below, Defendants’ request to compel the material in question is 19 granted. 20 By way of background, Plaintiffs (a group of authors) allege that Defendants’ ChatGPT 21 software relies on a large language model by which it is trained through “copying massive 22 amounts of text and extracting expressive information from it,” and that “[o]nce the large language 23 model has copied and ingested the text in its training dataset, it is able to emit convincingly 24 naturalistic text outputs in response to user prompts.” See First Amend. Compl. (“FAC”) (dkt. 25 120) at ¶ 2. Plaintiffs further allege that “when ChatGPT is prompted, ChatGPT generates 26 summaries of Plaintiffs’ copyrighted works – something only possible if ChatGPT was trained on 27 Plaintiffs’ copyrighted works.” Id. at ¶ 5. More specifically, Plaintiffs’ have alleged that “[w]hen 1 accurate summaries. These summaries are attached [to the FAC] as Exhibit B. The summaries get 2 some details wrong, which is expected, since a large language model mixes together expressive 3 material derived from many sources. Still, the rest of the summaries are accurate, which means 4 that ChatGPT retains knowledge of particular works in the training dataset and is able to output 5 similar textual content.” Id. at ¶ 51. Further, Exhibit-B to the FAC sets forth several prompts that 6 ask Chat GPT to summarize in detail various parts of Plaintiffs’ writings. See FAC Exh. B (dkt. 7 120-2) at 2-17, 20, 23, 26-27, 29, 35-38. The Exhibit also sets forth numerous queries to ChatGPT 8 – along with their responses – about a number of Plaintiffs’ works; the queries resemble questions 9 and answers that one might encounter in a literature class (e.g., What are the main themes of this 10 work? What are examples of nature and beauty in this work? What are some examples of the 11 immigrant experience in this work?). See id. at 18, 21, 24, 30, 33, 39. Additionally, the Exhibit 12 also sets forth the queries and responses on a number of occasions where ChatGPT was asked to 13 write a paragraph, or to compose a screenplay, either in the style of one of the Plaintiffs or “like” 14 one of their works. See id. at 19, 22, 25, 28, 31-32, 34, 40. 15 The current dispute concerns Defendants’ RFP 9, which seeks “[a]ll non-privileged 16 Documents and Communications relating to [Plaintiffs’] investigation of the claims alleged in the 17 Complaint.” Ltr. Br. (dkt. 153) at 1. Defendants then narrowed this request to encompass: “(a) the 18 OpenAI account information for individuals who used ChatGPT to investigate Plaintiffs’ claims; 19 and (b) the prompts and outputs for Plaintiffs’ testing of ChatGPT in connection with their pre-suit 20 ChatGPT testing, including prompts and outputs that did not reproduce or summarize Plaintiffs’ 21 works or otherwise support Plaintiffs’ claims, along with documentation of Plaintiffs’ testing 22 process.” Id. (emphasis added). Defendants maintain that Plaintiffs have “refuse[d] to respond in 23 full based on a claim of work product protection, offering to produce only ‘full threads of the 24 prompts and outputs’ that led to the examples in Exhibit-B to the Complaint.” Id. In short, 25 Defendants contend that Plaintiffs have “offered up only their preferred, cherry-picked results” by 26 refusing to tender the prompts and results that did not improperly reproduce or summarize 27 Plaintiffs’ works. Id. 1 prompt-and-output information for three reasons: (1) because Plaintiffs revealed or placed their 2 work product at issue during the course of the litigation by including allegations regarding how 3 ChatGPT responded to Plaintiffs’ prompts made from those accounts in the FAC and in Exhibit-B 4 thereto; (2) because Plaintiffs voluntarily disclosed the information in question to their adversary 5 in litigation; and, (3) because the account information and the totality of the prompt-and-response 6 data that Plaintiffs used to interrogate ChatGPT should, in fairness, be considered together with 7 Plaintiffs’ testing results (set forth in Exhibit-B) and that the totality of that data would be needed 8 for Defendants to subject Plaintiffs’ claims to meaningful adversarial testing. See id. at 1-3. 9 Specifically, as to the account data component of the information that Defendants seek, 10 Defendants submit that because “the ‘custom instructions’ feature allows users to ‘add preferences 11 or requirements’ for ‘ChatGPT to consider when generating its responses,’” Defendants “need[] 12 this discovery to test Plaintiffs’ allegations regarding ChatGPT’s behavior in response to the 13 ‘interrogation’ Plaintiffs chose to put at issue.” Id. at 3. 14 Plaintiffs assert that this material is shielded from discovery because it is attorney work 15 product – Plaintiffs add that Defendants “seek[] wide-ranging discovery into Plaintiffs’ counsels’ 16 investigatory files, regardless of whether the prompts and outputs were used in the complaint, and 17 regardless of whether the prompts and outputs ‘support Plaintiffs’ claims.’” Id. at 3. Plaintiffs 18 submit that they have agreed to tender prompts and outputs that support their claims and were set 19 forth in the FAC and in Exhibit-B thereto. Id. at 5. However, the “prompts and outputs that did not 20 reproduce or summarize Plaintiffs’ works or otherwise support Plaintiffs’ claims” are shielded 21 from discovery because to produce those would be to divulge the “thoughts and analysis 22 conducted by Plaintiffs’ lawyers in preparation of the litigation.” Id. at 3. Plaintiffs, therefore, 23 consider the prompts and responses that did not improperly reproduce or summarize their works to 24 be opinion work product. Id. at 4. 25 As to Defendants’ waiver argument, Plaintiffs contend that the disclosure of some of their 26 prompts and responses did not operate as a waiver regarding the undisclosed prompts and 27 responses because those disclosures (in the FAC and Exhibit-B) were necessary in order to satisfy 1 responses set forth in the FAC and in Exhibit-B “are not evidence which the jury will be asked to 2 rely on . . . [as] Plaintiffs will rely on other, more concrete evidence obtained through discovery to 3 identify the literary works OpenAI used to train their products.” Id. The way Plaintiffs see it, “the 4 fact that ChatGPT did not always produce a summary of Plaintiffs’ work when prompted is 5 immaterial because the summaries were simply used to plausibly allege that Defendants trained 6 their products on Plaintiffs’ Asserted Works[] [and] ‘[n]egative’ test results are not relevant for the 7 same reason the ‘positive’ test results are not particularly relevant at this stage.” Id. At bottom, 8 Plaintiffs state that Defendants seek to obtain “a free ride on the work of [their] attorney,” which 9 they urge the court to reject due to the suggestion that “this information is not relevant[] [a]nd 10 OpenAI can obtain the material itself – [because] it can interrogate ChatGPT for itself.” Id. 11 The Federal Rules of Civil Procedure
Free access — add to your briefcase to read the full text and ask questions with AI
1 2 3 4 UNITED STATES DISTRICT COURT 5 NORTHERN DISTRICT OF CALIFORNIA 6 EUREKA DIVISION 7 8 PAUL TREMBLAY, et al., Case No. 23-cv-03223-AMO (RMI)
9 Plaintiffs, ORDER RE: FOURTH DISCOVERY 10 v. DISPUTE
11 OPENAI, INC., et al., Re: Dkt. No. 153 12 Defendants.
13 14 Now pending before the court is a jointly-filed letter brief setting forth a discovery dispute 15 through which Defendants seek to compel certain discovery over which Plaintiffs have asserted a 16 work product privilege. See Ltr. Br. (dkt. 153) at 1-5. Pursuant to Federal Rule of Civil Procedure 17 78(b) and Civil Local Rule 7-1(b), the court finds the matter suitable for disposition without oral 18 argument. For the reasons stated below, Defendants’ request to compel the material in question is 19 granted. 20 By way of background, Plaintiffs (a group of authors) allege that Defendants’ ChatGPT 21 software relies on a large language model by which it is trained through “copying massive 22 amounts of text and extracting expressive information from it,” and that “[o]nce the large language 23 model has copied and ingested the text in its training dataset, it is able to emit convincingly 24 naturalistic text outputs in response to user prompts.” See First Amend. Compl. (“FAC”) (dkt. 25 120) at ¶ 2. Plaintiffs further allege that “when ChatGPT is prompted, ChatGPT generates 26 summaries of Plaintiffs’ copyrighted works – something only possible if ChatGPT was trained on 27 Plaintiffs’ copyrighted works.” Id. at ¶ 5. More specifically, Plaintiffs’ have alleged that “[w]hen 1 accurate summaries. These summaries are attached [to the FAC] as Exhibit B. The summaries get 2 some details wrong, which is expected, since a large language model mixes together expressive 3 material derived from many sources. Still, the rest of the summaries are accurate, which means 4 that ChatGPT retains knowledge of particular works in the training dataset and is able to output 5 similar textual content.” Id. at ¶ 51. Further, Exhibit-B to the FAC sets forth several prompts that 6 ask Chat GPT to summarize in detail various parts of Plaintiffs’ writings. See FAC Exh. B (dkt. 7 120-2) at 2-17, 20, 23, 26-27, 29, 35-38. The Exhibit also sets forth numerous queries to ChatGPT 8 – along with their responses – about a number of Plaintiffs’ works; the queries resemble questions 9 and answers that one might encounter in a literature class (e.g., What are the main themes of this 10 work? What are examples of nature and beauty in this work? What are some examples of the 11 immigrant experience in this work?). See id. at 18, 21, 24, 30, 33, 39. Additionally, the Exhibit 12 also sets forth the queries and responses on a number of occasions where ChatGPT was asked to 13 write a paragraph, or to compose a screenplay, either in the style of one of the Plaintiffs or “like” 14 one of their works. See id. at 19, 22, 25, 28, 31-32, 34, 40. 15 The current dispute concerns Defendants’ RFP 9, which seeks “[a]ll non-privileged 16 Documents and Communications relating to [Plaintiffs’] investigation of the claims alleged in the 17 Complaint.” Ltr. Br. (dkt. 153) at 1. Defendants then narrowed this request to encompass: “(a) the 18 OpenAI account information for individuals who used ChatGPT to investigate Plaintiffs’ claims; 19 and (b) the prompts and outputs for Plaintiffs’ testing of ChatGPT in connection with their pre-suit 20 ChatGPT testing, including prompts and outputs that did not reproduce or summarize Plaintiffs’ 21 works or otherwise support Plaintiffs’ claims, along with documentation of Plaintiffs’ testing 22 process.” Id. (emphasis added). Defendants maintain that Plaintiffs have “refuse[d] to respond in 23 full based on a claim of work product protection, offering to produce only ‘full threads of the 24 prompts and outputs’ that led to the examples in Exhibit-B to the Complaint.” Id. In short, 25 Defendants contend that Plaintiffs have “offered up only their preferred, cherry-picked results” by 26 refusing to tender the prompts and results that did not improperly reproduce or summarize 27 Plaintiffs’ works. Id. 1 prompt-and-output information for three reasons: (1) because Plaintiffs revealed or placed their 2 work product at issue during the course of the litigation by including allegations regarding how 3 ChatGPT responded to Plaintiffs’ prompts made from those accounts in the FAC and in Exhibit-B 4 thereto; (2) because Plaintiffs voluntarily disclosed the information in question to their adversary 5 in litigation; and, (3) because the account information and the totality of the prompt-and-response 6 data that Plaintiffs used to interrogate ChatGPT should, in fairness, be considered together with 7 Plaintiffs’ testing results (set forth in Exhibit-B) and that the totality of that data would be needed 8 for Defendants to subject Plaintiffs’ claims to meaningful adversarial testing. See id. at 1-3. 9 Specifically, as to the account data component of the information that Defendants seek, 10 Defendants submit that because “the ‘custom instructions’ feature allows users to ‘add preferences 11 or requirements’ for ‘ChatGPT to consider when generating its responses,’” Defendants “need[] 12 this discovery to test Plaintiffs’ allegations regarding ChatGPT’s behavior in response to the 13 ‘interrogation’ Plaintiffs chose to put at issue.” Id. at 3. 14 Plaintiffs assert that this material is shielded from discovery because it is attorney work 15 product – Plaintiffs add that Defendants “seek[] wide-ranging discovery into Plaintiffs’ counsels’ 16 investigatory files, regardless of whether the prompts and outputs were used in the complaint, and 17 regardless of whether the prompts and outputs ‘support Plaintiffs’ claims.’” Id. at 3. Plaintiffs 18 submit that they have agreed to tender prompts and outputs that support their claims and were set 19 forth in the FAC and in Exhibit-B thereto. Id. at 5. However, the “prompts and outputs that did not 20 reproduce or summarize Plaintiffs’ works or otherwise support Plaintiffs’ claims” are shielded 21 from discovery because to produce those would be to divulge the “thoughts and analysis 22 conducted by Plaintiffs’ lawyers in preparation of the litigation.” Id. at 3. Plaintiffs, therefore, 23 consider the prompts and responses that did not improperly reproduce or summarize their works to 24 be opinion work product. Id. at 4. 25 As to Defendants’ waiver argument, Plaintiffs contend that the disclosure of some of their 26 prompts and responses did not operate as a waiver regarding the undisclosed prompts and 27 responses because those disclosures (in the FAC and Exhibit-B) were necessary in order to satisfy 1 responses set forth in the FAC and in Exhibit-B “are not evidence which the jury will be asked to 2 rely on . . . [as] Plaintiffs will rely on other, more concrete evidence obtained through discovery to 3 identify the literary works OpenAI used to train their products.” Id. The way Plaintiffs see it, “the 4 fact that ChatGPT did not always produce a summary of Plaintiffs’ work when prompted is 5 immaterial because the summaries were simply used to plausibly allege that Defendants trained 6 their products on Plaintiffs’ Asserted Works[] [and] ‘[n]egative’ test results are not relevant for the 7 same reason the ‘positive’ test results are not particularly relevant at this stage.” Id. At bottom, 8 Plaintiffs state that Defendants seek to obtain “a free ride on the work of [their] attorney,” which 9 they urge the court to reject due to the suggestion that “this information is not relevant[] [a]nd 10 OpenAI can obtain the material itself – [because] it can interrogate ChatGPT for itself.” Id. 11 The Federal Rules of Civil Procedure define work product as “documents and tangible 12 things that are prepared in anticipation of litigation or for trial by or for another party or its 13 representative . . .” Fed. R. Civ. P. 26(b)(3)(A); see also United States v. Richey, 632 F.3d 559, 14 567-68 (9th Cir. 2011). Although work product is generally exempt from disclosure in discovery, 15 a party can still obtain work product of an opposing party if the requesting party shows a 16 “substantial need for the materials to prepare its case and cannot, without undue hardship, obtain 17 their substantial equivalent by other means.” Fed. R. Civ. P. 26(b)(3)(A). 18 There are two categories of work product. As to the first, fact work product is discoverable 19 upon a showing of “substantial need” and “undue hardship.” Fed. R. Civ. P. 26(b)(3). Substantial 20 need consists of the relative importance of the information in the documents to the party’s case 21 and the ability to obtain that information from other means, and “[t]he test is not only relevancy 22 but that there are no other available means by which to secure the same information.” AT&T Corp. 23 v. Microsoft Corp., 2003 U.S. Dist. LEXIS 8710, 2003 WL 21212614, at *6 (N.D. Cal. Apr. 18, 24 2003). “Undue hardship is demonstrable if witnesses are unavailable or cannot recall the events in 25 question.” Id. On the other hand, opinion work product is defined as material prepared by an 26 attorney that contains “mental impressions, conclusions, opinions, or legal theories of a party’s 27 attorney or other representative concerning the litigation.” Fed. R. Civ. P. 26(b)(3)(B). In other 1 application of facts to those theories, rather than the bare facts or legal theories alone. While it is 2 sometimes said that opinion work product is “virtually undiscoverable” (see e.g., Republic of 3 Ecuador v. Mackay, 742 F.3d 860, 869 n.3 (9th Cir. 2014)), even opinion work product is not 4 totally immune from discovery when an attorneys’ mental impressions are at issue in a case and 5 the need for the material is compelling. Holmgren v. State Farm Mut. Auto. Ins. Co., 976 F.2d 6 573, 577 (9th Cir. 1992). Thus, “[t]he privilege derived from the work product doctrine is not 7 absolute[] [and] [l]ike other qualified privileges, it may be waived.” United States v. Nobles, 422 8 U.S. 225, 236 (1975). Similar to the waiver of the attorney-client privilege, a litigant can waive 9 work product protection to the extent that the litigant reveals, or places the work product at issue, 10 during the course of litigation. United States v. Sanmina Corp. & Subsidiaries, 968 F.3d 1107, 11 1119 (9th Cir. 2020). 12 For instance, in Nobles, 422 U.S. at 237-38, a criminal defendant’s decision to present his 13 defense investigator as a witness waived work-product privilege over the investigator’s report 14 “with respect to matters covered in his testimony.” Id. at 236, see also id. at 239 (“[R]espondent 15 sought to adduce the testimony of the investigator and contrast his recollection of the contested 16 statements with that of the prosecution’s witnesses. Respondent, by electing to present the 17 investigator as a witness, waived the privilege with respect to matters covered in his testimony. 18 Respondent can no more advance the work-product doctrine to sustain a unilateral testimonial use 19 of work-product materials than he could elect to testify in his own behalf and thereafter assert his 20 Fifth Amendment privilege to resist cross-examination on matters reasonably related to those 21 brought out in direct examination.”). 22 The court will first note that the account settings and the negative test results are more in 23 the nature of fact work product than opinion work product. While Plaintiffs suggest that revealing 24 these facts would necessarily reveal “the thoughts and analysis conducted by Plaintiffs’ lawyers in 25 preparation of the litigation” (see Ltr. Br. (dkt. 153) at 3), the court finds this suggestion to be 26 exaggerated and unpersuasive. As stated above, opinion work product consists of the attorney’s 27 interpretation of legal theories and the application of facts to those theories, rather than the bare 1 the nature of bare facts. However, even assuming arguendo that to reveal these bare facts would 2 provide a revelatory glimpse into Plaintiffs’ counsels’ thought process such as to reveal counsels’ 3 interpretation of legal theories and the application of some facts to those theories, Plaintiffs cannot 4 avoid the notion that by placing a large subset of these facts in the FAC and in Exhibit-B, 5 Plaintiffs have waived the ability to assert work product protection. As was the case with the 6 investigator’s report in Nobles, 422 U.S. at 237-38 – here too, Plaintiffs’ negative results and the 7 account settings in question are “reasonably related” to the positive results that Plaintiffs have 8 placed at issue and included in the FAC and in Exhibit-B. Thus, an evaluation of those account 9 settings and negative test results is a necessary component of Defendants’ ability to understand 10 Plaintiffs’ positive results and to fairly subject those results to meaningful scrutiny. 11 It is therefore no answer for Plaintiffs to simply declare that the negative results are “not 12 relevant” simply because Plaintiffs did not include them in the FAC. As stated, Defendants’ 13 review of the negative test results, as well as the account setting used to interrogate ChatGPT, 14 appears likely to aid in Defendants’ understanding of Plaintiffs’ positive results as well as 15 Defendants’ ability to subject those positive test results to scrutiny. Or, to put it another way, the 16 court disagrees with Plaintiffs’ contention that the account settings and the prompts and outputs 17 that do not “support Plaintiffs’ claims” (see Ltr. Br. (dkt. 153) at 3) are not “particularly relevant” 18 (see id. at 2) because while they may not be relevant to Plaintiffs’ pursuit of their claims, they are 19 plainly relevant to Defendants’ defenses. Of course, the reason is that Plaintiffs have placed a 20 subset of their test results at issue in this case; and, having done so, Defendants should be entitled 21 to examine the entirety of Plaintiffs’ test results (including the negative results), as well as the 22 account settings used in interrogating ChatGPT. 23 Lastly, the court is unpersuaded by Plaintiffs’ argument that Defendants can simply 24 interrogate ChatGPT themselves. See id. at 5. Without knowing the account settings used by 25 Plaintiffs to generate their positive and negative results, and without knowing the exact 26 formulation of the prompts used to generate Plaintiffs’ negative results, Defendants would be 27 unable to replicate the same results. See e.g., Mushroom Assoc. v. Monterey Mushrooms, Inc., 1 cannot obtain substantially equivalent materials because the specific work product is directly at 2 || issue. For instance, although the discovering party could obtain its own test results concerning a 3 patent, the tests results known to counsel at the time advice is given are what is relevant when a 4 || party asserts the advice of counsel defense.”). Here, Plaintiffs’ FAC and Exhibit-B contain a series 5 of test results, and while Defendants can certainly produce some prompts and responses 6 || themselves, doing so will not be useful in understanding or scrutinizing Plaintiffs’ test results. On 7 the other hand, discovery into the account settings used by Plaintiffs, and discovery into Plaintiffs’ 8 || negative test results will achieve that end. 9 For the reasons stated herein, as well as those argued by Defendants (see Ltr. Br. (dkt. 153) 10 at 1-3), the request to compel the OpenAI account information for the individuals who used 11 ChatGPT to investigate Plaintiffs’ claims, and the prompts and outputs for Plaintiffs’ testing of 12 || ChatGPT in connection with their pre-suit ChatGPT testing, including prompts and outputs that 5 13 || did not reproduce or summarize Plaintiffs’ works or otherwise support Plaintiffs’ claims, along 14 || with documentation of Plaintiffs’ testing process, is GRANTED. 3 15 IT IS SO ORDERED. 16 || Dated: June 24, 2024 Hh
ROBERT M. ILLMAN 19 United States Magistrate Judge 20 21 22 23 24 25 26 27 28