Nuance Communications Inc. v. Omilia Natural Language Solutions, Ltd.

District Court, D. Massachusetts·Decided August 6, 2020·No. 1:19-cv-11438·Unknown

Opinion

UNITED STATES DISTRICT COURT DISTRICT OF MASSACHUSETTS ________________________________________ ) NUANCE COMMUNICATIONS, INC., ) ) Plaintiff, ) ) v. ) Civil Action ) No. 19-11438-PBS OMILIA NATURAL LANGUAGE SOLUTIONS, ) LTD., ) ) Defendant. ) ________________________________________)

MEMORANDUM AND ORDER ON CLAIM CONSTRUCTION August 6, 2020 Saris, D.J. Nuance Communications, Inc. accuses Omilia Natural Language Solutions, Ltd., of infringing U.S. Patent No. 6,999,925 (the “‘925 patent”). The parties dispute the claim construction of four terms: “a second language,” “multi-lingual speech recognizer,” “generating a second acoustic model,” and “automatically generate/ing.” The Court held a non-evidentiary Markman hearing on July 10, 2020 and reviewed technical tutorials and briefs submitted by both parties. BACKGROUND The ʼ925 patent, entitled “Method and Apparatus for Phonetic Context Adaptation for Improved Speech Recognition,” describes a method to “automatically generat[e] from a first speech recognizer a second speech recognizer which can be adapted to a specific domain.” ‘925 patent, col. 1, ll. 15-19 (Dkt. 84-1). The patented method “provide[s] for fast and easy customization of speech recognizers to a given domain,” such as “a certain language, a multitude of languages, a dialect or a set of dialects, [or] a certain task area” like medicine or banking. See id. at col. 2, ll. 21-23, col. 6, ll. 5-8. Speech

recognizers generally match audio input of human speech to phonemes, or abstractions of individual sounds. To do so, recognizers determine the probability that an audio file contains a certain sound, or phone, using the phone’s “phonetic context.” “Phonetic context” refers to sub-word units consisting of individual phones with neighboring phones. See ‘id. at col. 4, ll. 35-38. “[T]he collection of a large amount of training data and the subsequent training of a speech recognizer is both expensive and time consuming.” Id. at col. 1, ll. 61-63. A recognizer is trained using “labelled training data.” See id. at col. 1, ll. 54-55. The recognizer uses the training data to “construct[] a

binary decision network” with a “predefined number of leaves,” which are also known as “terminal nodes.” See id. at col. 4, ll. 44, 61, col. 5, l. 6. Several methods for “the customization of a general speech recognizer to a particular domain” predate the ‘925 patent. See id. at col. 5, ll. 28-67. These methods included “[1] ‘selecting’ a subset of the general speech recognizer decision network and the phonetic contexts or [2] simply ‘enhancing’ the decision network by . . . attaching a new sub-tree with new leaf nodes and further phonetic contexts.” Id. at col. 7, ll. 12-17. The ‘925 patent describes the first approach as sub-optimal because it could not “detect any new phonetic context . . .

relevant to a new domain but not present in the general recognizer’s inventory.” Id. at col. 5, ll. 59-61. The patent describes the second approach as sub-optimal because it “still require[d] the collection of a substantial amount” of training data. Id. at col. 5, ll. 50-51; see also id. at col. 1, l. 66- col. 2, l. 3 (describing a third “[c]onventional adaptation method[]” of “simply provid[ing] a modification of the acoustic model parameters”).. In contrast, the ‘925 patent method requires only a “small amount of domain specific adaptation data” to adapt a speech recognizer to a particular domain. Id. at col. 6, ll. 11-16. The patent “preserves the phonetic context information of the first

speech recognizer” and “simultaneously allows for the creation of new phonetic contexts.” Id. at col. 2, ll. 45-50. “This is achieved by . . . re-estimating the decision network and phonetic contexts based on domain-specific training data.” Id. at col. 6, ll. 16-20. Figure 1 reflects a preferred embodiment of the '925 patent. The specification makes clear “that the invention is not to limited to the precise arrangements and instrumentalities shown” in Figure 1. Id. at col. 2, 11. 63-64.

speaker independent, general purpase SPEECH RECOGNIZER (1)

(UNJSUPERVISED COLLECTION OF APPLICATION specific training speech (2)

phonetic context extraction (31)

context classification of application specific training speech (32)

damain dependent phonetic context detection

computation and/or adaptation of domain specific HMM parameters (5)

application specific requirements fulfilled?

YES low resource, application specific speech recognizer (6) FIGURE 41

The patent describes the embodiment in Figure 1 as follows. First, the adaptation data is “passed through the original decision network” to “obtain a partitioning of the adaptation data.” Id. at col. 7, ll. 52-59. The system then “insert[s], delet[es], or adapt[s]” phonetic contexts, “resulting in a new, re-estimated (domain specific) decision network.” Id. at col. 7,

ll. 8-12, 62-65. This method can “significantly improve the recognition rate within a given target domain” while “avoid[ing] an unacceptable decrease of recognition accuracy in the original recognizer’s domain.” Id. at col. 10, ll. 13-18. The patent provides that this method “can be used for the incremental and data driven incorporation of a new language into a true multi-lingual speech recognizer.” Id. at col. 9, ll. 20-22. Independent claims 1, 2, 14, 15, and 27 and dependent claims 12 and 25 of the ʼ925 patent use the disputed terms at issue here. The claims are as follows: 1. A computerized method of automatically generating from a first speech recognizer a second speech recognizer, said first speech recognizer comprising a first acoustic model with a first decision network and corresponding first phonetic contexts, and said second speech recognizer being adapted to a specific domain, said method comprising: based on said first acoustic model, generating a second acoustic model with a second decision network and corresponding second phonetic contexts for said second speech recognizer by re-estimating said first decision network and said corresponding first phonetic contexts based on domain-specific training data, wherein said first decision network and said second decision network utilize a phonetic decision [t]ree to perform speech recognition operations, wherein the number of nodes in the second decision network is not fixed by the number of nodes in the first decision network, and wherein said re-estimating comprises partitioning said training data using said first decision network of said first speech recognizer.

2. A computerized method of automatically generating from a first speech recognizer a second speech recognizer, said first speech recognizer comprising a first acoustic model wit[h] a first decision network and corresponding first phonetic contexts, and said second speech recognizer being adapted to a specific domain, said method comprising: based on said first acoustic model, generating a second acoustic model with a second decision network and corresponding second phonetic contexts for said second speech recognizer by re-estimating said first decision network and said corresponding first phonetic contexts based on domain-specific training data, wherein said first decision network and said second decision network utilize a phonetic decision tree to perform speech recognition operations, wherein the number of nodes in the second decision network is not fixed by the number of nodes in the first decision network, wherein said domain-specific training data is of a limited amount, and wherein the generating step further comprises the steps of: identifying at least one acoustic context from the domain-specific training data; and adding a node to the second decision network for the identified context independent of other generating step operations.

Free access — add to your briefcase to read the full text and ask questions with AI

Nuance Communications Inc. v. Omilia Natural Language Solutions, Ltd., (D. Mass. 2020).

Nuance Communications Inc. v. Omilia Natural Language Solutions, Ltd. (Nuance Communications Inc. v. Omilia Natural Language Solutions, Ltd.) — published by Counsel Stack Legal Research, free access to 12M+ legal documents.

Related

PODS, Inc. v. Porta Stor, Inc.
484 F.3d 1359 (Federal Circuit, 2007)
Collegenet, Inc. v. Applyyourself, Inc.
418 F.3d 1225 (Federal Circuit, 2005)
Thorner v. Sony Computer Entertainment America LLC
669 F.3d 1362 (Federal Circuit, 2012)
Fin Control Systems Pty, Ltd. v. Oam, Inc.
265 F.3d 1311 (Federal Circuit, 2001)
Whitserve, LLC v. Computer Packages, Inc.
694 F.3d 10 (Federal Circuit, 2012)
3m Innovative Properties v. Tredegar Corporation
725 F.3d 1315 (Federal Circuit, 2013)
Hill-Rom Services, Inc. v. Stryker Corporation
755 F.3d 1367 (Federal Circuit, 2014)
Golden Bridge Technology, Inc. v. Apple Inc.
758 F.3d 1362 (Federal Circuit, 2014)
Akzo Nobel Coatings, Inc. v. Dow Chemical Company
811 F.3d 1334 (Federal Circuit, 2016)
Trustees of Columbia Univ. v. Symantec Corporation
811 F.3d 1359 (Federal Circuit, 2016)
Poly-America, L.P. v. Api Industries, Inc.
839 F.3d 1131 (Federal Circuit, 2016)
Aylus Networks, Inc. v. Apple Inc.
856 F.3d 1353 (Federal Circuit, 2017)
Continental Circuits LLC v. Intel Corporation
915 F.3d 788 (Federal Circuit, 2019)