Dictionary: cmu-pronouncing-dictionary (https://github.com/words/cmu-pronouncing-dictionary) Upstream data: https://github.com/cmusphinx/cmudict Exact package version, entry count and generated SHA-256: dictionary-provenance.json ISC License Copyright (c) 2015 Zeke Sikelianos Permission to use, copy, modify, and/or distribute this software for any purpose with or without fee is hereby granted, provided that the above copyright notice and this permission notice appear in all copies. THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE. CMU upstream licence: Copyright (C) 1993-2015 Carnegie Mellon University. All rights reserved. Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met: 1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. The contents of this file are deemed to be source code. 2. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution. This work was supported in part by funding from the Defense Advanced Research Projects Agency, the Office of Naval Research and the National Science Foundation of the United States of America, and by member companies of the Carnegie Mellon Sphinx Speech Consortium. We acknowledge the contributions of many volunteers to the expansion and improvement of this dictionary. THIS SOFTWARE IS PROVIDED BY CARNEGIE MELLON UNIVERSITY ``AS IS'' AND ANY EXPRESSED OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL CARNEGIE MELLON UNIVERSITY NOR ITS EMPLOYEES BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. English usage frequencies: wordfreq 3.1.1, Copyright 2022 Robyn Speer. The frequencies field in pack-en.json is adapted from wordfreq English data, restricted to this dictionary. That adapted field is licensed CC BY-SA 4.0: https://creativecommons.org/licenses/by-sa/4.0/ Source: https://github.com/rspeer/wordfreq Rebuild: scripts/build_web_frequencies.py (wordfreq==3.1.1). No wordfreq application code is bundled into the website. Upstream attribution notice follows: wordfreq Copyright 2022 Robyn Speer # Attribution notes Robyn Speer must be credited as Robyn Speer, which is her maiden name, used on academic work. Crediting her as Elia Robyn Lake (her married name) will make the credit less effective, as it will not line up with other work. Crediting Robyn Speer by a different name than one of the above is a serious violation of the license, in which case you do not have permission to use, copy, or redistribute wordfreq. If you use wordfreq in academic work, you must cite it. See "Citing wordfreq" in README.md. # Included licenses `wordfreq` is freely redistributable under the Apache license (see `LICENSE.txt`), and it includes data files that may be redistributed under a Creative Commons Attribution-ShareAlike 4.0 license (). `wordfreq` contains data extracted from Google Books Ngrams () and Google Books Syntactic Ngrams (). The terms of use of this data are: Ngram Viewer graphs and data may be freely used for any purpose, although acknowledgement of Google Books Ngram Viewer as the source, and inclusion of a link to http://books.google.com/ngrams, would be appreciated. `wordfreq` also contains data derived from the following Creative Commons-licensed sources: - The Leeds Internet Corpus, from the University of Leeds Centre for Translation Studies () - Wikipedia, the free encyclopedia () - ParaCrawl, a multilingual Web crawl () It contains data from OPUS OpenSubtitles 2018 (), whose data originates from the OpenSubtitles project () and may be used with attribution to OpenSubtitles. It contains data from various SUBTLEX word lists: SUBTLEX-US, SUBTLEX-UK, SUBTLEX-CH, SUBTLEX-DE, and SUBTLEX-NL, created by Marc Brysbaert et al. (see citations below) and available at . I (Robyn Speer) have obtained permission by e-mail from Marc Brysbaert to distribute these wordlists in wordfreq, to be used for any purpose, not just for academic use, under these conditions: - Wordfreq and code derived from it must credit the SUBTLEX authors. - It must remain clear that SUBTLEX is freely available data. These terms are similar to the Creative Commons Attribution-ShareAlike license. Some additional data was collected by a custom application that watched the streaming Twitter API, in accordance with Twitter's Developer Agreement & Policy. This software gives statistics about words that were commonly used on Twitter; it does not display or republish any Twitter content, and does not contain any content from after Twitter's sale. ---