20 of 3278 resources
Textual collection
Arabic news corpus (United Kingdom) based on material crawled in 2018 created in the project Deutscher Wortschatz or Leipzig Corpora Collection. The project regularly collects and processes available documents from the Internet (typically in an annual cycle) and other sources. The results are corpora and corpora-based dictionaries for more than 250 languages, which provide statistical information about almost each word, example sentences and links to related words. Because of the huge amount of used text material containing several million sentences, information about almost every word can be provided. The service ranks among the most comprehensive information systems about the German language and provides the largest freely available amounts of data for many other languages. For copyright reasons, the data are provided as derived text formats that do not allow reconstruction of the original document structures.
Textual collection
Braunschweiger Zeitung 2012 ist Teil des Deutschen Referenzkorpus DeReKo. Die Korpora geschriebener Gegenwartssprache des IDS bilden die weltweit größte linguistisch motivierte Sammlung elektronischer Korpora mit geschriebenen deutschsprachigen Texten aus der Gegenwart und der neueren Vergangenheit. Sie enthalten belletristische, wissenschaftliche und populärwissenschaftliche Texte, eine große Zahl von Zeitungstexten sowie eine breite Palette weiterer Textarten und werden kontinuierlich weiterentwickelt. Aktueller Stand: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/archiv-1/ Abrufbar über KorAP: https://korap.ids-mannheim.de/ Abrufbar über Cosmas II: https://cosmas2.ids-mannheim.de/cosmas2-web/ Weitere Informationen: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/
Textual collection
The MCScript corpus is a large dataset of narrative texts and questions about these texts, intended to be used in a machine comprehension task that requires reasoning using commonsense knowledge. Our dataset complements similar datasets in that we focus on stories about everyday activities, such as going to the movies or working in the garden, and that the questions require commonsense knowledge, or more specifically, script knowledge, to be answered. We show that our mode of data collection via crowdsourcing results in a substantial amount of such inference questions. The dataset forms the basis of a shared task on commonsense and script knowledge organized at SemEval 2018 and provides challenging test cases for the broader natural language understanding community.
Textual collection
The ressource contains a semantic analysis of Spanish nonce-formations derived with potentially collective suffixes (based on a list of hapax legomena in esTenTen11). Metadata: List_hapaxlegomena_esTenTen11.xml Ressource: Liste_coll_HL_SP.xlsx
Textual collection
Gemeinfreie Digitalisate der Exilsammlungen
Textual collection
Dutch news subcorpus based on material from 2012 (300,000 sentences) created in the project Deutscher Wortschatz or Leipzig Corpora Collection. The project regularly collects and processes available documents from the Internet (typically in an annual cycle) and other sources. The results are corpora and corpora-based dictionaries for more than 250 languages, which provide statistical information about almost each word, example sentences and links to related words. Because of the huge amount of used text material containing several million sentences, information about almost every word can be provided. The service ranks among the most comprehensive information systems about the German language and provides the largest freely available amounts of data for many other languages. For copyright reasons, the data are provided as derived text formats that do not allow reconstruction of the original document structures.
Textual collection
Niederösterreichische Nachrichten 2007 ist Teil des Deutschen Referenzkorpus DeReKo. Die Korpora geschriebener Gegenwartssprache des IDS bilden die weltweit größte linguistisch motivierte Sammlung elektronischer Korpora mit geschriebenen deutschsprachigen Texten aus der Gegenwart und der neueren Vergangenheit. Sie enthalten belletristische, wissenschaftliche und populärwissenschaftliche Texte, eine große Zahl von Zeitungstexten sowie eine breite Palette weiterer Textarten und werden kontinuierlich weiterentwickelt. Aktueller Stand: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/archiv-1/ Abrufbar über KorAP: https://korap.ids-mannheim.de/ Abrufbar über Cosmas II: https://cosmas2.ids-mannheim.de/cosmas2-web/ Weitere Informationen: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/
Textual collection
The BKS Corpus consists of three subcorpora: (a) Comic Corpus, (b) Bosnian Interviews, (c) Novosadski Corpus of Spoken Language. The research interest of the SFB 441 project B8 lies in the use of the Bosnian/Croatian/Serbian v/t/n-deictics in different text classes.
Textual collection
Arabic news corpus (People’s Republic of China) based on material crawled in 2018 created in the project Deutscher Wortschatz or Leipzig Corpora Collection. The project regularly collects and processes available documents from the Internet (typically in an annual cycle) and other sources. The results are corpora and corpora-based dictionaries for more than 250 languages, which provide statistical information about almost each word, example sentences and links to related words. Because of the huge amount of used text material containing several million sentences, information about almost every word can be provided. The service ranks among the most comprehensive information systems about the German language and provides the largest freely available amounts of data for many other languages. For copyright reasons, the data are provided as derived text formats that do not allow reconstruction of the original document structures.
Textual collection
This experiment tests potential Principle C violations across clauses.
Textual collection
80.500 Artikel der Periodika „Das Andere Deutschland / La Otra Alemania“, Buenos Aires, 1939 - 1949, und „Aufbau / Reconstruction“, New York, 1934 – 1950
Textual collection
Bosnian Wikipedia subcorpus based on material from 2018 (300,000 sentences) created in the project Deutscher Wortschatz or Leipzig Corpora Collection. The project regularly collects and processes available documents from the Internet (typically in an annual cycle) and other sources. The results are corpora and corpora-based dictionaries for more than 250 languages, which provide statistical information about almost each word, example sentences and links to related words. Because of the huge amount of used text material containing several million sentences, information about almost every word can be provided. The service ranks among the most comprehensive information systems about the German language and provides the largest freely available amounts of data for many other languages. For copyright reasons, the data are provided as derived text formats that do not allow reconstruction of the original document structures.
Textual collection
Czech news subcorpus based on material from 2012 (1,000,000 sentences) created in the project Deutscher Wortschatz or Leipzig Corpora Collection. The project regularly collects and processes available documents from the Internet (typically in an annual cycle) and other sources. The results are corpora and corpora-based dictionaries for more than 250 languages, which provide statistical information about almost each word, example sentences and links to related words. Because of the huge amount of used text material containing several million sentences, information about almost every word can be provided. The service ranks among the most comprehensive information systems about the German language and provides the largest freely available amounts of data for many other languages. For copyright reasons, the data are provided as derived text formats that do not allow reconstruction of the original document structures.
Textual collection
This data set comprises adjusted data from a rating study in which participants read single sentences on the screen and rate their acceptability on a scale from 1-7.
Textual collection
Im Forschungsprojekt "Der deutsche Sprachraum aus der Sicht linguistischer Laien" wurden die laienlinguistischen Konzeptualisierungen zur deutschen Sprache untersucht. Wesentliche Ziele dieses Forschungsprojektes sind die Grundlagenforschung zur laienbezogenen Sprachkonzeption und die empirische Erhebung von Wissensbeständen linguistischer Laien. Dazu wurden in Deutschland, Österreich, der deutschsprachigen Schweiz, Liechtenstein, Luxemburg, Ostbelgien und Südtirol jeweils sechs Personen dreier Altersgruppen befragt. Momentan steht ein Korpus von sprachlichen und metasprachlichen Daten von 139 Gewährspersonen zur Verfügung. Die produzierten mental maps sind durch eine WebGIS-Anwendung abrufbar. Die im August 2017 erschienene Abschlusspublikation mit ausgewählten Ergebnissen beschließt das DFG-Projekt. Hundt, Markus; Palliwoda, Nicole; Schröder, Saskia (Hrsg.) (2017) Der deutsche Sprachraum aus der Sicht linguistischer Laien. Ergebnisse des Kieler DFG-Projektes. De Gruyter. DOI: https://doi.org/10.1515/9783110554212 Rezension: Stoltmann, Kai (2018): Rezension. In: Zeitschrift für Dialektologie und Linguistik (ZDL), 85/3. S. 365-368. Homepage: Wahrnehmungsdialektologie
Textual collection
Die Zeit (Online-Ausgabe), 2008 ist Teil des Deutschen Referenzkorpus DeReKo. Die Korpora geschriebener Gegenwartssprache des IDS bilden die weltweit größte linguistisch motivierte Sammlung elektronischer Korpora mit geschriebenen deutschsprachigen Texten aus der Gegenwart und der neueren Vergangenheit. Sie enthalten belletristische, wissenschaftliche und populärwissenschaftliche Texte, eine große Zahl von Zeitungstexten sowie eine breite Palette weiterer Textarten und werden kontinuierlich weiterentwickelt. Aktueller Stand: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/archiv-1/ Abrufbar über KorAP: https://korap.ids-mannheim.de/ Abrufbar über Cosmas II: https://cosmas2.ids-mannheim.de/cosmas2-web/ Weitere Informationen: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/
Textual collection
The dialogues of the Spanish DollDialogue study were elicited in a referential communication task in which it was likely that participants would act jointly to negotiate spatial relations. The dialogue partners exchanged information about furniture arrangements in a doll's house in order to furnish an empty doll's house in the same way. The verbal data and photos of the resulting furnishings mapped against the model arrangement allow to trace sequences of miscommunication.
Textual collection
Die Zeit (Online-Ausgabe), 2007 ist Teil des Deutschen Referenzkorpus DeReKo. Die Korpora geschriebener Gegenwartssprache des IDS bilden die weltweit größte linguistisch motivierte Sammlung elektronischer Korpora mit geschriebenen deutschsprachigen Texten aus der Gegenwart und der neueren Vergangenheit. Sie enthalten belletristische, wissenschaftliche und populärwissenschaftliche Texte, eine große Zahl von Zeitungstexten sowie eine breite Palette weiterer Textarten und werden kontinuierlich weiterentwickelt. Aktueller Stand: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/archiv-1/ Abrufbar über KorAP: https://korap.ids-mannheim.de/ Abrufbar über Cosmas II: https://cosmas2.ids-mannheim.de/cosmas2-web/ Weitere Informationen: https://www.ids-mannheim.de/digspra/kl/projekte/korpora/