Search and Browse – ELRA Catalogue

2006 CoNLL Shared Task – Arabic & Czech text

Arabic
Czech

ID: ELRA-W0087

2006 CoNLL Shared Task – Arabic & Czech consists of dependency treebanks used as part of the CoNLL 2006 shared task on multi-lingual dependency parsing. The Conference on Computational Natural Language Learning (CoNLL) is accompanied every year by a shared task intended to promote natural lan...

MEMBER	academic	commercial
Licence: Non Commercial Use - Non Standard Licence Terms

NON MEMBER	academic	commercial
Licence: Non Commercial Use - Non Standard Licence Terms

2007 CoNLL Shared Task - Arabic & English text

Arabic
English

ID: ELRA-W0123

ISLRN: 505-782-255-628-8

2007 CoNLL Shared Task - Arabic & English consists of dependency treebanks in two languages used as part of the CoNLL 2007 shared task on multi-lingual dependency parsing and domain adaptation. The languages covered in this release are: Arabic and English. The Conference on Computational Natur...

MEMBER	academic	commercial
Licence: Non Commercial Use - Non Standard Licence Terms

NON MEMBER	academic	commercial
Licence: Non Commercial Use - Non Standard Licence Terms

Amharic-English bilingual corpus text

Amharic
English

ID: ELRA-W0074

ISLRN: 590-255-335-719-0

The Amharic-English bilingual corpus contains parallel text from legal and news domains in Amharic script, in transliterated form and in English. The size of the corpus is of 232,653 words in Amharic and 291,701 in English. This parallel corpus contains documents from two domains, namely legal...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	2000.00 €
Licence: Commercial Use - ELRA VAR	2000.00 €	2000.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	4000.00 €
Licence: Commercial Use - ELRA VAR	4000.00 €	4000.00 €

Bilingual (Spanish-English) Speech synthesis HTS models audio

English
Spanish; Castilian

ID: ELRA-S0335

ISLRN: 277-380-359-561-3

This database contains Bilingual (English and Spanish) Festival HTS models. Models were trained with 9h of speech from 2 female bilingual speakers and 2 male bilingual speakers. Each speaker recorded 2h 15 min per language. The speech data can be found in the TC-STAR Bilingual Voice-Conversion S...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

Chinese-Vietnamese Parallel Corpus text

Chinese
Vietnamese

ID: ELRA-W0312

ISLRN: 128-772-037-486-0

The Chinese-Vietnamese Parallel Corpus consists of 200,000 sentence pairs, with an average length of 15 words per sentence. The corpus is provided in XML format and is annotated according to TEI-encoding guidelines.

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	200.00 €	400.00 €
Licence: Commercial Use - ELRA VAR	1400.00 €	1400.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	300.00 €	600.00 €
Licence: Commercial Use - ELRA VAR	2100.00 €	2100.00 €

Chinese-Vietnamese - PhraseBank with audio files audio

Chinese
Vietnamese

ID: ELRA-S0485

ISLRN: 428-557-564-826-7

Chinese-Vietnamese - PhraseBank with audio files of daily conversations spoken by native speakers containing 4002 sentence pairs. Scripts with Pinyin, Topic, Cat, Vietnamese translation with corresponding audio in Chinese and Vietnamese. Corpus in XML and WAV formats.

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	400.00 €	500.00 €
Licence: Commercial Use - ELRA VAR	900.00 €	900.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	600.00 €	750.00 €
Licence: Commercial Use - ELRA VAR	1350.00 €	1350.00 €

DA-EN Danish Ministry of Higher Education and Science 2 (Processed) text

Danish
English

ID: ELRA-W0157

ISLRN: 026-863-463-067-1

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts Danish-English from the Danish Ministry o...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

DA-EN Danish Ministry of Higher Education and Science 3 (Processed) text

Danish
English

ID: ELRA-W0155

ISLRN: 625-397-811-990-4

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts Danish-English from the Danish Ministry o...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

DA-EN Danish Ministry of Higher Education and Science 4 (Processed) text

Danish
English

ID: ELRA-W0172

ISLRN: 560-401-490-272-1

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts Danish-English from the Danish Ministry o...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

DA-EN Danish Ministry of Higher Education and Science (Processed) text

Danish
English

ID: ELRA-W0166

ISLRN: 222-781-852-505-9

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts Danish-English from the Danish Ministry o...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

ECPC Corpus (European Comparable and Parallel Corpora of Parliamentary Speeches Archive) – set 1 text

English
Spanish; Castilian

ID: ELRA-W0128

ISLRN: 036-939-425-010-1

The European Comparable and Parallel Corpora of Parliamentary Speeches Archive (ECPC), compiled at the Universitat Jaume I (Spain), is a collection of XML metatextually tagged corpora containing speeches from three European chambers (the European Parliament, the British House of Commons, and the ...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use, Share Alike - CC-BY-NC-SA	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use, Share Alike - CC-BY-NC-SA	0.00 €	0.00 €

Ema-lon Manipuri Corpus (including word embedding and language model) text

English
Manipuri

ID: ELRA-W0316

ISLRN: 588-170-827-016-7

The Ema-lon Manipuri Corpus consists of a set of resources for Manipuri language (locally known as Meiteilon) for the purpose of machine translation. The main source for these resources is the Sangai Express news website. The resources that constitute the present corpus are listed below: 1. EM C...

MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Attribution, Non Commercial Use - CC-BY-NC-4.0	0.00 €	0.00 €

English-Nepali Parallel Corpus text

English
Nepali (macrolanguage)

ID: ELRA-W0077

ISLRN: 853-487-663-161-6

The Nepali Monolingual written corpus is one of the 3 resources that constitute the Nepali National Corpus. The Nepali National Corpus was produced in 2006 in the framework of the project Bhasha Sanchar (“language communication”), also known as Nelralec, for Nepali Language Resources and Localiza...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €

English-Persian parallel Corpus text

English
Persian

ID: ELRA-W0051

ISLRN: 671-618-321-687-7

Please refer to ELRA-W0118 for the latest version of this corpus. This version consists of about 3,500,000 English and Persian (Farsi) words aligned at sentence level (about 100,000 sentences, distributed over 50,021 entries). The format of the files is Unicode. It has been originally created wi...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	500.00 €	2500.00 €
Licence: Commercial Use - ELRA VAR	2500.00 €	2500.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	600.00 €	3000.00 €
Licence: Commercial Use - ELRA VAR	3000.00 €	3000.00 €

English-Persian parallel corpus text

English
Persian

ID: ELRA-W0118

ISLRN: 074-825-114-781-7

The English-Persian parallel corpus contains more than 200,000 aligned sentences across a variety of text types from the domains of art, law, culture, science, religion, literature, medicine, idioms, politics and others. It is an extension of the English-Persian parallel corpus already distribute...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	1000.00 €	5000.00 €
Licence: Commercial Use - ELRA VAR	5000.00 €	5000.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	1200.00 €	6000.00 €
Licence: Commercial Use - ELRA VAR	6000.00 €	6000.00 €

English-Punjabi Code-Mixed Social Media Content text

English
Panjabi; Punjabi

ID: ELRA-W0319

ISLRN: 695-759-706-170-8

The English-Punjabi Code-Mixed Social Media Content corpus is composed is composed of 893,615 parallel sentences of English-Punjabi distributed over the following domains: - 82,341 parallel sentences of English-Punjabi code-mixed Agriculture Domain Data, - 59,158 parallel sentences of English-P...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

English-Vietnamese Parallel Corpus text

English
Vietnamese

ID: ELRA-W0124

ISLRN: 838-483-738-912-8

This is a corpus of 500,000 English-Vietnamese sentence pairs, built to develop SMT (Statistical Machine Translation) systems. The parallel corpus contains English documents translated by professional translators into Vietnamese. The source texts include books, dictionaries, newspapers, online ne...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	600.00 €	1200.00 €
Licence: Commercial Use - ELRA VAR	6000.00 €	6000.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	1000.00 €	2000.00 €
Licence: Commercial Use - ELRA VAR	8000.00 €	8000.00 €

English-Vietnamese Parallel Corpus text

English
Vietnamese

ID: ELRA-W0311

ISLRN: 893-470-491-825-6

The English-Vietnamese Parallel Corpus consists of 1,000,000 sentence pairs, with an average length of 20 words per sentence. The corpus is provided in XML format and is annotated according to TEI-encoding guidelines.

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	900.00 €	1800.00 €
Licence: Commercial Use - ELRA VAR	9000.00 €	9000.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	1500.00 €	3000.00 €
Licence: Commercial Use - ELRA VAR	12000.00 €	12000.00 €

EUROPARL Corpus Parallel Corpora: Portuguese-English text

English
Portuguese

ID: ELRA-W0090

ISLRN: 435-502-922-727-2

The EUROPARL Corpus (Portuguese-English subpart of the parallel corpora), was extracted from the proceedings of the European Parliament. It contains transcriptions of sessions dating back from 1996 to 2011, with a total of approximately 58,324,562 tokens of European Portuguese (L1) and 49,216,896...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €
Licence: Commercial Use - ELRA VAR	0.00 €	0.00 €

GeFRePaC - German French Reciprocal Parallel Corpus text

French
German

ID: ELRA-W0031

ISLRN: 086-761-267-762-3

The German-French Reciprocal Parallel Corpus (GeFRePaC) was produced by the Multilinguale Forschung/Multilingual Research Abteilung Lexik, Institut für Deutsche Sprache (Germany) through a funding from ELRA in the framework of the European Commission project LRsP&P (Language Resources Production ...

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	0.00 €

Corpus:
Lexical/Conceptual:
Tool/Service:
Language Description:

Text:
Audio:
Image:
Video:
Text Numerical:
Text N-Gram:

Resource Type:

Media Type:

55 Language Resources (Page 1 of 3)