LT Corpus – ELRA Catalogue

Last view: 2024-06-20

374 Last view: 2024-06-20

Last update: 2020-03-06

3 Last update: 2020-03-06

LT Corpus

View resource name in all available languages

Corpus LT

ISLRN: 569-208-468-863-2

ID:

ELRA-W0059

The LT Corpus is composed of 70 fiction texts from Portuguese renowned authors. The corpus contains 1,781,083 tokens. The texts date from before 1940.

The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The LT Corpus is a copyright free subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120.
The Corpus is delivered with the annotation manual of the CRPC, a metadata file and a narrative description of the resource.

View resource description in French

Le corpus LT est constitué de 70 textes de fiction d’auteurs portugais renommés. Le corpus contient 1 781 083 tokens. Les textes datent d’avant 1940.

Le corpus est livré en un seul fichier, dans deux formats différents. La version txt contient une phrase par ligne, un numéro d’identification pour chaque texte et aucune autre annotation. Les fichier cqpweb contient un token par ligne, suivi par une étiquette de partie du discours et le lemme correspondant, avec l’annotation des chunks NP. Le corpus LT est un sous-ensemble libre de copyright du Corpus de référence du portugais contemporain et suit le même schéma d’annotation (cf.: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120).
Le corpus est fourni avec le manuel d’annotation du corpus de référence, un fichier de meta-données et une description narrative de la ressource.

MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	2500.00 €
Licence: Commercial Use - ELRA VAR	2500.00 €	2500.00 €

NON MEMBER	academic	commercial
Licence: Non Commercial Use - ELRA END USER	0.00 €	3000.00 €
Licence: Commercial Use - ELRA VAR	3000.00 €	3000.00 €

DistributionAvailability start date 05/12/2012 Contact Person

Valérie Mapelli

text

Monolingual text corpusLanguages

Portuguese

Linguality

Linguality type: Monolingual

Size

no size available

Metadata

Created: 05/12/2005

Metadata Language: French, English (fr, en)

Version

Version: 1.0

Last Updated: 12/05/2012