{# Audit 04/10/2026 : « autre » n'est pas un code de langue ; SPHAERO n'est pas l'éditeur des documents qu'elle héberge ou référence. #} {# citation_pdf_url doit mener à un PDF : un lien vers une page DOI est pénalisé par Google Scholar (avant : tout lien externe). #}
Accès ouvert · CC BY-SA

New uses for old books: Description of digitised corpora-based on the Setswana language collection in the WITS Cullen Africana Collection

Article scientifique 2021 Anglais

Résumé

This paper described a corpus of 104 books separated from a larger collection of African Langaguge books. The books were catalogued into a standard library and archival metadata. A subset was digitised and cleaned. The books were then divided into five subsets and compared against each other and the entire Corpus. We have also created tables of collocates, words frequencies. We also performed basic statistics on those words(see tables in the appendix). We speculated that the Corpus as a whole could be roughly used as a general language register. We also give some examples of the characteristics of the genre subsets. The paper aims to introduce the Corpus to NPL researchers and offer it for further research.

Citer ce document

Rahlao, M., Lewin, N., & Surtee, T. G. (2021). New uses for old books: Description of digitised corpora-based on the Setswana language collection in the WITS Cullen Africana Collection. https://doi.org/10.55492/dhasa.v3i03.3819

Exporter : BibTeX · RIS (Zotero, Mendeley, EndNote)

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Licence et provenance

Licence : CC BY-SA

Notice moissonnée depuis OpenAlex le 05/09/2026. Le document reste hébergé par sa source.
Voir le document à la source →

Statistiques

Consultations : 1

Téléchargements : 0