The single online interface to the Greek Learner Corpus (GLC), the largest freely available error-annotated corpus of Greek as a second or foreign language, with rich learner metadata.
Use the tabs above to explore the corpora: descriptive Overview statistics, Learner Analytics, and the two subcorpora GLC v.1 and GLC v.2 (error annotations, search/concordancer and part-of-speech views).
≈ 450 written productions (~33,000 words) by adolescent learners, annotated with the UAM Corpus Tool. Its error scheme is the basis for GLC v.2.
Open the GLC v.1 tab for the annotation scheme, error browser, concordancer and POS views.
≈ 422,000 tokens of written and spoken learner Greek (36 first languages, CEFR A1–C2), plus a comparable native-speaker control subcorpus, with a multi-layered grammatical error annotation.
Open the GLC v.2 tab for the annotation scheme, error browser, concordancer and POS views.
annoteca, the annotation companion of the GLC: the place where the corpus annotations are edited, compared across annotators and curated towards a gold standard. The Gateway shows the results; annoteca is where they are made.
Request accessFor the full description of the corpus, its design and annotation, see the About GLC tab.
The Greek Learner Corpus was developed, until 2023, with the support of the Hellenic Foundation for Research and Innovation (H.F.R.I.) within the project Latent Aspects in L2 Acquisition (LAL2A); it is now maintained and extended independently, beyond that funded project.
The Greek Learner Corpus (GLC) is the largest online, freely available learner corpus of Greek as a second / foreign language (L2/FL). It brings together two error-annotated subcorpora (GLC I and GLC II) of written and spoken learner productions, each accompanied by a rich set of metadata relevant to L2/FL teaching and acquisition. GLC II was compiled within the research project Latent Aspects in L2 Acquisition (LAL2A) during 2019-2022, funded by the Hellenic Foundation for Research and Innovation (H.F.R.I./EL.ID.EK., project no. 3161), at the Aristotle University of Thessaloniki.
GLC v.1 (the second tab on the left) is a collection of about 450 written productions (~33,000 words) by adolescent learners of Greek, compiled at the Aristotle University of Thessaloniki under the supervision of Alexandros Tantos and Despina Papadopoulou (Tantos & Papadopoulou, 2014). It draws on the language-proficiency assessment tests ‘Let’s speak Greek I, II, III’, taken in the Reception Classes of Greek schools between 2010 and 2014, and was annotated with the UAM Corpus Tool. Its error-annotation scheme formed the basis for GLC v.2. GLC v.1 is also accessible through the Greek node of CLARIN, CLARIN-EL.
GLC v.2 (GLC II, the third tab) is the larger and more recent corpus. In its open release, compiled during 2019-2022, it comprises 1,101 written texts (197,713 word tokens) and 318 spoken interviews (~156,005 tokens), about 422,360 tokens in total, produced by adult learners from 36 first-language backgrounds (most frequently Russian, Arabic, Spanish, Serbian and Turkish) and spanning the CEFR levels A1–C2. It is a cross-sectional corpus, offering a snapshot of learner performance at beginner, intermediate and advanced levels. Productions were contributed by a number of institutes in Greece and abroad: the School of Modern Greek Language of the Aristotle University of Thessaloniki, the Greek-language department for women immigrants and refugees at the Lyceum Club of Greek Women (Athens), the private language school Peek at Greek (Thessaloniki), the Sismanoglio Foundation in Istanbul, and the branches of the Hellenic Foundation for Culture in Belgrade, Odessa, Florence and Alexandria.
The written productions were elicited through a Google Form organised in three sections (project description and informed consent, learner profile, and task profile), across three genres (description, narration and argumentation), with an upper time limit of one hour. Besides the digital submissions, several hand-written texts were collected and digitised following a well-defined protocol that ensured a standardised encoding of features such as eroded letters, deletions, segmentation and paragraph indicators.
Teachers and philologists with students learning Greek as a second / foreign language who would like to contribute their own students' productions are welcome to get in touch with us.
The spoken subcorpus was collected through recorded online (Zoom) interviews of 6–9 minutes on the same topics; half of the recordings have already been transcribed following the conventions of Conversation Analysis, and the remainder will be added to the Gateway progressively.
GLC v.2 is also accompanied by a comparable control subcorpus of written productions by L1-monolingual Greek speakers (~66,645 tokens), matched to the learners for age, sex and educational background and based on the same topics. This alignment enables direct comparison between learners’ and native speakers’ productions.
The corpus carries a multi-layered, stand-off grammatical error annotation, organised hierarchically as Error Domain → Error Category → Error Cluster. Four morphosyntactic domains are covered: Agreement, Gender Assignment, Aspect and Voice. For instance, Agreement is analysed into Gender (masculine / feminine / neuter), Number (singular / plural) and Case (nominative / accusative / genitive / vocative). The annotation was carried out by six trained linguists working in pairs, reaching an inter-annotator agreement of around 90%, and is documented in an annotation manual that supports consistency and reuse across future expansions of the corpus.
This web application, the Greek Learner Corpus Gateway, is the single interface to both GLC I and GLC II. It provides descriptive statistics for each corpus, error-profile plots at corpus and text level, the error annotations and metadata in searchable tabular form, advanced search and download across any combination of linguistic and extralinguistic variables, and additional dashboards, learner-analytics, concordance and part-of-speech views.
Beyond dissemination, the GLC data and its annotations are hosted on the annoteca platform, a dedicated annotation platform, currently under development, that supports collaborative corpus annotation workflows: multi-annotator annotation, inter-annotator-agreement measurement, and curation towards a gold standard. There, in the near future, the GLC annotation will be able to be maintained, reviewed and extended in a reproducible, multi-user environment, while the corpus serves as the foundation for AI-assisted annotation of new L2-Greek data.
GLC v.2 follows a clear-cut design (see the figure below) that lets researchers quickly judge whether the corpus meets their interests.
Organization of the Greek Learner Corpus II
Plain-language definitions of the indicators and the main data fields you meet in the tables and downloads.
Shown in the Dashboard, the CEFR development view, the per-text Text Profile cards and Compare Groups. Values are computed per text; group figures are means, except the error rate, which is pooled over the group's texts. Sentences are segmented on sentence-final punctuation and newlines.
| Errors / 1k tokens | Annotated errors per 1,000 word tokens (punctuation excluded). A proxy for accuracy. |
| Lexical diversity (MTLD) | Measure of Textual Lexical Diversity: how varied the vocabulary is, robust to text length (unlike a raw type/token ratio). Higher = more varied. Computed by scanning the token stream and counting how often the running type/token ratio falls to 0.72, averaged forward and backward. |
| Mean sentence length | Average number of words per sentence. A fluency / complexity proxy. |
| Lexical density | Share of content words (nouns, verbs, adjectives, adverbs) among all word tokens. Range 0–1; higher = more information-dense. |
| Subordination (SCONJ / sentence) | Subordinating conjunctions (SCONJ, e.g. «ότι», «όταν», «επειδή») per sentence. A syntactic-complexity proxy. |
| Clauses per sentence | Average number of clauses per sentence, counted from dependency relations (root + advcl + ccomp + xcomp + acl + csubj). Needs a dependency parse; left blank where one is not available. |
| Error-free sentences (%) | Percentage of a text's sentences that carry no annotated error. An accuracy measure. |
The error tables and the error browser (GLC v.1 and v.2).
| Error type | The broad grammatical category of the error (e.g. agreement, case, article, spelling). GLC v.1 uses the UAM CorpusTool scheme; GLC v.2 the Lionda-Miga morphosyntactic scheme. |
| Error subtype | A finer label under the type (e.g. case: accusative). One GLC v.1 error span can carry several stacked subtypes, ordered broad→specific. |
| Target_Form | The intended, grammatically correct form of the erroneous span (the correction). For GLC v.2 it is human-annotated; for GLC v.1 it is an LLM-proposed layer that is pending review. |
| Target_Lemma | The dictionary (lemma) form of the corrected word. |
| Offset (v.2) / Start–End (v.1) | The character position of the error span within the text. GLC v.2 Offset is 1-based; GLC v.1 Start/End are 0-based. |
| Relation / Relation_Type (v.2) | Links the two members of an agreement error (e.g. a noun and the article that fails to agree with it) and its type/direction. Only partly represented in the Gateway's flat schema; full relational fidelity lives in annoteca. |
Per-text learner background, used for filtering and for group comparison.
| Proficiency level (CEFR) | The learner's Common European Framework level (A1 to C2). |
| L1 (first language) | The learner's native language. |
| Genre / writing task | The text type or writing task (e.g. narrative, description, argument). In GLC v.2 these are the Genre / Essay-choice fields. |
| Country of origin | The learner's country of origin, as recorded in the metadata. |
| Sex / Gender | The learner's sex/gender as recorded in the metadata (GLC v.2: Sex; GLC v.1: Gender). |
The POS tables (POS tab) and the lexical measures. GLC v.1 POS is LLM-tagged and not yet audited; GLC v.2 POS had a high-confidence (Tier A) second-opinion audit applied.
| Form | The surface word form exactly as it appears in the text. It is never rewritten, even for errors (the intended form lives in Target_Form). |
| UPOS | Universal part-of-speech tag (NOUN, VERB, ADJ, DET, ...). A blank UPOS on fused tokens like «στο/στην» marks a Universal Dependencies multiword token (σε + article), not a tagging error. |
| XPOS | Language-specific (fine-grained) part-of-speech tag. |
| Lemma | The dictionary base form of the word. |
| FEATS (features) | Morphological features (Case, Gender, Number, Tense, Aspect, Voice, ...) in Universal Dependencies format. |
| Normalized (v.1) | A corrected-lemma layer (what the learner meant) used for normalized lexical frequency. GLC v.1 only; it is filled when the POS review is applied. |
If you use the Greek Learner Corpus or this Gateway in your research, please cite:
1 This chapter describes an earlier version of this platform.
A version-stamped, DOI-backed citation for dataset releases is planned.
The Greek Learner Corpus was developed, until 2023, with the support of the Hellenic Foundation for Research and Innovation (H.F.R.I.) within the project Latent Aspects in L2 Acquisition (LAL2A); it is now maintained and extended independently, beyond that funded project.
Texts
Tokens
Types
Errors
Errors / 1k tokens
Texts
Tokens
Types
Errors
Errors / 1k tokens
How the selected indicator develops across CEFR levels (per-text means; error rate is token-normalised). Clauses/sentence needs a dependency parse and is available for GLC v.2 only.
Share (%) of each error type within every level.
Each cell counts the sentences in which two error types occur together (darker = more frequent co-occurrence); the diagonal is excluded. Click a cell to list those sentences below.
The GLC v.1 grammatical error taxonomy, as defined in the UAM CorpusTool project. Categories keep their original Greek labels and are nested from the broadest error domains down to their finer subcategories; each carries a short gloss and, where the category occurs in the data, one real example drawn from GLC v.1.
AGREEMENT
Error type | Information | Examples |
gender-a | unidentified gender assignment | η λεκσι ικογενια από μόνο του |
gen-masc | erroneous use of masculine gender | στον ακτή |
gen-fem | erroneous use of feminine gender | είχα να δω τις γονείς μου |
gen-neut | erroneous use of neuter gender | η γεύση του καπνού είναι δυσάρεστο |
case | unidentified case | στη σκιά του σμαραγδένιο πράσινο |
case-nom | erroneous use of nominative case | τον χειμώνας |
case-gen | erroneous use of genitive case | η πρώτη στιγμής μιας μέρας |
case-acc | erroneous use of accusative case | εμένα περίμεναν οι καλύτερους ανθρώπους |
case-voc | erroneous use of vocative case |
|
num-sing | erroneous use of singular number | στην άλλη μέρες |
num-pl | erroneous use of plural number | εσείς ξέρεις ότι θέλουμε… |
pers-first | erroneous use of first person | υπάρχω μία φίλη μου |
pers-second | erroneous use of second person | ο James ξυπνάι μόνος σου |
pers-third | erroneous use of third person | το σπίτι μου που μόλις αγόρασε είναι ωραίο |
relation | ambiguous case: error in gender assignment or in gender & number agreement | η γάλα, τα μυρωδιά, τα τεραπή |
VOICE
Error type | Information | Examples |
active | erroneous use of active voice | Είναι δύσκολο να ενσωματώσουν σε μια ξένη κοινωνία |
passive | erroneous use of passive voice | Ετοιμαστήκαμε τις τσάντες μας από την προηγούμενη μέρα |
Verb_type/dep | erroneous use of deponent verb |
|
Verb_type/Unac | erroneous use of unaccusative verb | o σύζυγος της Τζοάνας διασκεδαζόταν |
Verb_type Unerg | erroneous use of unergative verb |
|
GENDER
Error type | Information | Examples |
gender-g | unidentified use of gender | χαρούμενες πάρτες |
syncretism |
| χαρούμενες πάρτες |
masc | erroneous gender assignment (masculine) | Ο δρόμος μας εκεί ήταν σαν όνειρος |
fem | erroneous gender assignment (feminine) | της κοινόχρηστας |
neut | erroneous gender assignment (neuter) | Τα γκέι δεν μπορούν όμως να έχουν… |
CASE
Error type | Information | Examples |
ccase | unidentified case |
|
nom | erroneous use of nominative case | να μάθω η Ελλεινική γλόσσα |
genitive | erroneous use of genitive case | Ήταν πολύ συγκινητικό να της δούμε έτσι ντυμένη |
acc | erroneous use of accusative case | Ήτανε μείον δεκαπέντε βαθμούς |
voc | erroneous use of vocative case |
|
ASPECT
Error type | Information | Examples |
form | Use of non-target form | κάθεμέρα μείνουσαμε στην παραλία την ολόκληρη μέρα |
imperf | Erroneous use of imperfective aspect | πρέπει να αγοράζουμε κατι για το σπίτι |
pperf | Erroneous use of perfective aspect | ήτανε πολύ δύσκολα να ξεκινήσω να μάθω την γλώσσα. |
perf | Erroneous use of perfect aspect | Ο κινέζικος κινηματογράφος έχει αναπτυχθεί όλο και πιο γρήγορα από το 2013 |