Abstract
Georgian is learned by more people every year [19], with almost no methodological literature specific to it: the major language platforms never built for the language, and most available materials predate the research this paper synthesizes. We review the five effects in second-language acquisition and memory research that have replicated across decades, languages, and laboratories: acquisition from comprehensible input [1], the necessity of produced output for grammatical accuracy [3], the superiority of retrieval over restudy [5], the superiority of distributed over massed practice [7], and the top-heavy distribution of word frequency that makes acquisition order a first-class design decision [8]. We then locate where the difficulty of Georgian actually lives, against the US Foreign Service Institute's classification of Georgian as a roughly 1,100-class-hour language [11]: not in the script, which is fully phonemic and small, but in verbal morphology, in a lexicon that shares almost nothing with European languages, and in a spoken register whose core vocabulary diverges sharply from written-corpus frequency lists, a gap we measured in our own published corpus work [16]. We state how these findings are operationalized in the EasyGeorgian curriculum, and we document in full the methodology of our free CEFR-aligned placement instrument, including its adaptive ability estimation and its deliberate limits. This is a narrative synthesis with a declared commercial stake, published so that our design choices can be checked against their sources, and corrected in public when the evidence moves.
1. Introduction
Language teaching has a folklore problem, and small languages have it worst. Where English or Spanish teaching is disciplined by decades of published research, competing schools, and commercial pressure, teaching materials for a language like Georgian are shaped by whatever tradition local classrooms inherited: script drills first, grammar tables early, translation as the default exercise. None of the mainstream platforms whose scale forces methodological scrutiny ever built for Georgian [19], so the corrective never arrived. Learners feel the result. The dominant complaints we hear, and the dominant failure pattern in the field, are months spent on the alphabet, verb paradigms memorized and never produced in speech, and vocabulary learned from lists that do not match what Georgians actually say.
This paper does three things. Section 2 states what the acquisition and memory literature actually agrees on, with sources, at the level of confidence each finding deserves. Section 3 applies it to Georgian specifically: what is genuinely hard, what is falsely feared, and what the language's typology changes about method. Sections 4 and 5 document how we operationalize the evidence, including the complete methodology of our placement instrument. We build and sell courses based on this synthesis, and that stake is declared here and reargued in Section 6, because the honest defense against motivated reading is publishing the reasoning where it can be attacked.
2. The evidence base
Second-language acquisition is a young field with noisy data, and famous ideas have failed under scrutiny: matching instruction to self-reported learning styles, for one, has no supporting evidence that survives controlled testing [14]. What follows is the short list of effects that have not failed: findings replicated across enough paradigms, populations, and decades that building against them is the conservative choice, not the fashionable one.
Table 1 · Five replicated effects and their design consequences
The synthesis in one table. Each row is expanded, with sources, in the section indicated.
| Effect | Representative evidence | Design consequence | § |
|---|---|---|---|
| Acquisition from comprehensible input | Krashen (1982); Asher (1969) [1][2] | Material sits just past the learner's level; listening precedes reading | 2.1 |
| Output drives accuracy | Swain (1985); DeKeyser (2007) [3][4] | Learners produce the language from the first session, under mild time pressure | 2.2 |
| Retrieval beats restudy | Roediger & Karpicke (2006) [5] | Every exercise is a prompt with a retrieval gap, never a re-read | 2.3 |
| Spacing beats massing | Ebbinghaus (1885); Cepeda et al. (2006) [6][7] | Scheduling is the system's job: items return at widening intervals | 2.4 |
| Frequency is top-heavy | Zipf (1935); Nation (2006) [15][8] | Vocabulary is taught in measured spoken-frequency order | 2.5 |
Persistence (Section 2.6) is deliberately excluded from the table: it is a moderator of everything above rather than a separate effect.
2.1 Comprehensible input
The strongest single claim in the field is Krashen's input hypothesis: learners acquire language by understanding messages in it, and the input that teaches sits just beyond current competence, close enough that context carries the unfamiliar parts [1]. Input far below the learner's level contains nothing new; input far above it is noise. The comprehension-first tradition predates the hypothesis: Asher's Total Physical Response experiments found that learners whose early hours were spent entirely on comprehension outperformed learners pushed to speak from the first minute, including on later speaking measures [2].
Two refinements matter for the present case. First, Krashen's later argument for narrow listening: input from the same speakers on recurring topics compounds, because each session reuses the last session's vocabulary and the learner's effort concentrates on the genuinely new [9]. Second, listening and reading are not interchangeable sources of input. Speech runs on a measurably different vocabulary than writing, and for Georgian the divergence is extreme (Section 3.4).
2.2 Output and automatization
Input alone has a documented failure mode. Swain's studies of Canadian immersion students found learners with thousands of hours of comprehensible input whose production remained grammatically broken, and her diagnosis became the output hypothesis: comprehension can succeed by skimming meaning from word order and context, so only the attempt to produce forces the learner to notice the grammar they do not control [3]. Skill acquisition theory supplies the mechanism for the rest: fluent speech is not knowledge but a performance skill, and performance skills automatize through repeated execution under mild time pressure, the way typing and sight-reading do [4]. The design consequence is unambiguous. If the target is speaking, the majority of practice must be speaking, from the first session; readiness is a result of production, not its prerequisite.
2.3 Retrieval practice
The testing effect is among the most replicated results in memory research. In Roediger and Karpicke's canonical design, students who recalled a text from memory retained dramatically more after a week than students who re-read it for the same time, and the re-readers were more confident while learning less [5]. The fluency of restudy is precisely the problem: memory strengthens in proportion to the effort of retrieval, not the comfort of exposure. For vocabulary learning the implication is blunt. A visible word list teaches little; a prompt, a beat of effortful silence, and then the produced form is what writes the item in. Serious practice systems therefore hide the answer until the learner has attempted it.
2.4 Distributed practice
Ebbinghaus drew the forgetting curve in 1885 [6], and the modern quantitative synthesis, covering more than 250 experiments, confirms what the curve implies: for a fixed number of repetitions, spreading them out produces far more durable memory than massing them, with the optimal gap growing as the memory matures [7] (EX-1). Cramming wins the test tomorrow and loses the language next year. Because no learner will schedule hundreds of per-item review dates by hand, spacing is only real when the system does it: either an explicit scheduler over discrete items, or repetition structured into the material itself.
2.5 Frequency and coverage
Word frequency in every measured language follows a sharply top-heavy distribution [15]: a few hundred words carry most running speech, and returns fall steeply from there. Nation's coverage research quantifies the consequence for learners: reaching roughly 98 percent coverage of running words, the level at which unassisted comprehension becomes comfortable, requires on the order of 6,000 to 7,000 word families for speech and 8,000 to 9,000 for writing, but the first thousand families do vastly more work per word than the sixth [8]. Acquisition order is therefore not a detail: the learner who knows the 2,000 most frequent words of speech outperforms, in every practical situation, the learner who knows 2,000 assorted textbook words. The precondition is a frequency list that measures the register the learner needs. Section 3.4 shows how badly that precondition fails for Georgian if taken on faith.
2.6 Persistence
The least glamorous variable decides the rest: every effect above is conditional on the learner still practicing months in. The motivation literature treats persistence not as a character trait but as something materials support or drain, through session length, visible progress, and content connected to the learner's actual reasons [10]. A theoretically optimal routine abandoned in week three is outperformed by a merely good routine sustained for a year. We treat persistence as a design constraint of the same rank as the five effects, while noting its evidence base is the softest of the six: it rests on individual-differences research rather than controlled memory experiments, and honest synthesis says so.
3. The Georgian instance
3.1 Locating the difficulty
The US Foreign Service Institute, which has trained diplomats in languages for seven decades and publishes the only long-running difficulty classification anchored to instructional outcomes, places Georgian in Category III: roughly 44 weeks and 1,100 classroom hours to professional working proficiency, alongside Russian, Finnish, and Turkish, and below only the 2,200-hour Category IV languages [11].
Table 2 · FSI difficulty categories, with Georgian located
Classroom hours to professional working proficiency for a native English speaker, per the Foreign Service Institute's published training estimates.
| Category | Approx. class hours | Example languages |
|---|---|---|
| I | 600–750 | Spanish, French, Italian, Dutch |
| II | ~900 | German, Indonesian, Swahili |
| III | ~1,100 | Georgian, Russian, Finnish, Turkish, Hebrew |
| IV | ~2,200 | Arabic, Mandarin, Japanese, Korean |
Source: US Department of State, Foreign Service Institute, foreign language training estimates [11]. Hours are institutional estimates for intensive classroom study, not self-study equivalents.
An aggregate difficulty number, however, is methodologically useless until it is decomposed, because learners systematically fear the wrong parts of this language and spend their hours accordingly. The decomposition below is the empirical core of this paper's Georgian-specific argument.
3.2 The script
Mkhedruli, the modern Georgian script, is 33 letters, unicameral, with an almost perfectly one-to-one letter-to-phoneme mapping: no case distinctions, no contextual letter forms, and effectively no spelling ambiguity in either direction [12]. As a decoding task it is among the easiest scripts a learner of any language will meet, far closer to a new cipher for known sounds than to a new writing system in the sense that Chinese characters or Arabic orthography are. The traditional practice of spending the first weeks of a course on handwriting drills therefore inverts the difficulty profile: it front-loads the easiest component of the language at the point of highest learner motivation, and defers contact with the hard components. In our curriculum the script is a single-afternoon objective handled by a free standalone resource, and course time begins on the spoken language immediately.
3.3 The verb
The Georgian verb is where Category III is earned. A single finite form can index subject and object person simultaneously (polypersonal agreement), carries preverbs that fuse direction with aspect, distributes across eleven traditional screeves in three series, and splits its case alignment: nominative subjects in the present series, ergative in the aorist, dative in the perfect [13]. Class membership and irregularity mean the paradigm resists the tabular treatment that works acceptably for, say, Spanish conjugation. This is precisely the grammar type for which the evidence of Sections 2.1 and 2.2 matters most: complex, semi-regular morphology is acquired robustly through massive exposure to whole forms in meaningful contexts plus repeated production of those forms, with explicit paradigm study functioning as an organizing aid afterwards, not as the acquisition route. Our courses accordingly teach verb forms as produced chunks inside sentences, sequence the screeves by measured spoken frequency rather than by paradigm completeness, and leave table lookup to a free reference conjugator. Independent corroboration that this morphology is objectively the hard core of the language comes from an unexpected direction: in our AI benchmark, verb morphology and rare-paradigm production are exactly the dimensions where otherwise near-perfect frontier systems drop furthest [17].
3.4 Vocabulary without cognates
A Kartvelian lexicon offers an English speaker almost no free vocabulary: the cognate discount that lets a Spanish learner recognize thousands of words on sight is close to zero for Georgian. Every one of Nation's thousands of word families (Section 2.5) must actually be learned, which multiplies the value of both frequency ordering and systematic spaced retrieval relative to a European language. It also raises the cost of a wrong frequency list, and for Georgian the standard lists are wrong in two documented ways. First, the most widely copied "Georgian frequency list" online is not Georgian: our corpus work decoded it as a Bulgarian frequency list corrupted through a broken character encoding, and derivative flashcard decks continue to circulate [16]. Second, even legitimate written-corpus rankings diverge from speech at the exact ranks a learner needs first. In our published study, built on a lemmatized spoken corpus with teacher review, core conversational words sit thousands of ranks deep in the written list.
Table 3 · Spoken rank versus written rank, five everyday words
Rank of the lemma in our spoken corpus versus the standard written-corpus frequency ranking. Words a conversation cannot proceed without are effectively invisible to written-frequency ordering.
| Word | Gloss | Spoken rank | Written rank |
|---|---|---|---|
| არა ara | no, not | 16 | 8,662 |
| შემიძლია shemidzlia | I can | 30 | 9,599 |
| მჭირდება mch'irdeba | to need | 71 | 5,189 |
| ალბათ albat | probably | 130 | 9,603 |
| მართლა martla | really | 143 | 9,512 |
Source: What Georgians Actually Say, EasyGeorgian Research 2026, Table 2; full corpus methodology and CC-BY dataset at the study page [16].
3.5 Phonology before orthography
Georgian contrasts ejective consonants with their aspirated counterparts, a phonemic distinction English lacks entirely, and permits consonant clusters of a length European learners have no articulatory habit for [13]. These contrasts do not survive transliteration: the official romanization marks ejectives with an apostrophe that informal romanizations simply drop, collapsing distinct phonemes into one letter. A learner whose first contact with a word is its Latin-script rendering therefore anchors a wrong sound that later listening must overwrite. This is a second point where our benchmark work supplies independent evidence of the trap's reality: most tested AI systems, the source of a growing share of informal learning material, fail to produce the official romanization correctly, with systematic apostrophe-dropping as the dominant error [17]. The methodological consequence is audio-first sequencing not as a stylistic preference but as damage control: the ear meets the form before any romanization can mis-anchor it.
4. Operationalization
A synthesis earns its keep only in the design decisions it forces. Ours, stated as commitments that can be checked against any of our materials: input stays comprehensible and narrow, with the same voices across a course and new material embedded in known material (2.1); every lesson demands spoken production from the first session, under mild time pressure, roughly a hundred retrievals per half-hour session (2.2, 2.3); repetition is scheduled by the system, inside lesson scripts at widening intra- and inter-lesson gaps and, for discrete vocabulary, by an SM-2 spaced-repetition scheduler (2.4); vocabulary follows our measured spoken-frequency ranking, not a written list and not topical chapters (2.5); sessions are sized to survive real weeks, and progress is always visible (2.6); the script is dispatched in an afternoon by a free resource (3.2); verb forms are taught as chunks in context with paradigms available as reference, not curriculum (3.3); and no learner meets a word in romanization before meeting it in audio (3.5).
Two disclosures belong next to that list. These commitments are instantiated in commercial products, which is a reason to scrutinize the synthesis, not to excuse it. And the commitments are evidence-based, not evidence-proven: no randomized trial of our specific curriculum exists (Section 6), so what we claim is fidelity to the strongest available findings, and publication here is the mechanism by which that fidelity can be audited.
5. Measuring placement: the level test
Placement is methodology's first decision: material pitched a level wrong wastes months on either side (2.1). When we could find no free, instant, CEFR-aligned [18] placement instrument for Georgian, we built one and made it public. Georgia has an official Reference Level Description for the language (A1–B2) [20] and a state certification exam arriving from 2026; a proctored certificate answers a different question than the one a self-directed learner has, which is where to start. This section is the instrument's full methodology, preserved from its original standalone publication.
5.1 Instrument design
The design borrows from the best-studied placement instruments, DIALANG's vocabulary-routed architecture in particular [21], adapted to a click-only web session of about eight minutes.
Table 4 · Instrument structure
Five stages, click-only, about eight minutes in total. Typing is never scored: too many capable learners lack a Georgian keyboard, and scoring it would measure hardware.
| Stage | ~Time | Task | Scoring role |
|---|---|---|---|
| Screener | 30 s | Two self-report items plus script decoding: hear a word, pick its spelling among near-identical strings | Routing; non-readers go to a spoken track (audio-to-picture, A0/A0+ result) instead of failing grammar items |
| Vocabulary | 2.5 min | Rapid yes/no recognition: 36 real words stratified across spoken-frequency bands, 12 phonotactically plausible pseudowords | Vocabulary-size estimate with false-alarm correction, LexTALE-style [22], plus a deliberately weak starting point for the adaptive stage (Section 5.2). Pseudowords are the honesty check |
| Adaptive core | 3.5 min | 14 items selected near current ability: gap-fills, error identification, short reading, dialogue response, tile-assembly sentence building | Primary ability estimate; one deliberate stretch item probes above level |
| Listening | 2 min | 5 audio items from single words to short monologues, one replay allowed | Listening accuracy read-out; feeds the same ability estimate |
| Typed bonus | optional | Two typed-recall prompts with a prominent skip | Unscored in this version |
No mid-test right/wrong feedback (it induces anxiety and answer-gaming), no visible timer, and every scored item carries an "I don't know yet" option (Section 5.2).
5.2 Ability estimation
Ability is estimated with a constrained Elo-style update, the approach the adaptive-education literature recommends when no pretested item-response population exists [23]. Item difficulties are seeded from the construct map's band anchors and word-frequency ranks, then held fixed during live sessions. The ability estimate updates after every response with a decreasing step size, items are selected near the current estimate with light randomization, and the final score smooths the last estimates to reduce noise. The vocabulary stage sets the starting estimate, and that routing is deliberately weak: yes/no word recognition is the cheapest signal in the instrument and the poorest proxy for what a learner can actually do, so it nudges the starting point and the adaptive stage does the placing.
The first 43 completed sessions, in August 2026, produced the instrument's first revision. We report it in full, including the part that reflects badly on the original design. The vocabulary stage had been mapping its score to a starting ability through a step function, so a single yes/no answer at the wrong moment could shift a taker's starting point by as much as 0.8 on the ability scale, most of a CEFR band. The step that did the damage was the one at the top: either side of it, one answer moved the starting point by 0.7, and that step sat directly beneath the cutoff for the highest band we report. Its top value also placed a taker close enough to the top band that they had only to avoid falling out of it, and all 22 sessions that reached that value finished in the top band, with no exceptions. The mapping is now interpolated and its ceiling sits in the middle of B1.
The same data was then tested for the more obvious suspicion, that the seeded item difficulties were themselves wrong, and did not support it. Comparing bands within a single session, where the same person answers both sets and the comparison therefore does not depend on the ability scale being correct, the seeded spacing holds: items written to C1 sit 0.98 above items written to B2 against 1.0 as assigned, and A2 items sit 1.16 above A1 items against 1.1. Only the B2-over-B1 step is materially off, and it rests on seven sessions. The difficulties were left alone.
What the data does show is a ceiling. The hardest items in the bank are answered correctly about 92 percent of the time by the people who reach them, and 9 of the 17 items written to the top band have never been answered incorrectly by anyone. Above roughly B2 the instrument has run out of questions rather than mislabeled the ones it holds, which is why the top band is open and reports as B2+ (Section 5.4). Re-estimation of difficulties from response data happens offline and is reported here when it changes the instrument, under a versioned item-bank changelog. At current volumes that is occasional rather than scheduled: the per-item sample is still thin, and leaving a seeded value in place is better than fitting noise.
The "I don't know yet" option is engineered rather than decorative. Random guessing on a four-option item injects variance into the estimate; the IDK update is calculated to be exactly expectation-equivalent to a random guess with zero variance, so honest skipping never costs anything relative to gambling, gambling never gains anything relative to skipping, and the estimate stabilizes faster. On multiple-choice items a committed wrong answer costs marginally more than an IDK. The property holds mathematically and is asserted by the test suite.
5.3 Content anchoring
Two anchors hold the content to external standards. The construct map, which assigns grammar to bands, is checked against the official Reference Level Description for Georgian [20] and the published level-requirement papers for Georgian as a foreign language [24], with contested grammar arbitrated against Hewitt's reference grammar [13]. Roughly half the item bank adapts sentences already verified by Georgian teachers for our course materials; the remainder was authored fresh. The pseudowords in the vocabulary stage are generated from Georgian phonotactics, checked for collisions against our full word-form corpus, and reviewed so that none is a real word in any register or dialect.
5.4 What the test does not measure
The limits are deliberate. No speaking or pronunciation scoring: automated assessment of Georgian speech is not currently good enough to score respectably, so we do not pretend to. No writing assessment: the typed bonus stays unscored. Recognition, not production: every scored item is a choice among presented options, so a learner who has studied Georgian grammar closely can place above the level at which they can speak or write it, and several early takers told us exactly that about their own results. No result above B2+: the top band is open because the hardest questions in the bank are answered correctly about 92 percent of the time by the people who reach them (Section 5.2), so any label we printed above that line would be extrapolation rather than measurement, and separating C1 from C2 needs productive tasks in any case. Not a certificate: the result is an indicative placement with zero stakes. Band cutoffs are provisional, seeded from the official level descriptions and frequency data, and revised as sessions accumulate, with the first revision reported in Section 5.2. The result screen says so. Sessions are logged anonymously (random session id, responses, timing; no account, no fingerprinting), an email is stored only on explicit request for the full report, and accuracy statistics against learners of known level will be published as the data matures.
6. Limitations
Four limitations bound this paper's claims, stated in descending order of importance. First, the conflict of interest: we sell courses built on this synthesis, which gives us a motive to read the literature in our own favor. The mitigations are that every empirical claim is cited to its source, the synthesis commits us publicly to positions that can be checked against those sources, and corrections reach us at [email protected] and are versioned. Second, this is a narrative synthesis, not a systematic review or meta-analysis: we selected the effects by replication weight across the field's standard literature, not by exhaustive search protocol, and a different reviewer could weight differently, particularly on motivation (2.6), where the evidence is observational. Third, generalization: the underlying experiments overwhelmingly involve English-adjacent target languages and literate adult participants; no controlled acquisition study specific to Georgian as a foreign language exists to our knowledge, so Section 3 is an argued application of general findings to Georgian's typology, corroborated where possible by our own measured data [16][17], not a body of Georgian-specific experimental results. Fourth, no randomized trial of the EasyGeorgian curriculum itself exists; outcome claims for our specific products are therefore fidelity claims, not efficacy claims, and we say so wherever the distinction matters. The placement instrument carries its own limitations, stated in Section 5.4.
7. Conclusion
Nothing in this paper is a secret method. Comprehensible input, forced early production, retrieval over restudy, engineered spacing, and measured-frequency vocabulary are the boring consensus of half a century of research; they are simply consensus that Georgian, skipped by every platform whose scale forces methodological discipline, never received. The language's actual difficulty profile, a trivial script attached to a genuinely hard verb, a cognate-free lexicon, and a spoken register that written frequency lists cannot see, makes the consensus matter more here, not less. We publish the synthesis, the Georgian-specific measurements, and the placement instrument's full design so that all three can be cited, checked, and corrected. The evidence base will move; when it does, this page moves with it, in public.
Authorship and roles
Published by EasyGeorgian (Lasse N., Tamar N.). Research synthesis, drafting, and instrument design by Claude (Anthropic) operating under editorial direction, with every cited claim checked against its published source; Georgian language examples and course-derived materials verified by Georgian teachers as part of the underlying course pipeline; corpus measurements and benchmark results drawn from our separately published studies [16][17], each with its own methodology and data release. We state the division of labor because a methodology paper's trustworthiness begins with who did what.
References
- Krashen, S. (1982). Principles and Practice in Second Language Acquisition. Pergamon Press. Author's edition at sdkrashen.com
- Asher, J. (1969). The Total Physical Response approach to second language learning. The Modern Language Journal, 53(1).
- Swain, M. (1985). Communicative competence: some roles of comprehensible input and comprehensible output in its development. In Gass, S. & Madden, C. (eds), Input in Second Language Acquisition. Newbury House.
- DeKeyser, R. (2007). Skill acquisition theory. In VanPatten, B. & Williams, J. (eds), Theories in Second Language Acquisition. Routledge.
- Roediger, H. L. & Karpicke, J. D. (2006). Test-enhanced learning: taking memory tests improves long-term retention. Psychological Science, 17(3). doi.org
- Ebbinghaus, H. (1885). Über das Gedächtnis: Untersuchungen zur experimentellen Psychologie. Duncker & Humblot.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. & Rohrer, D. (2006). Distributed practice in verbal recall tasks: a review and quantitative synthesis. Psychological Bulletin, 132(3). doi.org
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? The Canadian Modern Language Review, 63(1).
- Krashen, S. (1996). The case for narrow listening. System, 24(1).
- Dörnyei, Z. (2005). The Psychology of the Language Learner: Individual Differences in Second Language Acquisition. Lawrence Erlbaum.
- US Department of State, Foreign Service Institute. Foreign language training: category rankings. state.gov (accessed July 2026)
- Unicode Consortium / Wikipedia, Georgian scripts (Mkhedruli inventory and unicameral status). en.wikipedia.org (accessed July 2026)
- Hewitt, B. G. (1995). Georgian: A Structural Reference Grammar. John Benjamins.
- Pashler, H., McDaniel, M., Rohrer, D. & Bjork, R. (2008). Learning styles: concepts and evidence. Psychological Science in the Public Interest, 9(3).
- Zipf, G. K. (1935). The Psycho-Biology of Language. Houghton Mifflin.
- EasyGeorgian Research (2026). What Georgians Actually Say: a spoken-frequency study of Georgian. CC-BY dataset and methodology. easygeorgian.com/spoken-georgian-frequency
- EasyGeorgian Research (2026). The Georgian AI Benchmark, 2026 edition. CC-BY dataset and methodology. easygeorgian.com/georgian-ai-benchmark
- Council of Europe (2001). Common European Framework of Reference for Languages: Learning, Teaching, Assessment. Cambridge University Press. coe.int
- EasyGeorgian Research (2026). The State of Georgian Learning 2026. easygeorgian.com/state-of-georgian-learning
- Ministry of Education and Science of Georgia / Council of Europe (2006). Reference Level Description for the Georgian Language, A1–B2. Tbilisi.
- Alderson, J. C. (2005). Diagnosing Foreign Language Proficiency: The Interface between Learning and Assessment. Continuum.
- Lemhöfer, K. & Broersma, M. (2012). Introducing LexTALE: a quick and valid Lexical Test for Advanced Learners of English. Behavior Research Methods, 44. doi.org
- Pelánek, R. (2016). Applications of the Elo rating system in adaptive educational systems. Computers & Education, 98.
- Shavtvaladze, N. (2023). Requirements and description for the A2 level of teaching Georgian as a foreign language. International Journal of Multilingual Education, 22; with the companion A1 and B1 papers in Language and Culture.
(License CC-BY-4.0. Cite as "How Georgian Is Learned: An Evidence Synthesis, EasyGeorgian Research 2026" with a link to this page. Corrections: [email protected], applied in public with a versioned changelog.)