MNLP - Exam Questions
All 53 exam-style questions for Multilingual Natural Language Processing, one set per note. Each answer opens with the key points a grader looks for, then a full model answer. Write or say your answer before revealing. The short drilling cards are on the flashcards page.
Back to the Multilingual Natural Language Processing home.
L01 Overview
Exam questions
Distinguish multilingual from crosslingual NLP. Give examples of each and explain why the English-centric nature of LLMs keeps both open problems.
Key points: multilingual means systems for multiple languages; crosslingual means transferring information across languages; multilingual examples: NER, parsing, classification; crosslingual examples: MT, QA, unlearning, reasoning; LLMs English-centric and fragile on non-English text.
- Multilingual: building systems that work for multiple languages. Examples: NER for multiple languages (can it be trained without being language-specific?), language-independent parsing (are there universal features?), cross-language classification (can it be trained without language-specific data?).
- Crosslingual: transferring information or knowledge across languages. Examples: machine translation, crosslingual QA, crosslingual unlearning, crosslingual reasoning.
- LLMs: trained predominantly on Internet resources, so they are English-centric. Multilingual models such as Llama 3.1 and DeepSeek do better on high-resource languages; dedicated multilingual models such as Aya-Expanse tend to lag behind. LLMs are still fragile on non-English text.
Losing marks: treating the two terms as synonyms. One is about covering many languages, the other about moving information between them.
Contrast pre-deep-learning NLP with deep-learning NLP. What did neural networks gain, what did they cost, and why do insights now transfer between fields?
Key points: before: per-application methods weighing engineered features; gains: state of the art, little or no feature engineering, few network types; costs: data hunger, hard-to-trace errors, spectacular failures; transfer via uniform network types and few application-specific features.
- Before: a different methodology per application. Classical ML (SVMs, decision trees, generative Bayesian models, discriminative max-ent models) weighed hand-engineered features: POS tags, morphology, parse trees, named entities, taxonomies, argument roles.
- Advantages of deep learning: state-of-the-art performance across NLP tasks; very little or no feature engineering; a limited repertoire of network types covers most or all tasks.
- Disadvantages: needs large amounts of training data; errors are hard to trace (opacity); can fail spectacularly.
- Transfer: the same success appears in computer vision, handwriting recognition, speech, robotics and IR. Because network types are uniform and application-specific features are few, insights carry over between these areas.
Link to originalWhat roles do large language models now play in NLP, and what are their limits? Use the Arabic to English translation comparison as evidence.
Key points: LLMs originated in and now dominate NLP; end-to-end QA, summarization, MT; substitute humans in annotation, LLM-as-judge evaluation, dialog; fragile on non-English text; Mistral Small 3.2 hallucinated the embassy’s location.
- LLMs originated within NLP and now dominate research.
- Uses: end-to-end NLP (QA, summarization, MT); substituting or supplementing humans in data annotation, evaluation (LLM-as-judge) and dialog.
- Limit: still fragile on non-English text, because training data is English-centric.
- Evidence (Monz’s Arabic to English example): GPT-5.6 Terra, Gemini 3.6 Flash and Claude Sonnet 4.6 all render the sentence as the embassy in South Sudan calling for an investigation into an attack that killed Ethiopian peacekeepers, differing only in wording. Mistral Small 3.2 produced “The Ethiopian embassy in Khartoum, where the Ethiopian peacekeepers were detained, is closed”: it catastrophically hallucinated the embassy’s location, and its sentence no longer matches the other three.
L02 Multilinguality and Writing Systems
Exam questions
State the four English-first assumptions built into standard NLP pipelines. For each, give languages where it fails and what goes wrong concretely.
Key points: whitespace word boundaries fail for Chinese, Japanese, Thai; word-order grammar fails for case-marking Russian, Finnish, Korean; one word one meaning fails for agglutinative and compounding languages; left-to-right fails for Arabic, Hebrew, Urdu, Persian.
- Words are separated by whitespace. Fails for Chinese, Japanese, Thai, Lao, Khmer, Burmese: there is no word delimiter, so whitespace tokenization returns one token for a whole Chinese sentence. Thai uses spaces to end clauses or sentences and has no dedicated full stop, so a space means something different from what the tokenizer assumes.
- Grammatical relations are encoded by word order. Fails for languages with rich case marking (Russian, Finnish, Korean): who did what to whom is marked on the noun, word order is free, and position-based features carry much less information.
- One word = one unit of meaning. Fails for agglutinative languages (Finnish, Turkish, Hungarian) and compounding languages (German, Dutch, Swedish): Bundesverfassungsgericht is one token holding federal + constitutional + court.
- Text flows left to right. Fails for Arabic-script and Semitic languages (Arabic, Hebrew, Urdu, Persian): rendering runs right to left, so the first character in the string is not the leftmost on screen.
Argue that "script is not language" holds in both directions, and explain what each direction breaks in an NLP pipeline.
Key points: one script serves many languages (Latin, Arabic); script detection does not identify the language; one language uses several scripts (Serbian, Uzbek); language does not fix the script, so train/test scripts can share no characters.
- One script, many languages: Latin (English, German, French, most European languages, several African, Turkic and Asian languages), Arabic (Arabic, Persian, Urdu, Pashto, Sorani), Cyrillic (Russian, Belarusian, Serbian), Devanagari (Hindi, Marathi, Nepali), Ethiopic (Amharic, Tigrinya). So detecting the script does not identify the language: Latin narrows nothing, and Arabic script leaves several languages that are not even all in one family.
- One language, many scripts: Serbian (Cyrillic and Latin), Kurdish (Latin and Arabic), Uzbek (Latin and Cyrillic), Punjabi (Gurmukhi and Arabic). So detecting the language does not tell you the script: a Serbian corpus can mix both, even in one document, and a model trained on one script and tested on the other sees no shared characters for the same language.
- Some scripts are language-specific (Armenian, Georgian, Greek, Sinhala, Khmer), which is the only case where the two coincide.
- Also: one sentence can mix several writing systems (Japanese uses Katakana, Kanji and Hiragana together), so a single “script of this text” label can fail.
Why is the English-centric bias of NLP structural, and why do claims of "language independence" deserve suspicion?
Key points: resources exist first or only in English; non-English data are by-products or translations of English resources; coverage follows commercial and military interest, not speakers; pipelines encode English properties; language independence is rarely tested.
- Almost every resource exists first, often only, in English: annotated data, unannotated data, toolkit coverage (dictionaries, analyzers), benchmarks. Most publications also focus on English.
- Non-English resources arise mostly as by-products (translations, product reviews) or as translations of English annotated resources (so the annotation scheme stays English-shaped), and rarely as dedicated efforts.
- Coverage tracks commercial interest and military relevance, and speaker count is a poor predictor: over 7,000 languages, more than 100 with over 10 million speakers, yet Dutch is only medium resource.
- Standard pipelines encode properties of English (whitespace tokenization, fixed word order, closed vocabularies) as if they were properties of language.
- “Language independent” usually means tested on English plus a very small set of languages: a method gets the label because nobody checked.
- The bias is understandable, since resource creation costs time and money, which is exactly why it does not fix itself.
Compare alphabetic, syllabic and logographic writing systems and explain what each, plus the abjad and Hangul cases, implies for NLP.
Key points: alphabetic: sounds, small inventory; syllabic: whole syllables; logographic: words or morphemes, thousands of symbols; abjad omits vowels, so the model must disambiguate; Hangul: alphabetic jamo in syllable blocks.
- Alphabetic: each symbol is a consonant or vowel sound, small inventory (Latin, Cyrillic, Hangul).
- Syllabic: each symbol is a whole syllable (Hiragana, Katakana).
- Logographic: each symbol is a word or morpheme directly, little sound information (Chinese Hanzi, Japanese Kanji).
- Inventory size sets the cost of a character: tens of symbols for an alphabet, thousands for a logographic script, which a character-level vocabulary must budget for.
- Abjad (Arabic): consonantal skeleton, vowels omitted. كتب can be kataba, kutiba or kutub, so the orthography is lossier than the language and the model must resolve the ambiguity from context. No tokenizer can recover information that was never written.
- Hangul breaks the neat taxonomy: alphabetic by symbol-to-sound mapping (each jamo is a consonant or vowel), syllabic by visual arrangement (jamo grouped into square blocks), and letter shapes depict tongue and lip positions.
Explain the difference between a grapheme, a codepoint and a byte, and show how confusing them corrupts tokenization, deduplication and span annotation.
Key points: grapheme is the perceived character; one or more codepoints per grapheme; 1 to 4 UTF-8 bytes per codepoint; NFC/NFD mismatch silently breaks dedup and search; mixed breaking conventions misalign spans.
- Grapheme: what the reader sees as one character. Codepoint: Unicode’s abstract number for a character, one or more per grapheme. Byte: UTF-8 stores each codepoint in 1 to 4 bytes. They coincide only in ASCII.
- Example: NFD é is 1 grapheme, 2 codepoints (U+0065 U+0301), 3 bytes (65 CC 81).
- Tokenization: byte-level tokenizers see the byte stream. Byte-level splits can produce non-existent characters; codepoint splits can strand a diacritic from its base letter.
- Deduplication, search, equality, prefix matching assume identical codepoint sequences, so NFC and NFD copies of the same text silently fail to match. No error is raised, the numbers are just wrong.
- Span annotation (SQuAD-style QA): offsets recorded under one breaking convention and read under another give shifted or truncated answers. Counts across corpora with different conventions are not comparable.
- Fix: normalize consistently in the pipeline, and pick one breaking convention and enforce it everywhere. Consistency matters more than which level you pick.
Crawled multilingual data mixes normalization forms. Which Unicode normalization form would you use for search indexing, for comparing rendered text, and for generation targets, and why?
Key points: canonical reversible, compatibility lossy; NFD for comparing rendered text; NFKD for search indexing and analysis; avoid compatibility forms for generation; normalize consistently.
- The four forms form a grid: canonical (reversible) NFC and NFD; compatibility (lossy) NFKC and NFKD; composed (C) versus decomposed (D).
- Comparing rendered text: NFD, which splits into base characters plus combining marks and preserves semantic and visual equivalence. It is non-lossy, so you can decompose and recompose.
- Search indexing or NLP analysis: NFKD. Its lossiness is a feature: the ligature fi and the pair fi become the same, ① becomes 1.
- Generation: avoid compatibility forms. If training targets are NFKD-normalized, the model can never learn to produce the character the user actually typed (for example the long s ſ becomes plain s).
- Whatever you choose, apply it consistently in the pipeline (for example decompose then compose), otherwise dedup, search and equality checks are silently inaccurate.
Losing marks: getting the analysis versus generation rule backwards, or claiming NFKD is reversible.
Explain byte inflation for byte-level tokenizers, compute it for "hello", "café" and "中文", and argue why it is a fairness problem.
Key points: UTF-8 bytes over perceived characters; 1.0x, 1.25x, 3.0x; more bytes means more aggressive splitting; CJK users pay more tokens for the same content.
- , characters meaning what a reader perceives.
- “hello”: 5 bytes / 5 = 1.0x. “café” (NFC, é is 2 bytes): 5 / 4 = 1.25x. “中文” (3 bytes each): 6 / 2 = 3.0x.
- More bytes per character means more aggressive splitting, so the same meaning costs more tokens; byte splits can also produce fragments that are not characters at all.
- Fairness: a Latin-script user pays about 1 byte per character, a CJK user 3. The same sentence costs CJK users more tokens: shorter effective context window, more compute for identical content, and less linguistically meaningful units. The English-first assumption reaches the byte layer.
Link to originalEncode U+4E2D (中) in UTF-8 by hand, stating which byte rules you use and why UTF-8's design is self-synchronising.
Key points: 15 bits exceed 11, so 3-byte form; 1110yyyy 10yyyyxx 10xxxxxx; result E4 B8 AD; lead and continuation byte rules; self-synchronising.
- 0x4E2D =
0100111000101101: 15 significant bits, more than the 11 payload bits of the 2-byte form, so use the 3-byte form1110yyyy 10yyyyxx 10xxxxxxwith payload bits.- Split 4 + 6 + 6:
0100 | 111000 | 101101.- Bytes:
11100100 10111000 10101101(hex E4 B8 AD).- Rules: a byte starting
0is a single byte (ASCII); a byte starting11begins a multi-byte sequence, with the number of leading 1s giving its length; a byte starting10is a continuation byte.- Self-synchronising: no byte value is ambiguous about its role, so from any offset you can tell whether you are at a character start and walk back by skipping
10xxxxxxbytes. UTF-8 therefore survives truncation, concatenation and byte-level search.
L03 Morphology and Word Formation
Exam questions
Name the four morphological types, give example languages for each, and illustrate each with one analysed word or sentence. What distinguishes agglutinative from fusional?
Key points: isolating: no inflection, word order and particles (Chinese, Vietnamese); agglutinative: one morpheme, one meaning (Turkish ev-ler-iniz-den); fusional: one affix, several fused features (Russian knig-ami); polysynthetic: a sentence in one word (Inuktitut); agglutinative and fusional differ in separability, not density.
- Isolating (Chinese, Vietnamese): words do not change form; grammar comes from word order and particles. Chinese 我 昨天 看 了 一 本 书 “I read a book yesterday”: 看 never changes, past comes from the particle 了 (PFV) and 昨天 “yesterday”.
- Agglutinative (Turkish, Finnish): morphemes strung together, one morpheme, one meaning, visible boundaries. Turkish ev-ler-iniz-den = house + plural + your + from, “from your houses”.
- Fusional (Russian, Spanish, Arabic): one affix encodes several features that cannot be separated. Russian knig-ami “with books”: -ami = instrumental and plural at once. Spanish habl-é: person, number, tense, aspect in one vowel.
- Polysynthetic (Inuktitut, Mohawk): verb, subject, object and modifiers (negation, tense, mood, instrument, location) packed into one word. Qangatasuukkuvimmuuriaqalaaqtunga “I will have to go to the airport”.
- The axis: isolating and polysynthetic are the two ends of a scale of how much meaning one word carries. Agglutinative and fusional sit in the middle and differ on a separate question: can the packed meaning be cut apart?
Losing marks: defining agglutinative as “more morphemes”. It means separable morphemes. Fusional languages can be just as dense; defining the classes by density misclassifies Spanish. Also: claiming -ami marks feminine gender. It is the instrumental plural for all three genders.
Why do fixed word-level vocabularies break for morphologically rich languages and for multilingual models, and what is the practical fix?
Key points: fixed-size matrices and softmax make drive memory; inflection multiplies forms and productivity keeps the vocabulary open; either explodes or OOV soars with every rare form as
<unk>; a shared multilingual vocabulary fragments rich languages; fix: subword tokenization with zero OOV.
- Neural weight matrices are fixed size; the output layer is a matrix plus a softmax over at every position, so drives memory.
- Realistic sizes: English about 200K words, Russian about 1M, because Russian marks case on every noun so each lemma has many more surface forms.
- Inflectional features multiply: forms per lemma = product of slot sizes. Productivity (doomscrolling, doomscrolled, doomscrolls, doomscroller) means no vocabulary can be closed, so OOV never goes away.
- Two horns: keep all forms and explodes (embedding matrix and softmax with it); or use a vocabulary sized for English and get massive OOV rates, with every rare form mapped to
<unk>and given equal probability.- Word indices hide relations (move 2863 vs moved 87542) and give nothing for unseen words (hammered). Morphology could build meaning for unseen words; context alone cannot.
- Multilingual: a shared vocabulary leaves English words whole while Finnish gets fragmented into meaningless pieces, so sentences cost different numbers of tokens: a bias built into the tokenizer before training.
- Fix: morphological analysers are language dependent and need per-language expertise, so the practical compromise is subword tokenization: statistically useful pieces, language-agnostic, learned from data, zero OOV by falling back to characters.
Run Forward and Backward Maximum Matching on 研究生活 with vocabulary {研究, 研究生, 生活, 生, 活}. Why do the results differ, how does bidirectional matching resolve it, and what are the method's limits?
Key points: FMM gives 研究生 | 活; BMM gives 研究 | 生活; genuine ambiguity leaks through the greedy scan direction; bidirectional: accept if they agree, else prefer fewer words or fewer single-character words; limits: fixed vocabulary, OOV silently split, no learning from data.
- FMM: at , 研究生活 is not in , 研究生 is (3 chars), emit it. At , 活 is in . Result 研究生 | 活.
- BMM: at , the longest suffix in is 生活 (2 chars), prepend it. At , 研究 is in . Result 研究 | 生活.
- Why: same string, same dictionary, only the scan direction changed. The string is genuinely ambiguous and a greedy rule lets the ambiguity leak through. Neither algorithm is buggy.
- Bidirectional: if FMM and BMM agree, accept with high confidence. If not, tie-break: fewer total words, or fewer single-character words. Here both give 2 words (tie), and BMM wins on single-character words (0 against FMM’s 1 for 活).
- Limits: needs a precompiled per-language vocabulary; new words (names, neologisms, typos) can never be matched; cannot improve from data; OOV words are silently chopped into single characters; genuine ambiguity cannot be fixed by any direction heuristic.
Reformulate Chinese word segmentation as sequence labelling and explain how an HMM with Viterbi solves it. Give the recursion, the complexity, and what the approach still cannot do.
Key points: BMES tags with a boundary after every E and S; hidden tags emit characters, transitions and emissions estimated by MLE; Viterbi recursion with backpointers; against ; first-order, no long-range context (CRFs), unsupervised training worse.
- BMES tags: B (first char of multi-char word), M (interior of 3+ char word), E (last char of multi-char word), S (single-char word). Boundary after every E and S.
- HMM: hidden tag sequence generates the observed characters. Transitions , emissions . Decode .
- Training: MLE counts from a segmented corpus: , .
- Viterbi: , with a backpointer to the winning ; backtrace from the best final cell.
- Complexity: , i.e. for 4 tags, against brute-force tag sequences.
- Gains over maximum matching: no fixed dictionary, ambiguity resolved by probabilistic evidence, trainable from any segmented corpus.
- Still cannot: first-order Markov assumption (higher orders cost parameters); no longer-range context (motivates CRFs); unsupervised Baum-Welch (EM) training degrades quality. Neural sequence labellers (LSTMs, Transformers) are used in practice.
Distinguish inflection, derivation and compounding with examples, and explain why productivity of these processes matters for NLP.
Key points: inflection keeps POS, changes features (see, saw); derivation changes POS and/or meaning (happy, happiness); compounding joins words (Tischbein); productivity means no vocabulary can be closed, so OOV never goes away.
- Inflection: keeps the part of speech, changes grammatical features (tense, number). see, saw, seen; book, books. Another form of the same dictionary entry.
- Derivation: changes POS and/or meaning, giving a new dictionary entry. happy → happiness (noun), happy → happily (adverb); perfect → imperfect (meaning only).
- One-line test: does the part of speech survive? see / saw inflection, happy / happiness derivation.
- Compounding: several words into one. German Tisch + Bein = Tischbein “table leg”; sicher + gehen = sichergehen “make sure”. German writes it as one token where English uses two.
- Productivity: the processes apply iteratively and to new combinations. doomscroll is a new compound already carrying inflection (-ing, -ed, -s) and derivation (-er). No fixed vocabulary can be closed over a productive process, which is the formal reason OOV never disappears.
What is non-concatenative morphology, and why can no string-cutting segmenter (from maximum matching to byte-pair encoding) fully handle it? Use German, Arabic and Inuktitut.
Key points: change inside the root (apophony, Buch/Bücher); circumfix ge-…-en is one morpheme in two pieces; Arabic interleaves root and pattern and omits short vowels; Inuktitut boundary rules rewrite morphemes; segmenters cut contiguous pieces, so they solve only the concatenative case.
- Concatenative: word = concatenation of morphemes; string splitting can recover them.
- Non-concatenative: the change happens inside the root. Apophony: German Buch → Bücher, Haus → Häuser (vowel change plus suffix); English foot / feet, sing / sang / sung. No cut separates “book” from “plural” in Bücher.
- Circumfix (related problem): German ge-seh-en: ge- and -en are one morpheme in two pieces, so neither stripping the prefix nor the suffix alone works.
- Arabic: consonantal root k-t-b “write” poured into a pattern ya- C1 C2 -u- C3 -u = yaktubu “he writes”. Morphemes are interleaved. Short vowels are not written, so يكتب cannot distinguish yaktubu from yuktabu “it is written”: part of the information is absent from the string.
- Inuktitut: phonological rules at boundaries rewrite morphemes (-k-mut-uq surfaces as -mmuu-), so the word cannot be split back into its morphemes by string matching.
- Every segmentation method cuts a string into contiguous pieces, so subword segmentation is a solution to the concatenative case only.
Link to originalWord segmentation and subword segmentation are often confused. Define each, give a language where each is needed, and say how they are solved.
Key points: word segmentation finds boundaries in unmarked scripts (Chinese, Thai); solved by maximum matching or BMES labelling (HMM, CRF, neural); subword segmentation splits delimited words in rich languages (Finnish); solved statistically, e.g. BPE; same cutting problem, different target units.
- Word segmentation (tokenization): establishing word boundaries in scripts that do not mark them (Chinese, Japanese, Thai, Burmese). Must happen before POS tagging, parsing or anything else. Solved by maximum matching (FMM, BMM, bidirectional) or statistical BMES labelling (HMM plus Viterbi, CRFs, neural taggers). Benchmarked in the SIGHAN bake-offs.
- Subword segmentation: splitting words that are already well delimited into units smaller than the word, because morphologically rich languages (Finnish) blow up fixed vocabularies. Solved statistically, e.g. byte-pair encoding.
- Both decide where to cut a character string; they differ in what the pieces are meant to be: whole words against units smaller than words.
L04 Subword Segmentation
Exam questions
Compare BPE and Unigram LM as subword segmentation methods: mechanism, segmentation quality and downstream effect.
Key points: BPE bottom-up, greedy, one segmentation per word; Unigram LM top-down likelihood pruning, global best, n-best and sampling; Unigram LM closer to gold segmentations (CELEX2 F1 19.3 against 30.3); downstream about one point in English, 12.3 on Japanese TyDi QA; SentencePiece is the toolkit, not the algorithm.
Must hit three layers:
- Mechanism. BPE (Sennrich et al., 2016) is bottom-up: start from characters and greedily merge the most frequent adjacent pair until the merge budget is spent. A merge is never undone, and each word has exactly one segmentation. Unigram LM (Kudo, 2018) is top-down: start from a huge seed vocabulary of substrings, fit a unigram LM by EM, and prune the tokens whose removal costs the least likelihood until the target size is reached. It scores whole segmentations, , takes the global best by dynamic programming, and can return n-best or sampled segmentations.
- Segmentation quality (Bostrom and Durrett, 2020, same data and same vocabulary size): F1 against CELEX2 English morphology is 19.3% (BPE) against 30.3% (Unigram LM); against MeCab Japanese word segmentation, 73.8% against 77.2%. Unigram LM recovers suffixes (
ly,ed,s); BPE leaves word-initial single capitals and frequency artefacts such asn|an|ote|chn|ology.- Downstream (identical pretraining and model, only the tokenizer differs): Unigram LM wins the English tasks by roughly one point (SQuAD 1.1 EM 80.6 to 81.8) and Japanese TyDi QA by 12.3 EM and 12.3 F1 (EM 41.4 to 53.7).
- Vocabulary use: Unigram LM produces longer segments on average and uses its vocabulary more effectively; BPE spends the bottom of its vocabulary on near-useless tokens.
Losing marks: calling SentencePiece the algorithm (it is the toolkit that implements both); reading the low absolute English F1 as “tokenizers fail at English” when the meaningful number is the ratio between the methods.
Why does neural machine translation need subword segmentation? Argue from the weaknesses of word-level, character-level and rule-based (FST) alternatives.
Key points: vocabulary bottleneck, representation depends on frequency and contexts; word-level: huge vocabularies, no sharing, unseen words become
<UNK>; character-level: meaningless units, long sequences; FSTs: laborious, language dependent, no ranking of ambiguities; subwords: 30K to 128K, near-zero OOV, learned from frequencies.
- The vocabulary is one of the main bottlenecks of training an NMT system. How well a word is represented depends on its frequency and on the number of different contexts it occurs in, both properties of the training corpus.
- Most new surface forms are word formations: inflection (
tall,taller,tallest) and compounding (Donaudampfschifffahrtsgesellschaft). A word-level vocabulary treatstallandtallestas unrelated integers.- Word-level: 300K to 500K entries and still OOVs, huge embedding and softmax matrices, every unseen word collapses to
<UNK>(“Kyrgyzstan” in a test sentence becomes an unspecified place).- Character-level: near-zero OOV, but each symbol carries almost no meaning and sequences are long and slow.
- Hand-built FSTs: accurate, but extremely laborious, totally language dependent and give no preference among ambiguous analyses, so they do not scale to a hundred languages.
- Subwords (typically 30K to 128K): learned from frequencies with no annotation, vocabulary size tunable, near-zero OOV, moderate sequence length. The price is that pieces are not clean morphemes.
Train BPE on the corpus
the tall man is taller than the tallest manfor 10 merges, breaking ties by first occurrence. Give the merges and comment on the result.Key points: six-way tie at 3,
t+hseen first;thandoes not count towardst+a; correct merge order ending intall+e; finaltalle randtalle s t; greedy merge destroystall|erand is never undone.
- Word types:
the2,tall1,man2,is1,taller1,than1,tallest1. Split each into characters and append</w>.- Initial counts: six pairs tie at 3:
t+h(2 fromthe, 1 fromthan),t+a,a+l,l+l,a+n,n+</w>.t+his encountered first, so it wins.- Merges in order:
t+h(3),t+a(3),ta+l(3),tal+l(3),a+n(3),an+</w>(3),th+e(2),the+</w>(2),m+an</w>(2),tall+e(2).- Final corpus:
the</w>,tall </w>,man</w>,i s </w>,talle r </w>,th an</w>,the</w>,talle s t </w>,man</w>.- Comment: the frequent words
theandmanbecame single tokens, which is desired. Buttallerandtallestended astalle + randtalle + s + twhere the morphological split istall + erandtall + est. The greedytall + emerge was locally the most frequent move and destroyed the morpheme boundary, and BPE can never undo a merge.Losing marks: counting
thantowardst+a(inthanthetis followed byh).Segment
catsby bottom-up DP with p(c)=.02, p(a)=.04, p(t)=.03, p(s)=.05, p(ca)=.002, p(at)=.0015, p(ts)=.0004, p(cat)=.003, p(ats)=.00005, p(cats)=.0001.Key points: each cell initialised with the unsplit probability;
tsandatssplit,catstays whole; givescat | s, ; rare whole word loses to common pieces; valid because segments are independent.
- Length 1: is the character probability: 0.02, 0.04, 0.03, 0.05.
- Length 2:
castays whole, 0.002 (split gives 0.0008);atstays whole, 0.0015 (split 0.0012);tssplits, , so .- Length 3:
catstays whole, 0.003 (splits give 0.00003 and 0.00006);atsis best asat|s, , beating unsplit 0.00005 anda|ts0.00006, so .- Length 4: unsplit
cats0.0001; : ; : ; : . Winner , .- Read back: splits into and ; is unset, so the result is
cat | swith .- Points to make: the rare whole word
cats(0.0001) loses to two common pieces, which BPE cannot do once it has mergedcats. The morphologically correct split appears without any morphology in the model, because pluralsis very frequent. Combining the best scores of two halves is valid only because the unigram model treats segments as independent.Explain fertility and why one subword vocabulary serves languages unequally. Include byte-level BPE in your answer.
Key points: fertility = subword tokens / words; drivers: vocabulary size, word length, morphological richness; four consequences: reassembly, no universal scheme, API cost, context; byte-level BPE: zero OOV but multi-byte scripts get shorter subsegments; report fertility per language.
- Fertility = subword tokens / words, the average number of segments per word. 1.0 means every word is one token. English: 1.343 (BPE) and 1.318 (Unigram LM).
- Drivers: merge count or target vocabulary size (the one factor you control), average surface word length, morphological richness. Isolating Latin-script languages are close to 1 for common words; agglutinative languages (Finnish, Turkish, Hungarian) have many more, individually rarer word forms, which get split further.
- Consequences: (1) the model must reassemble meaning from many small parts; (2) no tokenization scheme dominates across all languages and scripts, and a multilingual vocabulary is a compromise for all of them, hence the push to 128K; (3) commercial APIs price by token, so high-fertility languages pay more for the same content; (4) fixed context windows fill up faster.
- Byte-level BPE: the 256-byte base alphabet guarantees zero OOV, but UTF-8 uses 1 byte for ASCII, 2 for Latin with diacritics, Cyrillic and Greek, 3 for most CJK, Devanagari and Thai, and 4 for many emoji. Those scripts spend merges just getting from bytes back to characters and end up with shorter subsegments. The cause is UTF-8 encoding length, which has nothing to do with the language itself.
- Methodology: when comparing across languages, measure and report fertility per language, because sequence length is a confound.
Link to originalWalk through Kudo's (2018) Unigram LM vocabulary construction algorithm and correct the pruning rule in the version printed as Algorithm 2 in MNLP.
Key points: seed with all substrings occurring more than once, not crossing words; fit unigram LM by EM, segmentation latent; = likelihood drop without ; prune at most a fraction per round, refit; correction: remove smallest , never prune single characters.
- Seed with all substrings that occur more than once in and do not cross word boundaries (millions of candidates: this is the top-down start).
- While : fit a unigram LM to by EM (segmentation is latent: the E-step computes expected subword counts over all segmentations, the M-step renormalises them into ).
- For each token : , where is the LM without and strings that used are re-segmented.
- Remove tokens, a hyperparameter, then refit. Pruning is gradual because each is computed assuming all other tokens are still present.
- Fit the final unigram LM and return and .
The error: the printed version removes the tokens with the highest . Removing a token cannot raise the likelihood, so and a large marks a valuable token: taken literally the printed rule prunes the most useful pieces first, and it is wrong. Correct rule: remove the tokens with the smallest (equivalently, keep the top-scoring pieces), which is Kudo’s intent and what SentencePiece implements. Also, Kudo never prunes single-character tokens, so every string stays segmentable; the printed lines would allow a character to be removed.
L05 Static Embeddings
Exam questions
Does unsupervised bilingual lexicon induction work? Answer with the mechanism, the headline result and the conditions under which it fails, with numbers.
Key points: shared space hypothesis and the MUSE pipeline (adversarial, CSLS selection, Procrustes refinement); matches supervised on close pairs; falls short on distant pairs; collapses for case-rich mixed-marking languages, across domains and across algorithms, not because of data size; identical-word baseline beats it.
Must hit:
- Mechanism. The shared space hypothesis says separately trained embedding spaces are approximately isomorphic, so one linear (orthogonal) map should align them. MUSE (Conneau et al., 2018) bootstraps a rough adversarially (a linear generator against an MLP discriminator, with re-orthogonalisation), picks the checkpoint with the unsupervised CSLS criterion, then refines with Procrustes () on a synthetic dictionary of mutual CSLS nearest neighbours, repeated.
- Headline result. Fully unsupervised matches or beats supervised Procrustes-CSLS on close, well-resourced pairs (P@1): en-es 81.7 against 81.4, es-en 83.3 against 82.9, en-fr 82.3 against 81.1, en-de 74.0 against 73.5.
- Where it falls short. Distant pairs: en-ru 44.0 against 51.7, en-zh 32.5 against 42.7.
- Where it collapses (Søgaard et al., 2018). Mixed- or double-marking, case-rich languages: EN-ET 0.00, EN-FI 0.09, EN-EL 0.07. EN-FI stays at 0.0 even when retrained on 1.7 billion Finnish words, so data size is not the explanation. Mismatched domains: 0.0 to 0.13 in every cross-domain en-es cell. Mismatched algorithms: Spanish CBOW against English skipgram gives 0.00 to 0.13.
- A free baseline beats it. A seed dictionary of identically spelled words beats adversarial training on every pair involving English (EN-ES 82.62 against 81.89, EN-ET 31.45 against 0.00).
- Verdict. It works when the languages are typologically similar, the corpora share a domain and both sides use the same embedding algorithm. Those are the conditions it was evaluated under, and they are the opposite of the low-resource setting that motivates it.
Losing marks: giving only the Conneau et al. headline numbers, or blaming the failures on data size alone.
Derive the negative sampling objective from a binary classification set-up, then derive its gradient with respect to the input vector and interpret it.
Key points: real-or-noise binary classification with negatives; and ; log objective over and ; gradient pulls to , pushes it from negatives; cost instead of .
- Why. A full softmax needs a score for all words per training example. Replace “which word is the context?” with “is this (word, context) pair real or noise?”: 1 positive pair from the data, negative pairs from a noise distribution.
- Likelihood. , with the observed pairs, the noise pairs, and , where (column of ) and (row of ).
- Sigmoid identity. , so the objective becomes , and after logs
- Per-example loss. .
- Gradient. With :
- Interpretation. The positive coefficient is negative, so a descent step moves towards ; each negative coefficient is positive, so moves away from . Both shrink to zero once the classifier gets the pair right, so the biggest corrections come from negatives the model mistakes for real contexts. Cost per example drops from dot products to .
Compare Word2Vec (CBOW and Skip-gram) with GloVe: what each optimises, what information each uses, how each handles frequency imbalance, and what the final embedding is.
Key points: CBOW and Skip-gram predict within local windows; GloVe is weighted least squares on global log counts; subsampling against ; final embedding or against .
- CBOW: predict the centre word from the average of context embeddings, , , cross entropy, SGD.
- Skip-gram: predict each of the context words from the centre word, , context words assumed independent given ; in practice trained with negative sampling.
- GloVe (Pennington et al., 2014): weighted least-squares regression on log co-occurrence counts, .
- Information used: Word2Vec sees individual local context windows, treating each as an independent event, so it keeps “rediscovering” the same association. GloVe counts once over the whole corpus (global statistics) and then fits gradient updates to those counts.
- Frequency imbalance: Word2Vec subsamples frequent words (and draws negatives from ); GloVe uses the explicit weighting with , .
- Interpretability: Word2Vec is indirect (it learns to predict contexts); GloVe directly fits log counts.
- Final embedding: Word2Vec typically , or ; GloVe . Both sum the two vector sets.
- Context: prediction-based methods tend to outperform count-based ones; count-based ones use global information directly. GloVe is the hybrid.
Why does constraining a cross-lingual mapping to be orthogonal help? Give the objective, the two problems it fixes with a derivation for each, its closed-form solution, and the evidence.
Key points: with unit vectors; rigid map stops overfitting the seed pairs; aligns training with retrieval; by SVD; unconstrained degrades with dimension, orthogonal better everywhere.
- Objective (Xing et al., 2015). subject to , with all embeddings normalised to unit length. An orthogonal only rotates and reflects: no scaling or shearing.
- Problem 1, overfitting. An unconstrained can stretch, shear and rotate, so it can warp the space to fit the seed pairs and fail to generalise to the words outside the dictionary (the majority). Orthogonal preserves geometry: , so the source space moves rigidly, which is exactly what the shared space hypothesis says should suffice.
- Problem 2, objective mismatch. Training minimises Euclidean distance, retrieval uses cosine. For unit vectors and orthogonal : , so minimising distance is maximising cosine.
- Solution. , : one SVD of a matrix, exact, no learning rate.
- Evidence (en-es, P@1). Unconstrained (Mikolov et al., 2013) falls from 30.43% at 300 dimensions to 20.69% at 700, because has free parameters for the same dictionary. Orthogonal is better everywhere and rises slightly, 38.99% to 41.04%.
Losing marks: claiming Euclidean training and cosine retrieval always agree (they coincide only for unit vectors under an orthogonal map), or calling the unconstrained least-squares map “Procrustes” (in the results tables Procrustes means the orthogonal SVD solution).
Explain the hubness problem in cross-lingual word retrieval and how CSLS corrects it, including which term of the formula does the work.
Key points: hubs are nearest neighbours of many points, worse in high dimensions; ; penalises hubs; constant per source word; better retrieval, only to set.
- Hubness: in high-dimensional spaces a few vectors become the nearest neighbour of many points regardless of real similarity. It is a general property of high-dimensional geometry and worsens with dimension. A hub sits near the centre of the cloud and is moderately close to everything.
- Effect on translation: one target hub becomes the nearest-neighbour “translation” of many unrelated source words: nn(cat) = nn(car) = nn(house) = thing. Nearest neighbour is asymmetric:
thingcan be the neighbour ofcatwithoutcatbeing the neighbour ofthing.- CSLS (Conneau et al., 2018): , with , the mean cosine of to its nearest target vectors, the mean cosine of to its nearest mapped source vectors.
- Which term works: penalises a candidate that is close to lots of things (a hub has high ). is constant across candidates for a fixed , so it cannot change which wins; it matters when scores are compared across source words, e.g. ranking pairs to build a dictionary.
- Effect: significantly better retrieval with no parameter tuning beyond (Conneau et al. use ). Over plain NN on the same Procrustes map: en-es 77.4 to 81.4, fr-en 76.1 to 82.4.
Describe the full MUSE pipeline of Conneau et al. (2018) for aligning two embedding spaces without any bilingual signal, and justify each design choice.
Key points: linear generator against MLP discriminator; linear so point correspondence follows the distribution match; label smoothing, 50k frequent words, re-orthogonalisation; CSLS-based unsupervised model selection; Procrustes refinement on mutual nearest neighbours.
- Adversarial step. Generator = the mapping (a single linear matrix); discriminator = an MLP (2 hidden layers of 2048, ReLU, input dropout) that tells mapped source vectors from real target vectors . They train in alternation: discriminator step with frozen, generator step with labels flipped and the discriminator frozen.
- Linear generator, deliberately. A deep generator could match the target distribution while scrambling which source word lands where. Only a constrained, near-rigid map forces point-level correspondence to come with the distribution match.
- Label smoothing . Targets stop the discriminator becoming overconfident, which would give the generator vanishing gradients.
- Only the 50k most frequent words are fed to the discriminator: rare words have poorly estimated embeddings, so frequent words give a cleaner signal.
- Re-orthogonalisation after each step, with , keeps near orthogonal so it cannot drift into a warping general linear map.
- Model selection without a dictionary. The adversarial loss is unreliable. Instead: for the 10k most frequent source words, find each one’s CSLS nearest target, average the cosines, keep the checkpoint with the highest value.
- Refinement. Build a synthetic dictionary of high-confidence mutual CSLS nearest neighbours among frequent words, re-solve by SVD (Procrustes), repeat. Translate with CSLS.
- Contribution of each part (en-es P@1): adversarial alone 69.8, + CSLS 75.7, + refinement 79.1, both 81.7.
What is fastText, how does it differ from Skip-gram with negative sampling, and what does the evidence of Bojanowski et al. (2017) say about where it helps and where it does not?
Key points: sum of character n-grams of
<w>(3 to 6) plus the word; same SGNS loss, only the centre word decomposed; vectors for OOV words; gains on morphologically rich languages, rare words and syntactic analogies; no gain on semantic analogies.
- Problem it addresses: Word2Vec and GloVe give one opaque vector per word type, so
run,runs,running,runnershare no structure; morphologically rich languages (Finnish, Turkish, Russian) split each lemma into many rare forms; OOV words get no vector.- Model: the centre word vector is the sum of its character n-gram vectors, , where holds all n-grams of
<w>for plus the whole word<w>.<and>mark word boundaries so prefixes and suffixes are distinct n-grams.- Same as SGNS: the loss , with context vectors still per word.
- Different: every n-gram of gets the same gradient (since ); the output is the n-gram table , and any word, seen or not, gets a vector by summing its n-grams.
- Word similarity: sisg (fastText) best or tied-best on 9 of 10 datasets; biggest gains in morphologically rich languages (Russian HJ 59 to 66 over sg); smaller gains on English Rare Words (43 to 47). Building OOV vectors from n-grams (sisg against sisg-) adds up to 6 points. Only loss: English WS353 (71 against cbow’s 73), frequent words where whole-word vectors are already good.
- Analogies: gains are syntactic (Czech 52.8 to 77.8, German +11.9, Italian +11.2, English +4.8 over sg). Semantic analogies do not improve and sometimes drop (German 66.5 to 62.3), because capital-country relations have nothing to do with character overlap.
How can the degree of isomorphism between two embedding spaces be measured without a dictionary, what did it show for English and Finnish, and what can it not show?
Key points: Laplacian eigenvalues of nearest-neighbour graphs; = sum of squared eigenvalue differences; no dictionary needed; high for EN-FI, EN-ET, EN-EL and tracks adversarial failure; low does not prove isomorphism.
- Why a measure is needed: isomorphism is a strict true/false criterion; Søgaard et al. (2018) want a degree.
- Eigenvector similarity: build a nearest-neighbour graph per language (adjacency matrix ), take the degree matrix , form the Laplacian , keep the largest of its eigenvalues, and compute .
- Reading it: larger Laplacian eigenvalues mean denser connectivity; larger means more structurally different graphs. Eigenvalues are invariant to relabelling the nodes, so no alignment is needed.
- Results: EN-ES 2.07 is the lowest; the three pairs where adversarial alignment fails completely have the largest values (EN-ET 6.61, EN-FI 7.33, EN-EL 5.01). tracks adversarial success.
- Finnish explanation: a lemma spreads across dozens of inflected forms that cluster tightly by meaning overlap, while English gives a flatter, more evenly connected graph. Different connectivity means high , and no rotation can fix it, since a rotation preserves the neighbour graph and so the eigenvalues.
- Limit: near-isomorphic graphs have low , but a low does not imply near-isomorphism (non-isomorphic cospectral graphs exist). can rule isomorphism out, never in.
Losing marks: concluding that a low proves the spaces are isomorphic.
How should word embeddings be evaluated? Cover the evaluation axes, the word similarity and analogy tasks, and the limits of 2D visualisations.
Key points: intrinsic against extrinsic, qualitative against quantitative; word similarity by Spearman, WS-353 relatedness against SimLex-999 similarity; analogies via scored by accuracy; 2D projections distort, t-SNE hyperparameters matter.
- Two independent axes: intrinsic (how good the embeddings are by themselves) against extrinsic (usefulness in downstream tasks such as summarisation, MT, IR); qualitative (inspect selected examples, neighbours, plots) against quantitative (an overall task-dependent score). A 2D plot is intrinsic and qualitative; a similarity correlation is intrinsic and quantitative; BLEU after plugging embeddings into MT is extrinsic and quantitative.
- Word similarity: rank word pairs by cosine, compare with human judgements by Spearman correlation. WS-353 annotates relatedness, SimLex-999 similarity;
coffee/cupis related but not similar, so the two benchmarks can rank the same embeddings differently. Both mix kinds of similarity.- Analogies: a : b :: c : X, answer the word closest by cosine to (input words excluded), scored by accuracy; semantic and syntactic sets.
- Visualisation: project to 2D with PCA (linear, top-variance directions) or t-SNE (non-linear, keeps local neighbours, gives up global distances). Useful for showing sense clusters or a consistent gender direction.
- Caveats: non-linear projections distort distances; t-SNE hyperparameters change cluster sizes, distances and even apparent clusters (Wattenberg 2016). Plots complement quantitative and extrinsic evaluation and do not substitute for it.
Link to originalDerive the closed-form solution of the unconstrained linear mapping , state when it exists, and explain why this mapping gets worse as the embedding dimension grows.
Key points: loss ; zero gradient gives ; needs invertible, ; SGD as the scalable alternative; parameters overfit, so accuracy falls with dimension.
- Matrix form. Stack the dictionary pairs as columns: is , is . Loss .
- Gradient. (per pair: ).
- Set to zero. , so .
- Existence. () must be invertible, which needs at least linearly independent source vectors, i.e. pairs.
- Closed form against SGD. Closed form is exact; SGD (used by Mikolov et al., 2013) is approximate, scales better with large dimensions and dictionaries, and never forms or inverts .
- Dimension. has free parameters for the same seed dictionary, so more dimensions means more room to overfit: en-es P@1 drops from 30.43% (300 dimensions) to 25.76% (500) to 20.69% (700). The unconstrained map can stretch and shear to fit the anchors and generalises poorly.
L06 Contextual Embeddings
Exam questions
Why can mBERT transfer a task from one language to another when nothing in its training is cross-lingual? Argue from the experimental evidence.
Key points: only parameters and subword vocabulary are shared; vocabulary overlap is a minor cause (cross-script transfer, fake English costs about a point); structure and word order matter; depth matters; no single factor, and mixed-language inputs fail.
- What mBERT lacks. Same architecture and loss as BERT (MLM + NSP) on 104 Wikipedias; no parallel data, no language-ID embedding, no language-specific parameters, no alignment loss, NSP pairs always within one language. Languages share only the parameters and the subword vocabulary.
- Vocabulary overlap is a minor cause. Pires et al. (2019): mBERT transfers across scripts (Hindi to Urdu POS 85.9, English to Bulgarian 87.1), and its zero-shot NER F1 is roughly flat in entity-wordpiece overlap, while English BERT’s F1 is near 0 at low overlap and rises with it. K et al. (2020): fake English removes all subword overlap and costs only 0.5 to 1.4 XNLI points.
- Structure matters. Transfer is best within the same word-order type (SVO to SVO 81.55, SVO to SOV 66.52) and rises with the number of shared WALS features. Permuting all words during pre-training drops XNLI by 8.4 (Spanish), 16.5 (Hindi) and 12.1 (Russian), yet transfer stays well over chance.
- Depth matters. At roughly constant parameter count, the gap between fake-English and Russian XNLI shrinks from 21.6 (1 layer) to 11.3 (24 layers): deeper networks learn more language-independent representations.
- Bottom line: no single factor explains transfer, or the lack of it.
- The limit. When premise and hypothesis are in different languages, accuracy falls under both monolingual settings, so the shared space is not truly language-neutral. Rajaee and Monz hypothesise that transfer of heuristics (such as premise-hypothesis word overlap) contributes to cross-lingual generalisation; mixed-language inputs remove that shortcut.
Losing marks: calling shared wordpieces the main cause; claiming mBERT saw parallel data or an alignment objective.
Explain why static word embeddings are inadequate and how contextual embeddings fix this. What design tension do contextual models face, and how can it be managed?
Key points: one vector per word type (); contextual models give each occurrence its own vector; tension: enough context to separate senses without losing the word’s own contribution; managed by choice of layer, choice of task, residual connections.
- Static embeddings, whether a by-product of a task (embedding layer of a classifier or MT system) or trained directly (Word2Vec with negative sampling, limited window), give one vector per word type: for the vehicle, the verb “to practise” and “train of thought”. Three senses, two parts of speech, one vector.
- In the by-product route the network did see the whole sentence, but only the input embedding layer was kept, so the context computed in deeper layers was thrown away.
- Contextual embeddings keep it: a deep network reads the whole sentence and every occurrence gets its own vector (a hidden state in that word’s column).
- Tension: integrate enough context to distinguish occurrences of the same word, without capturing so much that the word’s own contribution becomes unclear (every position becomes “the meaning of the sentence”).
- Managed by: (1) choosing the right layer(s): lower layers are closer to the word, higher closer to the sentence; (2) choosing the right training task(s): per-token outputs keep each top state about its word, a single pooled output pushes the top layer toward the sentence; (3) tightening connections between layers, e.g. residual connections, which carry the word’s own information upward alongside the added context.
Describe how BERT is pre-trained and fine-tuned: input representation, both pre-training objectives, the role of
[CLS], and how task heads attach.Key points: input = token + segment + position embeddings; MLM: 15% selected, 80/10/10, loss on selected positions only; NSP: 50% IsNext, 50% random, classified from
[CLS]; head on[CLS]for classification, on token outputs for tagging/spans; short fine-tuning, BERT optionally updated.
- Model: encoder part of the Transformer, bidirectional (every position attends left and right). Devlin et al., NAACL 2019.
- Input:
[CLS] A [SEP] B [SEP], with (token + segment A/B + position embedding). WordPiece tokens.- MLM: select 15% of tokens; replace 80% with
[MASK], 10% with a random token, leave 10% unchanged; loss on the selected positions only, predicting the original token.- NSP: 50% B really follows A (IsNext), 50% B is random (NotNext); classified from , the final-layer output at
[CLS].[CLS]: never masked, not tied to a word, represents the whole sequence, trained through a classification loss.- Fine-tuning: small randomly initialised head; connect it to the token outputs for tagging or span extraction, to for classification; train on the task with or without updating BERT. Batch 16 or 32, 2 to 4 epochs.
- Self-supervised pre-training can use huge amounts of task-irrelevant data; supervised fine-tuning then works with little task data.
Losing marks: saying the loss covers all tokens, or only the positions showing
[MASK]; saying BERT-base has 16 attention heads (it has 12).Describe how mBERT is trained and what it deliberately lacks. Explain its language sampling scheme and the problems a single shared multilingual vocabulary causes.
Key points: same architecture and MLM + NSP loss on 104 Wikipedias; no parallel data, language ID or alignment loss, NSP within one language; smoothed sampling with ; one shared 120k WordPiece vocabulary; favours high-resource languages and Latin script, higher fertility for low-resource languages, uncased strips accents.
- Released on GitHub in 2018 with no paper. Exact BERT architecture and loss (MLM + NSP), on the concatenated Wikipedias of 104 languages.
- Absent: parallel data (at least intentionally), language-ID embedding, language-specific parameters such as adapters, any alignment loss. NSP pairs are always in one language; no language mixing within a sequence (a batch can mix languages).
- Sampling: Wikipedia sizes are extremely skewed (English alone exceeds dozens of small languages combined), so languages are drawn with . : proportional, big languages dominate; : uniform, tiny languages are memorised; : up-weights low-resource and down-weights high-resource languages.
- Vocabulary: one shared WordPiece vocabulary of 120k (English BERT 30k), which is why mBERT has about 178M parameters against 110M. CJK characters are split into single characters before WordPiece. The sampling smoothing is also applied to the counts used to build the vocabulary.
- Problems: even smoothed, the vocabulary favours high-resource languages and Latin script; low-resource and morphologically rich languages get much higher fertility (more pieces per word); the uncased variant lowercases and strips accents, which damages languages that depend on diacritics.
Calculate: English has 1000 units of training text and Swahili 10. Using , give both sampling probabilities for , and , and interpret each.
Key points: : 0.0099 / 0.990, English dominates; : 0.038 / 0.962, Swahili up-weighted; : 0.5 / 0.5, small language memorised; ratio becomes .
- : , . Proportional to size; English dominates.
- : , , so and . Swahili gets about 3.9 times its raw share; English still gets 96%.
- : . Swahili’s share grows about 50-fold over its raw share, and each unit of Swahili text is drawn 100 times as often as each unit of English ( against ): the model would memorise the small language.
- General rule: the ratio between two languages goes from to , because with is concave and compresses large values more. Here 100:1 becomes about 25:1 at .
Compare XLM-R with mBERT: what changed in training, and how do the two compare within a language and across languages on XNLI and cross-lingual QA?
Key points: MLM only, dynamic masking, CC-100, 250k SentencePiece Unigram vocabulary; better within a language on XNLI and QA; better across languages on XNLI; worse across languages on QA (36.8 against 44.2 F1); better per language, not better at relating languages.
- XLM-R (Conneau et al. 2020) is a RoBERTa-style extension of mBERT: MLM only (NSP dropped: contributes little, sometimes hurts), dynamic masking, CC-100 (2.5TB filtered CommonCrawl, 100 languages, two orders of magnitude more data than Wikipedia), 250k Unigram LM vocabulary via SentencePiece (better fertility, bigger embedding matrix).
- Within a language XLM-R is better on both tasks: XNLI 74.2 against 65.7 (+8.5), XSQuAD F1 72.0 against 64.4 (+7.6).
- Across languages on XNLI it is also better: 64.8 against 54.5 (+10.3); gap within to across 9.4 against 11.2.
- Across languages on QA it is worse: 36.8 against 44.2 F1, a within-across gap of 35.2 against 20.2. Thai: within QA rises from 40.0 to 66.5, across barely moves (26.1 to 29.6).
- XLM-R’s QA matrix is “own language, or English question”: strong diagonal and English-question column (58.2 to 75.0), little in between. mBERT degrades more smoothly.
- Conclusion: more data and a bigger vocabulary made each language better on its own without making the model better at relating two languages inside one input.
What happens when the two inputs of one task are in different languages? Give the evidence for B-BERT, mBERT and XLM-R, and a possible explanation.
Key points: B-BERT mixed pairs fall below both monolingual conditions; mBERT across well below within (54.5 against 65.7); XLM-R worse across on QA; shared space not language-neutral; heuristics hypothesis: word-overlap shortcut fails across languages.
- Standard zero-shot transfer is English fine-tuning, then monolingual testing in another language. Cross-lingual inference within a task means e.g. an English premise with a Spanish hypothesis.
- K et al. (2020), B-BERT on XNLI: fake English pairs 78.5 to 79.3, target-language pairs 59.6 to 70.9, mixed pairs 45.7 to 61.1, under both monolingual conditions (Hindi 45.7, 13.9 points under its own zero-shot score). A fake-English hypothesis beats a fake-English premise.
- Rajaee and Monz (2024), mBERT on XNLI, 15 languages: within 65.7 against across 54.5 on average. A Swahili hypothesis is near chance (40 to 42.5).
- XLM-R: better within and across on XNLI, but across-language QA falls to 36.8 F1 against mBERT’s 44.2.
- Interpretation: the model can do the same task in another language but is poor at relating two languages inside one input, so the space is not language-neutral.
- Heuristics hypothesis: transfer of heuristics can contribute to cross-lingual generalisation. In SNLI full word overlap means entailment 94.7% of the time; a model can learn “high overlap means entailment”, which works in any single language but fails when premise and hypothesis are in different languages, where overlap is near zero. (How exactly this explains the drop is an interpretation, not an established result.)
What do layer-wise analyses reveal about where mBERT holds language-neutral information? Use translation retrieval and parameter freezing, and link to the choice of layer.
Key points: translation retrieval follows an inverted U peaking at layers 6 to 8; middle layers most language-neutral, top layers language-specific; Feat (no fine-tuning) below the best fine-tuned setting on every task; freezing lower layers helps, freezing up to layer 9 hurts; ties to choosing the right layer.
- Translation retrieval (Pires et al.): per layer, mean-pool hidden states (excluding
[CLS],[SEP]), shift English vectors by the mean EN-to-DE difference, retrieve the nearest German sentence by . Accuracy follows an inverted U: 20 to 33% at layer 1, peak 71 to 76% at layers 6 to 8, falling in the last layers.- Reading: low layers are close to language-specific surface tokens; middle layers are the most language-neutral (one constant offset maps one language onto the other); top layers are shaped by predicting words in a specific language.
- Freezing (Wu and Dredze 2019): fine-tune on English with the lowest layers frozen. Feature-based use without fine-tuning is below the best fine-tuned setting on every task average (by 1.9 to 14.8 points), though freezing up to layer 9 is even lower on NER and POS. Freezing the embeddings up to layer 3 or 6 helps or is neutral (best at layer 0/3 for NER, 3 for POS, 6 for MLDoc and XNLI), since English-only fine-tuning would pull lower layers toward English. Freezing up to layer 9 hurts (NER 67.3 against 74.3): upper layers must adapt to the task.
- Both measure the “choose the right layer” strategy for balancing word and context information.
Link to originalExplain how subword segmentation interacts with masked language modelling and with multilingual models: WordPiece, whole-word masking, and the vocabulary choices of mBERT and XLM-R.
Key points: WordPiece: count-ratio merges,
##, greedy longest match, whole word to[UNK]; suffix pieces easy to predict, so whole-word masking at the same rate, improves SQuAD and MNLI; mBERT 120k shared vocabulary biased to high-resource/Latin; XLM-R 250k Unigram LM, better fertility; overlap contributes little to transfer.
- WordPiece (BERT, mBERT): bottom-up merges by ;
##marks non-initial pieces; inference is greedy longest match; a word with an unmatchable character becomes[UNK]as a whole.- Masking problem: pieces carrying morphology (tense, number, case) are easy to predict from their visible stem, so per-token masking wastes many masked positions on trivially recoverable suffixes.
- Whole-word masking: mask all pieces of a word together at the same overall rate (about 15%); improves BERT-Large on SQuAD 1.1 (uncased F1 91.0 to 92.8) and MultiNLI (86.05 to 87.07).
- mBERT: 120k shared WordPiece vocabulary with smoothed counts, CJK split into characters, still biased toward high-resource languages and Latin script, higher fertility for low-resource and morphologically rich languages; the uncased variant strips accents.
- XLM-R: 250k Unigram LM vocabulary via SentencePiece, fewer pieces per word, at the cost of a larger embedding matrix.
- Wordpiece overlap itself contributes little to transfer (fake English costs 0.5 to 1.4 XNLI points).
L07 Crosslingual NLP
Exam questions
Why does multilingual pretraining (mBERT, XLM-R) not by itself make a model crosslingual, and what evidence shows that explicit crosslingual objectives fix it?
Key points: MLM uses only same-language context, so alignment is an emergent side effect; standard zero-shot keeps the input in one language; within against across on mixed-language XNLI and QA; XLM-R beats mBERT within but loses across on QA; InfoXLM’s TLM and XLCo close most of the gap.
Must hit:
- The mechanism. Multilingual models train jointly on many languages with one model and one subword vocabulary, but MLM predicts a word only from same-language context. When predicting word in language , unrelated context in is of no use, so nothing in the objective links languages. Any alignment is an emergent side effect.
- Why it looked fine. Standard zero-shot benchmarks (fine-tune on English, test on another language) keep all parts of the input in one language, and mBERT does reasonably there.
- The harder test. Put premise and hypothesis (XNLI) or context and question (XSQuAD) in different languages and compare within (both parts in one language) with across (mixed).
- The numbers. QA F1 within/across: mBERT 64.4/44.2, XLM-R 72.0/36.8, InfoXLM 73.8/64.5. XLM-R is much better than mBERT within languages (on average, though not in every language: English QA within is 84.5 for mBERT against 84.2) but worse across languages.
- The fix. InfoXLM adds TLM and XLCo on parallel data: within improves only +1.2 (XNLI) and +1.8 (QA) over XLM-R, across improves +5.5 and +27.7. The QA gap shrinks from 35.2 to 9.3.
Losing marks: treating “multilingual” and “crosslingual” as the same thing, or citing only standard zero-shot scores as proof of crosslingual ability.
You have English-only training data for NLI and need a system for 14 other languages. Compare translate-train, translate-test and zero-shot transfer, using the XLM results on XNLI.
Key points: translate-train translates the training set into each language (best, but costly); translate-test translates the test set into English; zero-shot fine-tunes on English only; zero-shot XLM (MLM+TLM) beats translate-test; TLM gives the gain over MLM alone.
- Translate-train: machine-translate the English training set into each language and fine-tune on the translation. XLM (MLM+TLM) average 76.7, the best of the three, but needs an MT system and a translated training set for every language.
- Translate-test: machine-translate each test example into English and apply an English model. XLM (MLM+TLM) average 74.2.
- Zero-shot crosslingual transfer: fine-tune on English only, test directly on each language. XLM (MLM) 71.5, XLM (MLM+TLM) 75.1.
- Key point: a single zero-shot model that never saw non-English NLI data beats the translate-test pipeline (75.1 against 74.2), and TLM is what gets it there (+3.6 over MLM alone).
- Zero-shot XLM (MLM+TLM) also beats mBERT and LASER in every language where they are reported.
Losing marks: mixing up which side gets translated (train set into the target language, or test set into English).
Explain translation language modeling (TLM): its input, its loss, why it forces crosslingual alignment, and how it differs from MLM and CLM.
Key points: TLM is MLM over a sentence concatenated with its translation; same loss, different input; a masked word is predictable from its translation, which forces alignment; positions restart at 0 and language embeddings mark the halves; CLM predicts the next word, MLM uses monolingual context.
- XLM (Conneau and Lample, 2019) has three objectives: CLM (predict the next word from a prefix), MLM (predict masked words from monolingual context, as in BERT), TLM (predict masked words in a sentence concatenated with its translation). TLM is always paired with MLM or CLM; NSP is dropped.
- Loss: TLM is MLM applied to : The loss is the same as MLM’s; only the input changes.
- Why it aligns: in “the [MASK] [MASK] blue” / “[MASK] rideaux étaient [MASK]”, the English context barely constrains curtains, but the unmasked French rideaux gives it away. The cheapest way to lower the loss is to attend across languages and learn rideaux ↔ curtains. Context in becomes useful for predicting in .
- Two input details: position embeddings restart at 0 in the second sentence (so position cannot separate the halves, and corresponding words get similar position signals); language embeddings (en, fr) mark which half is which.
- Evidence: zero-shot XNLI average 71.5 (MLM) to 75.1 (MLM+TLM).
How would you build a parallel corpus from the web, and how do LASER embeddings make sentence alignment possible?
Key points: OPUS and natural sources (UN, EU, multilingual news and countries); document alignment then sentence alignment, or direct mining from Common Crawl; LASER: BiLSTM encoder max-pooled into one sentence vector; language-ID decoder makes the vector language-independent; vecalign aligns by LASER similarity at scale.
- Sources first: natural by-products (multilingual news such as Xinhua, the UN and EU, websites of multilingual countries such as Canada and Belgium) and OPUS, a large research collection of parallel corpora.
- Crawl pipeline: (1) document alignment (find parallel documents), (2) sentence alignment inside them. Or mine sentences directly from raw crawls such as Common Crawl.
- Why sentence alignment is separate: segments do not follow paragraph boundaries (in the NHK swine fever example, two English paragraphs map into one Chinese paragraph, which must be split at a sentence boundary).
- Measuring equivalence: old methods use dictionary overlap and relative length; the modern method compares LASER sentence embeddings (Artetxe and Schwenk, 2019).
- LASER: BPE embeddings, stacked BiLSTM encoder, max pooling into one sentence vector; an LSTM decoder translates from that vector alone and is told the output language by a language ID embedding. So the vector must encode meaning and has no reason to encode the input language. Translations end up as near neighbours.
- vecalign uses LASER similarities to match sentences, with an efficient search that scales to massive data sets.
Explain neural machine translation as conditional language modeling, from the RNN encoder-decoder to the Transformer encoder-decoder.
Key points: seq2seq drops and ; conditional LM ; decoder input is the target shifted by one; RNN is an information bottleneck; Transformer cross-attention over all encoder outputs removes it.
- Seq2seq drops two assumptions of sequence labeling: that corresponds to , and that . MT needs both dropped (Hiermit hörte sie nicht auf / She did not stop with this: different lengths, reordering, one-to-many, many-to-one).
- It is conditional language modeling: , with the encoder output and the decoder’s prefix. The sequence probability is the product of these terms over .
- Data: source ids ; decoder input =
<s>+ sentence; target = sentence +</s>( shifted by one).- RNN encoder-decoder (Sutskever et al., 2014, LSTMs): . Everything about the source must pass through one fixed-size vector: an information bottleneck.
- Transformer: each decoder layer has a target context layer (masked self-attention over the target prefix), a source-target context layer (attention over all top-layer encoder outputs), and a feed-forward layer, each with a residual connection. Every decoder position in every layer can read every source position, which removes the bottleneck.
Describe parent-child transfer learning for low-resource NMT. What should be frozen, and which factors decide whether transfer pays off?
Key points: train a parent on a high-resource pair, continue on the child with the same target language; freeze only the target embeddings, train the rest; related parents help, but structure transfers too (French’); transfer must flow from the larger corpus to the smaller; inner layers carry most of the benefit.
- Procedure (Zoph et al., 2016): train a parent on a high-resource pair (French→English, 300M English tokens); copy all parameters into the child (Uzbek→English, 1.8M tokens); child source words take over rows of the parent’s source embedding matrix; freeze the English embeddings; continue training with strong regularisation (dropout 0.5).
- Gains: Hausa +4.5, Turkish +5.6, Uzbek +3.7, Urdu +8.6 BLEU; the smallest corpus (Urdu) gains most.
- Freezing: train everything up to and including attention (Uz→En dev 15.0), but keep the target embeddings frozen (unfreezing them drops to 14.7, then 13.7). The general rule “freeze the decoder” overstates it: training the target RNN raises 11.8 to 14.2.
- Relatedness: a Spanish child gets 31.0 with a French parent, 29.8 with German, 16.4 with none. But French’ (scrambled vocabulary) still gains 13.3 to 20.0, so structure transfers and shared words are not the only factor.
- Shared vocabulary (Kocmi and Bojar, 2018): unrelated parents (Czech, Russian) help English→Estonian as much as related Finnish.
- Direction: transfer must flow from the larger corpus to the smaller; reversed, it helps little or hurts.
- What transfers (Aji et al., 2020): inner layers carry most of the benefit; parent embeddings alone are worse than training from scratch.
Describe BART: its architecture, its pretraining noise functions, which noise works best, and how it is fine-tuned for classification, span prediction and translation.
Key points: bidirectional encoder plus autoregressive decoder reconstructing corrupted text; five noise functions; text infilling best, rotation and shuffling alone poor; classification from the last decoder state, start and end labels for spans; MT via a new random source encoder, only tuned BART beats the baseline.
- Architecture: bidirectional encoder (as BERT) plus autoregressive decoder (as GPT). The encoder reads a corrupted document, the decoder reconstructs the original through cross-attention. Suited to tasks needing both, such as translation and summarisation.
- Noise functions: token masking, token deletion, text infilling (a span replaced by one
[MASK]), sentence permutation, document rotation.- Best: text infilling (SQuAD 90.8, best XSum and ConvAI2 perplexity); deletion beats masking on all generation tasks; rotation and sentence shuffling alone are poor (SQuAD 77.2 and 85.4). Infilling plus shuffling gives the best CNN/DM perplexity (5.41).
- Classification: same input to encoder and decoder; label predicted from the last decoder hidden state. Span prediction (SQuAD): label each token, predict start and end of the answer.
- MT: a randomly initialised source encoder replaces BART’s embedding layer; BART can be frozen or updated; the source vocabulary can differ from BART’s. Ro→En: baseline 36.80, Fixed BART 36.29, Tuned BART 37.96.
- Results: matches RoBERTa on SQuAD (94.6 F1 on 1.1), best on all ROUGE columns for CNN/DM and XSum.
Why does crosslingual QA expose the weakness of purely multilingual models much more than XNLI does? Use the XLM-R and InfoXLM results.
Key points: XNLI can rely on a sentence-level gist; QA needs word-level matching across languages; XLM-R QA collapses unless the question is in English; InfoXLM fills the matrix; InfoXLM’s across gain is far larger on QA than on XNLI.
- XNLI compares two sentences; a coarse sentence-level gist in a shared space is often enough.
- Extractive QA requires finding the exact answer span, which means matching the question’s words to specific context words. With a Hindi question and Arabic context that is word-level matching between two non-English languages, never asked for by monolingual MLM, and exactly what TLM trains.
- XLM-R QA: diagonal strong (63.7 to 84.2), English-question column 58.2 to 75.0, but most other off-diagonal cells collapse (Arabic context with non-English question 14.6 to 36.7; Chinese context 16.0 to 32.0). Across average 36.8, under mBERT’s 44.2.
- InfoXLM QA: every cell at least 51.7; Arabic and Chinese contexts with non-English questions now 51.7 to 63.8.
- Gain of InfoXLM over XLM-R in the across score: +5.5 on XNLI, +27.7 on QA.
Which factors predict whether crosslingual transfer will succeed? Support each with evidence.
Key points: language relatedness; closeness to English; representation of the language and its script; an explicit crosslingual training signal; parent data size and direction, and which parameters transfer.
- Language relatedness: Nepali perplexity 157.2 alone, 140.1 with English, 115.6 with Hindi; Spanish NMT child 31.0 with a French parent against 29.8 with German.
- Being close to English: English dominates pretraining and fine-tuning data, so the English column is the brightest in every XNLI heatmap and English cells are the darkest in the WikiMatrix BLEU grid; directions between two non-English languages are mostly 5 to 25 BLEU.
- Representation of the language and its script: Swahili hypotheses leave mBERT barely over chance (40.2 to 42.5); mBERT cannot handle Thai questions in QA (18.8 to 23.4), while XLM-R has no such Thai problem.
- An explicit crosslingual training signal: TLM and XLCo on parallel data close most of the within/across gap (InfoXLM).
- Amount and direction of data in NMT transfer: transfer helps when it flows from a large parent to a small child; reversed it helps little or hurts. With a shared vocabulary and a large parent, relatedness matters less (Kocmi and Bojar).
- Which parameters transfer: inner layers carry most of the benefit (Aji et al.).
Link to originalDescribe multilingual NMT as in Johnson et al. (2017): training data, how the output language is chosen, what is shared, and how zero-shot translation arises.
Key points: one system with all parameters shared, transfer in parallel; English-centric data (, ) mixed within batches; target language tag prepended to the source; one WordPiece vocabulary and shared embeddings; non-English pairs are zero-shot.
Must hit:
- One single system for all translation directions, trained jointly: knowledge is transferred in parallel, as opposed to parent-child transfer in stages.
- Data: a collection of parallel corpora, English-centric: only and directions, because that is where the resources are. Directions are mixed during training, even within a batch.
- Target language tag: every training example gets a tag such as
<2es>prepended to the source as fed to the encoder. It names the target language; the source language is never named.- Shared everything: all parameters shared for all language combinations, one WordPiece vocabulary for all languages (32k, 64k), shared source, target and output embeddings. The architecture is a standard encoder-decoder.
- Zero-shot: training covers many-to-one and one-to-many; asking for a non-English pair (many-to-many, testing only) such as De→Fr is a direction never seen in training and relies entirely on transfer.
- Balancing: languages are sampled with weights, between proportional (high-resource dominate) and uniform (high-resource suffer).
Losing marks: saying the tag names the source language, or that the model trains on non-English pairs.