MNLP - Flashcards
All 421 flashcards for Multilingual Natural Language Processing: 52 exam-style questions and 369 recall cards, one set per note. Study everything at once, a single lecture, or tick any mix. Progress lives only in this browser tab and disappears when you close it.
Back to the Multilingual Natural Language Processing home.
L01 Overview
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Distinguish multilingual from crosslingual NLP. Give examples of each and explain why the English-centric nature of LLMs keeps both open problems.
- Multilingual: building systems that work for multiple languages. Examples: NER for multiple languages (can it be trained without being language-specific?), language-independent parsing (are there universal features?), cross-language classification (can it be trained without language-specific data?).
- Crosslingual: transferring information or knowledge across languages. Examples: machine translation, crosslingual QA, crosslingual unlearning, crosslingual reasoning.
- LLMs: trained predominantly on Internet resources, so they are English-centric. Multilingual models such as Llama 3.1 and DeepSeek do better on high-resource languages; dedicated multilingual models such as Aya-Expanse tend to lag behind. LLMs are still fragile on non-English text.
Losing marks: treating the two terms as synonyms. One is about covering many languages, the other about moving information between them.
Contrast pre-deep-learning NLP with deep-learning NLP. What did neural networks gain, what did they cost, and why do insights now transfer between fields?
- Before: a different methodology per application. Classical ML (SVMs, decision trees, generative Bayesian models, discriminative max-ent models) weighed hand-engineered features: POS tags, morphology, parse trees, named entities, taxonomies, argument roles.
- Advantages of deep learning: state-of-the-art performance across NLP tasks; very little or no feature engineering; a limited repertoire of network types covers most or all tasks.
- Disadvantages: needs large amounts of training data; errors are hard to trace (opacity); can fail spectacularly.
- Transfer: the same success appears in computer vision, handwriting recognition, speech, robotics and IR. Because network types are uniform and application-specific features are few, insights carry over between these areas.
What roles do large language models now play in NLP, and what are their limits? Use the Arabic to English translation comparison as evidence.
- LLMs originated within NLP and now dominate research.
- Uses: end-to-end NLP (QA, summarization, MT); substituting or supplementing humans in data annotation, evaluation (LLM-as-judge) and dialog.
- Limit: still fragile on non-English text, because training data is English-centric.
- Evidence (Monz’s Arabic to English example): GPT-5.6 Terra, Gemini 3.6 Flash and Claude Sonnet 4.6 all render the sentence as the embassy in South Sudan calling for an investigation into an attack that killed Ethiopian peacekeepers, differing only in wording. Mistral Small 3.2 produced “The Ethiopian embassy in Khartoum, where the Ethiopian peacekeepers were detained, is closed”: it catastrophically hallucinated the embassy’s location, and its sentence no longer matches the other three.
Name the four ways information can be conveyed or stored, with an example of each, and say where written and spoken language fall.
- Structured: tables, databases.
- Unstructured: language, images and videos.
- Continuous signals: spoken language and audio, images and video.
- Discrete signals: written language.
So language is unstructured; spoken language is a continuous signal and written language a discrete one.
Why is human language the medium of choice for complex information, and what is the catch for NLP?
A limited repertoire of words allows infinite expressivity. The catch: after decades of research, formally modelling language has proven surprisingly hard.
Define NLP: which three fields does it sit between, and which four aspects of language does it try to model?
NLP sits at the intersection of computer science, artificial intelligence and linguistics. Its goal is to model aspects of human language algorithmically and formally:
- Word formation (morphology)
- Sentence structure (syntax/grammar)
- Sentence meaning (semantics)
- Document or discourse structure
List six core NLP tasks and what each one does.
- Text categorization: assign documents to categories.
- Document summarization: extract or generate condensed versions.
- Machine translation: translate between languages.
- Question answering: return actual answers rather than ranked documents.
- Named entity recognition: identify persons, organizations, dates, locations.
- Sentiment analysis: estimate the attitude in reviews, either positive/negative or fine-grained.
How does question answering differ from information retrieval? Give an example.
IR returns a ranked list of documents; QA returns the actual answer. For “When was the Cuba Crisis?”, QA returns “1962” instead of documents that might contain it.
What does named entity recognition output? Show how an entity is tagged in a sentence.
NER identifies spans that name persons, organizations, dates and locations and labels each with its type. Example: “President [Biden]PER has received …”, where the span “Biden” is tagged as a person (PER).
Describe the pre-deep-learning approach to NLP: how methods related to applications, which ML methods and features were used, and what the ML actually did.
- Different application, different methodology.
- Methods: SVMs, decision trees, generative Bayesian models, discriminative max-ent models.
- Features: POS tags, morphology, parse trees, named entities, taxonomies, argument roles.
- Role of ML: to weigh the importance of individual (hand-designed) features for prediction.
Name the three multilingual scenarios and the question each one raises.
- NER for multiple languages: can it be trained without being language-specific?
- Language-independent parsing: can universal features be identified?
- Cross-language classification: can it be trained without language-specific data?
Name the four crosslingual scenarios and what each one does with information.
- Machine translation: makes information understandable across languages.
- Crosslingual QA: extracts information from resources in other languages.
- Crosslingual unlearning: manipulates information across languages.
- Crosslingual reasoning: combines information across languages.
Link to originalWhy are LLMs English-centric, and how do general multilingual LLMs compare with dedicated multilingual models?
They are trained predominantly on Internet resources, which are English-dominated. General multilingual models (Llama 3.1, DeepSeek) perform better on high-resource languages; dedicated multilingual models (Aya-Expanse) tend to lag behind.
L02 Multilinguality and Writing Systems
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
State the four English-first assumptions built into standard NLP pipelines. For each, give languages where it fails and what goes wrong concretely.
- Words are separated by whitespace. Fails for Chinese, Japanese, Thai, Lao, Khmer, Burmese: there is no word delimiter, so whitespace tokenization returns one token for a whole Chinese sentence. Thai uses spaces to end clauses or sentences and has no dedicated full stop, so a space means something different from what the tokenizer assumes.
- Grammatical relations are encoded by word order. Fails for languages with rich case marking (Russian, Finnish, Korean): who did what to whom is marked on the noun, word order is free, and position-based features carry much less information.
- One word = one unit of meaning. Fails for agglutinative languages (Finnish, Turkish, Hungarian) and compounding languages (German, Dutch, Swedish): Bundesverfassungsgericht is one token holding federal + constitutional + court.
- Text flows left to right. Fails for Arabic-script and Semitic languages (Arabic, Hebrew, Urdu, Persian): rendering runs right to left, so the first character in the string is not the leftmost on screen.
Argue that "script is not language" holds in both directions, and explain what each direction breaks in an NLP pipeline.
- One script, many languages: Latin (English, German, French, most European languages, several African, Turkic and Asian languages), Arabic (Arabic, Persian, Urdu, Pashto, Sorani), Cyrillic (Russian, Belarusian, Serbian), Devanagari (Hindi, Marathi, Nepali), Ethiopic (Amharic, Tigrinya). So detecting the script does not identify the language: Latin narrows nothing, and Arabic script leaves several languages that are not even all in one family.
- One language, many scripts: Serbian (Cyrillic and Latin), Kurdish (Latin and Arabic), Uzbek (Latin and Cyrillic), Punjabi (Gurmukhi and Arabic). So detecting the language does not tell you the script: a Serbian corpus can mix both, even in one document, and a model trained on one script and tested on the other sees no shared characters for the same language.
- Some scripts are language-specific (Armenian, Georgian, Greek, Sinhala, Khmer), which is the only case where the two coincide.
- Also: one sentence can mix several writing systems (Japanese uses Katakana, Kanji and Hiragana together), so a single “script of this text” label can fail.
Why is the English-centric bias of NLP structural, and why do claims of "language independence" deserve suspicion?
- Almost every resource exists first, often only, in English: annotated data, unannotated data, toolkit coverage (dictionaries, analyzers), benchmarks. Most publications also focus on English.
- Non-English resources arise mostly as by-products (translations, product reviews) or as translations of English annotated resources (so the annotation scheme stays English-shaped), and rarely as dedicated efforts.
- Coverage tracks commercial interest and military relevance, and speaker count is a poor predictor: over 7,000 languages, more than 100 with over 10 million speakers, yet Dutch is only medium resource.
- Standard pipelines encode properties of English (whitespace tokenization, fixed word order, closed vocabularies) as if they were properties of language.
- “Language independent” usually means tested on English plus a very small set of languages: a method gets the label because nobody checked.
- The bias is understandable, since resource creation costs time and money, which is exactly why it does not fix itself.
Compare alphabetic, syllabic and logographic writing systems and explain what each, plus the abjad and Hangul cases, implies for NLP.
- Alphabetic: each symbol is a consonant or vowel sound, small inventory (Latin, Cyrillic, Hangul).
- Syllabic: each symbol is a whole syllable (Hiragana, Katakana).
- Logographic: each symbol is a word or morpheme directly, little sound information (Chinese Hanzi, Japanese Kanji).
- Inventory size sets the cost of a character: tens of symbols for an alphabet, thousands for a logographic script, which a character-level vocabulary must budget for.
- Abjad (Arabic): consonantal skeleton, vowels omitted. كتب can be kataba, kutiba or kutub, so the orthography is lossier than the language and the model must resolve the ambiguity from context. No tokenizer can recover information that was never written.
- Hangul breaks the neat taxonomy: alphabetic by symbol-to-sound mapping (each jamo is a consonant or vowel), syllabic by visual arrangement (jamo grouped into square blocks), and letter shapes depict tongue and lip positions.
Explain the difference between a grapheme, a codepoint and a byte, and show how confusing them corrupts tokenization, deduplication and span annotation.
- Grapheme: what the reader sees as one character. Codepoint: Unicode’s abstract number for a character, one or more per grapheme. Byte: UTF-8 stores each codepoint in 1 to 4 bytes. They coincide only in ASCII.
- Example: NFD é is 1 grapheme, 2 codepoints (U+0065 U+0301), 3 bytes (65 CC 81).
- Tokenization: byte-level tokenizers see the byte stream. Byte-level splits can produce non-existent characters; codepoint splits can strand a diacritic from its base letter.
- Deduplication, search, equality, prefix matching assume identical codepoint sequences, so NFC and NFD copies of the same text silently fail to match. No error is raised, the numbers are just wrong.
- Span annotation (SQuAD-style QA): offsets recorded under one breaking convention and read under another give shifted or truncated answers. Counts across corpora with different conventions are not comparable.
- Fix: normalize consistently in the pipeline, and pick one breaking convention and enforce it everywhere. Consistency matters more than which level you pick.
Crawled multilingual data mixes normalization forms. Which Unicode normalization form would you use for search indexing, for comparing rendered text, and for generation targets, and why?
- The four forms form a grid: canonical (reversible) NFC and NFD; compatibility (lossy) NFKC and NFKD; composed (C) versus decomposed (D).
- Comparing rendered text: NFD, which splits into base characters plus combining marks and preserves semantic and visual equivalence. It is non-lossy, so you can decompose and recompose.
- Search indexing or NLP analysis: NFKD. Its lossiness is a feature: the ligature fi and the pair fi become the same, ① becomes 1.
- Generation: avoid compatibility forms. If training targets are NFKD-normalized, the model can never learn to produce the character the user actually typed (for example the long s ſ becomes plain s).
- Whatever you choose, apply it consistently in the pipeline (for example decompose then compose), otherwise dedup, search and equality checks are silently inaccurate.
Losing marks: getting the analysis versus generation rule backwards, or claiming NFKD is reversible.
Explain byte inflation for byte-level tokenizers, compute it for "hello", "café" and "中文", and argue why it is a fairness problem.
- , characters meaning what a reader perceives.
- “hello”: 5 bytes / 5 = 1.0x. “café” (NFC, é is 2 bytes): 5 / 4 = 1.25x. “中文” (3 bytes each): 6 / 2 = 3.0x.
- More bytes per character means more aggressive splitting, so the same meaning costs more tokens; byte splits can also produce fragments that are not characters at all.
- Fairness: a Latin-script user pays about 1 byte per character, a CJK user 3. The same sentence costs CJK users more tokens: shorter effective context window, more compute for identical content, and less linguistically meaningful units. The English-first assumption reaches the byte layer.
Encode U+4E2D (中) in UTF-8 by hand, stating which byte rules you use and why UTF-8's design is self-synchronising.
- 0x4E2D =
0100111000101101: 15 significant bits, more than the 11 payload bits of the 2-byte form, so use the 3-byte form1110yyyy 10yyyyxx 10xxxxxxwith payload bits.- Split 4 + 6 + 6:
0100 | 111000 | 101101.- Bytes:
11100100 10111000 10101101(hex E4 B8 AD).- Rules: a byte starting
0is a single byte (ASCII); a byte starting11begins a multi-byte sequence, with the number of leading 1s giving its length; a byte starting10is a continuation byte.- Self-synchronising: no byte value is ambiguous about its role, so from any offset you can tell whether you are at a character start and walk back by skipping
10xxxxxxbytes. UTF-8 therefore survives truncation, concatenation and byte-level search.Which four kinds of NLP resource exist first or only in English, and by which three routes do non-English resources come into being?
- Resources: annotated/labeled training data, unannotated training data, coverage of existing toolkits (dictionaries, analyzers), benchmarks/competitions.
- Routes: (1) naturally occurring by-products such as translations and product reviews, never made for NLP; (2) translations of existing annotated English resources, so the annotation scheme is English-shaped even when the text is not; (3) less so dedicated efforts, the route that would actually fix the problem.
How many languages are there, how many are large, and what actually determines how well a language is covered by NLP resources?
Over 7,000 spoken languages; more than 100 have over 10 million speakers. Coverage depends on commercial interest and military relevance; speaker count predicts it badly. Dutch, spoken by roughly 24 million people in a rich country, is only medium resource.
Give Monz's three resource tiers: what exists in each and which languages belong where, including the gap inside the top tier.
- High resource: large labeled and unlabeled data, many tools. English and Chinese (Mandarin), then an explicit big gap, then French, Spanish, Arabic, Russian.
- Medium resource: limited labeled and reasonable unlabeled data, some tools. German, Italian, Dutch, Korean, Japanese.
- Low resource: no labeled data, limited unlabeled data, no tools. All the remaining roughly 7,000 languages.
- The gap means “high resource” itself covers two very different situations.
Using 我不知道她去哪儿了。 ("I don't know where she went"), name three ways Chinese defeats English-style processing assumptions, and say what is shared with English.
- No spaces: whitespace tokenization yields one token, against 6 for the English sentence.
- No tense marking on the verb 去 (go): English changes go to went, Chinese leaves the verb alone.
- Past tense comes from the sentence-final particle 了, at the end of the clause, arbitrarily far from the verb, and a tokenizer treats it as an unrelated token. Models assuming morphological features live on the word they modify break.
- Characters are not words: 9 characters give 7 words plus punctuation (知道 “know” and 哪儿 “where” are two characters each).
- Shared: both the matrix clause (I NEG know) and the embedded clause (she go where) are SVO, so word order is not the difference.
Define morphology, morpheme, root and affix, with an example of each.
- Morphology: the study of the internal structure of words and how they are formed from smaller units of meaning.
- Morpheme: the smallest unit of meaning or grammatical function. dogs = dog + -s (plural).
- Root/base: the core part carrying the primary meaning: friend in unfriendly.
- Affix: prefixes (front) and suffixes (back) that modify a root’s meaning or grammatical role: un- and -ly in unfriendly.
How many surface forms does English build from "happy", and why did that make closed word-level vocabularies seem reasonable?
Eight: happy, happier, happiest, unhappy, unhappier, unhappiest, happiness, unhappiness. With so little inflection, a vocabulary of a few tens of thousands of word types covers English acceptably, so a closed vocabulary looked like a sensible design decision.
The Russian noun kniga ("book"): how many case and number cells, how many distinct forms, which forms are syncretic, and what does syncretism mean for NLP?
- 6 cases (nom, gen, dat, acc, inst, prep) x 2 numbers = 12 cells, but only 9 distinct forms.
- knígye covers dat.sg and prep.sg; knígi covers gen.sg, nom.pl and acc.pl.
- For NLP it cuts both ways: fewer types to learn, but the surface form is ambiguous between grammatical functions, so case cannot be read off the string.
What does "agglutinative" mean? Break Finnish onnellisimmillanikin ("even at my happiest") into its morphemes.
Agglutinative: morphemes stack, each contributing one piece of meaning, and the word can grow without bound.
- onnellis: stem, “happy”
- -imm-: superlative marker
- -i-: plural oblique-case marker
- -lla: adessive case (“at/on/in”)
- -ni: 1st person singular possessive (“my”)
- -kin: “even”
Six morphemes, one orthographic word; English needs four words, one of them a preposition.
How are Finnish onnellinen ("happy") and onnellisempi ("happier") formed?
- onnellinen = onni (“happiness”) + -llinen (adjective-forming suffix, roughly “having the quality of”), literally “having happiness”.
- onnellisempi: onnellinen has an oblique stem onnellis (adjectives in -nen often swap this for -s before suffixes), plus -empi (-mpi is the comparative suffix, with a linking vowel).
Why can more training data not fix a fixed word-level vocabulary for Finnish, and what does this motivate?
Agglutination produces combinatorially many word forms like onnellisimmillanikin, most of which a fixed vocabulary will never have seen however much data you collect. This motivates subword modelling.
What is the directionality trap for right-to-left scripts, and which operations does it break?
Codepoints are stored in logical order (Unicode is “written left to right”), and the visual right-to-left order is produced by the bidirectional algorithm at render time. So in Arabic and Hebrew “the first character of the string” and “the leftmost character on screen” differ. Code assuming they are the same breaks character offsets, span annotation and truncation, invisibly to a developer who cannot read the script.
Order the linguistic units from small to large and give what each level studies.
- Characters: size of the character set, writing systems.
- Words: word formation (morphology), lexical meaning/semantics.
- Clauses/sentences: grammar or syntax, sentence semantics.
- Paragraphs: discourse structure, anaphora, ellipsis.
Every higher level inherits whatever goes wrong at the character and word levels.
From "Amsterdam is the capital of the Netherlands" in eleven scripts, what four observations matter for NLP?
- Japanese mixes three writing systems in one sentence: アムステルダム and オランダ are Katakana, 首都 is Kanji, は, の and です are Hiragana.
- Thai has no spaces anywhere and no full stop at the end.
- Sentence terminators differ:
:(Armenian as reproduced; the proper Armenian full stop is ։, U+0589), ። (Amharic), 。 (Japanese), । (Bengali), nothing (Thai),.in the Russian, Georgian, Pashto, Inuktitut, Somali and Vietnamese lines. A[.!?]regex finds the end of only 6 of 11.- Same script, different languages: Somali and Vietnamese are both Latin, and Vietnamese adds tone marks and a stroked đ.
Define the three basic types of writing system and give examples of each.
- Alphabetic: each symbol represents a consonant or vowel sound; typically a small inventory. Latin, Cyrillic, Hangul.
- Syllabic: symbols represent whole syllables (vowel and consonant combinations). Japanese Hiragana and Katakana.
- Logographic: each symbol represents words or morphemes directly; typically little sound information. Chinese Hanzi, Japanese Kanji.
What does each script encode in water (Latin), पानी (Devanagari), 水 (Han), كتب (Arabic) and 한글 (Korean)?
- Latin water: alphabetic, sounds spelled out with letters.
- Devanagari पानी pani: alphasyllabic, consonant plus an inherent vowel.
- Han 水 shui: logographic, one symbol is one morpheme.
- Arabic كتب ktb: root only, vowels omitted.
- Korean 한글 (Hangul): letter shapes encode articulatory features.
What is an abugida? Explain with how पानी ( pani) is built.
The typological name for the Devanagari pattern: each consonant letter carries an inherent vowel that other marks override. पानी is pā + nī, two akshara: प (pa) with its inherent a overridden by the vowel sign ा (ā), and न (na) overridden by ी (ī). It is not spelled p + ā + n + ī as separate letters.
What is an abjad? Explain with the Arabic root k-t-b and the readings of كتب.
The typological name for the Arabic pattern: the consonantal skeleton is written and (short) vowels are left out. The root k-t-b underlies kataba (“he wrote”), kitāb (“book”) and maktab (“office”). The written form كتب itself can be read kataba (“he wrote”), kutiba (“it was written”) or kutub (“books”). kitāb and maktab are spelled differently (كتاب, مكتب) because long vowels and the m- prefix are written.
Which languages are written with more than one script, and which scripts are language-specific?
- Multiple scripts: Serbian (Cyrillic and Latin), Kurdish (Latin and Arabic), Uzbek (Latin and Cyrillic), Punjabi (Gurmukhi and Arabic).
- Language-specific scripts: Armenian, Georgian, Greek, Sinhala, Khmer.
What did ASCII provide, how did code pages extend it, and what were their two fatal disadvantages?
- ASCII (1963): 128 codes, enough for English letters, digits and punctuation.
- Code pages: the same 128 (or 256) codes in a byte, interpreted differently per language or script: Latin-1 (Western Europe), Shift-JIS (Japanese), GBK (Chinese), KOI8-R (Russian).
- Disadvantages: (1) text decoded with the wrong table becomes gibberish (mojibake), and since the bytes are legal in both tables no error is raised; (2) no multiple scripts in one document.
What happens when the UTF-8 bytes of Привет are decoded as Windows-1252, and why is this kind of error dangerous for corpora?
The bytes
D0 9F D1 80 D0 B8 D0 B2 D0 B5 D1 82decode as UTF-8 to Привет but as Windows-1252 to Привет. Same bytes, no error raised, so the garbage silently enters training data. (Strict Latin-1 would map9F,80,82to invisible C1 control characters; the visible garbage is the Windows-1252 superset’s output, which is what most real mojibake looks like.)Define Unicode and a codepoint, give two example codepoints, and state the two properties that carry most practical weight.
- Unicode: a single universal character set, maintained by the Unicode Consortium, in which every character in every supported script gets exactly one unique number.
- Codepoint: that number, written
U+number:U+0041is Latin “A”,U+4E2Dis Han 中.- (1) Unicode defines meaning only: how a codepoint is stored as bytes is a separate encoding question (UTF-8, UTF-16, UTF-32). (2) Unicode is a superset of ASCII:
U+0000toU+007Fare identical to ASCII.Give the scale of Unicode: code space, assigned codepoints, scripts, planes, and codepoints per plane, and show the arithmetic that links them.
- Over 1.1 million possible codepoints, 150,000+ assigned, 168 scripts, 17 planes spanning
U+0000toU+10FFFF.- Each plane holds codepoints, and .
- Watch the typo 65,356 that circulates in course material: the correct figure is 65,536.
Name the Unicode planes with their ranges and contents, and say why plane 0 matters most.
- 0, BMP (
0000–FFFF): most modern languages, the most common CJK Unified Ideographs, symbols such as currency, a Private Use Area.- 1, SMP (
10000–1FFFF): non-CJK historic ideographic scripts, modern scripts, symbols such as musical notation, emoji.- 2, SIP (
20000–2FFFF): additional CJK ideographs, mostly historical, uncommon or variants.- 3, TIP (
30000–3FFFF): historical CJK ideographs not in the BMP or SIP.- 4 to 13 (
40000–DFFFF): unassigned.- 14, SSP (
E0000–EFFFF): tags and variation sequence selectors.- 15 to 16, SPUA-A/B (
F0000–10FFFF): Private Use Area.Almost everything normally handled is in plane 0; anything outside needs more than 16 bits, which makes UTF-16 awkward and emoji four bytes in UTF-8.
What does a spread of codepoints such as a (U+0061), ã (U+00E3), & (U+0026), ⻩ (U+2EE9), ❁ (U+2741) and the dog face (U+1F436) show about the word "character" in Unicode?
That one namespace and numbering scheme covers ASCII letters, a precomposed accented letter (Latin-1 Supplement), ASCII punctuation, a CJK radical (the “yellow” radical, visually a form of 黄, yet a separate codepoint), a Dingbats ornament, and an emoji in plane 1.
Distinguish codepoint, character and glyph, and show with "café" why a naive string comparison fails.
- Codepoint: abstract number such as
U+00E9, Unicode’s unit of identity. Character: what a reader perceives as one unit, possibly several codepoints. Glyph: a stylistic rendering of a character.- NFC café:
U+0063 U+0061 U+0066 U+00E9(4 codepoints, precomposed é). NFD café:U+0063 U+0061 U+0066 U+0065 U+0301(5 codepoints, e + COMBINING ACUTE ACCENT).- They render identically but are not equal, differ in length and hash differently.
What are the NFC, NFD, NFKC and NFKD forms of ẛ̣ (long s with dot above plus combining dot below), and what does the comparison show?
- NFC: ẛ + ◌̣ =
U+1E9B U+0323- NFD: ſ + ◌̣ + ◌̇ =
U+017F U+0323 U+0307- NFKC: ṩ =
U+1E69- NFKD: s + ◌̣ + ◌̇ =
U+0073 U+0323 U+0307Canonical forms keep the long s ſ, since ſ and s are different characters under canonical equivalence. Compatibility forms replace it with plain s: information destroyed deliberately. (Simpler case: ã is
U+00E3in NFC andU+0061 U+0303in NFD.)Contrast NFD and NFKD: what each does, which is lossy, when to use each, and what the K stands for.
- NFD (canonical decomposition): splits into base characters and combining marks, preserving semantic and visual equivalence. Non-lossy: decompose then compose loses nothing. Use to compare rendered text.
- NFKD (compatibility decomposition): breaks characters into their most basic parts, including splitting ligatures. Lossy: visual styling is normalized to a base form that composition cannot reconstruct. Use for search indexing or NLP analysis, not necessarily generation.
- K = compatibility, from German Kompatibilität, because C was already taken by composition.
Compare UTF-8, UTF-16 and UTF-32: bytes per codepoint, ASCII compatibility, and where each is used.
- UTF-8: 1 to 4 bytes, ASCII compatible. The web, Linux, most file formats, virtually all modern NLP tooling; what people usually mean by “unicode”.
- UTF-16: 2 or 4 bytes, not ASCII compatible. Java, JavaScript strings, Windows internals, .NET.
- UTF-32: exactly 4 bytes, not ASCII compatible. Rare; simple for indexing but memory-hungry.
Give the four UTF-8 byte patterns with their codepoint ranges and payload sizes.
U+0000–U+007F:0xxxxxxx(7 bits)U+0080–U+07FF:110yyyxx 10xxxxxx( bits)U+0800–U+FFFF:1110yyyy 10yyyyxx 10xxxxxx( bits)U+10000–U+10FFFF:11110zzz 10zzyyyy 10yyyyxx 10xxxxxx( bits)Examples: A (
U+0041) is 1 byte, é (U+00E9) 2, 中 (U+4E2D) 3, grinning face (U+1F600) 4.Encode é (U+00E9) in UTF-8 step by step.
- 0xE9 = 233 =
11101001: 8 significant bits, more than the 7 of the 1-byte form, so use the 2-byte form (11 payload bits).- Pad to 11 bits and split 5 + 6:
00011 | 101001.- Byte 1 =
110+00011=11000011; byte 2 =10+101001=10101001(hex C3 A9).Encode the grinning face emoji (U+1F600) in UTF-8 step by step.
- 0x1F600 needs 17 bits, beyond the 16 of the 3-byte form, so use the 4-byte form (21 payload bits).
- Pad to 21 bits:
000011111011000000000; split 3 + 6 + 6 + 6:000 | 011111 | 011000 | 000000.- Bytes:
11110000 10011111 10011000 10000000(hex F0 9F 98 80).Worked example: encode COMBINING ACUTE ACCENT (U+0301) in UTF-8, and check it against the bytes of NFD "café".
- 0x301 =
1100000001: 10 significant bits, more than 7, at most 11, so the 2-byte form.- Pad to 11 bits, split 5 + 6:
01100 | 000001.- Byte 1 =
110+01100=11001100(CC); byte 2 =10+000001=10000001(81).So NFD café is
63 61 66 65 CC 81: 6 bytes, where the last two together are the accent.Worked example: decode the UTF-8 byte pair D0 9F (the first two bytes of Привет) to a codepoint.
D0=11010000: starts with110, so it opens a 2-byte sequence, payload10000.9F=10011111: starts with10, a continuation byte, payload011111.- Concatenate:
10000011111= 0x41F, so the codepoint is U+041F, the first letter П.State the byte inflation formula, define its terms, and give the three standard values.
Numerator: length of encoded as UTF-8. Denominator: number of characters a reader perceives.
- “hello”: 5 / 5 = 1.0x
- “café”: 5 / 4 = 1.25x
- “中文”: 6 / 2 = 3.0x
Worked example: compute the byte inflation of Привет, and of "café" in NFD, and say what the second result shows.
- Привет: its UTF-8 form
D0 9F D1 80 D0 B8 D0 B2 D0 B5 D1 82is 12 bytes for 6 characters, so 12 / 6 = 2.0x (each Cyrillic letter costs 2 bytes).- NFD café:
63 61 66 65 CC 81is 6 bytes for 4 perceived characters, so 6 / 4 = 1.5x, against 1.25x in NFC.- The normalization form alone changes what a byte-level tokenizer sees for the same visible text.
What two consequences does byte-level representation have for byte-level tokenizers?
- Byte inflation: the more bytes a character needs, the more aggressively tokens tend to be split, so a 3-byte script is chopped into more pieces than a 1-byte one for the same meaning.
- Non-existent characters: cutting a multi-byte sequence (for example a 3-byte sequence after byte 2) leaves pieces that are not valid characters, so the model learns over fragments no reader would recognise.
Which preprocessing steps break on unnormalized crawled text, how do they fail, and what is the fix (including the Python and command-line tools)?
- Crawled Unicode can mix forms (some pages NFC, some NFD) with nothing announcing which.
- Steps assuming identical codepoint sequences: deduplication, search, string equality and prefix matching. They fail silently: dedup keeps two copies differing only in form, and the reported corpus size is wrong.
- Fix: always normalize consistently in the pipeline, for example decompose then compose.
- Python
unicodedatacan normalize (unicodedata.normalize("NFC", s)) and test a form (unicodedata.is_normalized("NFC", s)).- For code-page text (GB2312, KOI8-R):
iconv -f GB2312 -t UTF-8, which fails if the from-encoding is wrong. That failure is useful, since silent wrong decoding is what creates mojibake.Compare breaking a string at byte, codepoint and grapheme level, using "café" (NFD).
- Bytes:
63 | 61 | 66 | 65 | CC | 81, 6 units, the last two are not characters. Byte breaking can split codepoints.- Codepoints: c | a | f | e | ◌́, 5 units, the accent is stranded from its base letter.
- Graphemes: c | a | f | é, 4 units, what the reader sees. Ideal, but grapheme boundaries are very language-specific, so it is the expensive option.
Where do mixed breaking strategies cause misalignments, and what is the rule to follow?
- Span annotation systems (SQuAD-style QA): spans stored as character offsets under one convention and read under another give shifted or truncated answers.
- Counting across corpora with different strategies: length statistics, token counts and coverage numbers become incomparable, with no warning.
- Rule: choose one convention and enforce it across every corpus, tool and metric. Any string breaking or length computation must be consistent across documents.
What is the Turkish-I problem?
Lowercasing is not language-independent. Turkish has two letter pairs, dotted İ / i and dotless I / ı, so lowercase(I) = ı in Turkish but i in English. Running
.lower()on Turkish text with the wrong locale changes which word you are looking at.Why can sorting not be done by codepoint, and how do Swedish and German order å, ä, ö differently?
- Codepoints of accented letters are not necessarily near their base letter’s, so codepoint order puts å nowhere near a.
- Swedish: a … z, then å, ä, ö as separate letters after z.
- German: ä sorts as ae, ö as oe, ü as ue, interleaved with the plain vowels.
- So sorting is a locale-dependent operation.
How should punctuation be tokenized? Give the steps for
He'd say "I promise you!", but then disappear.and the general-category approach.
- Agreed principle: punctuation adjacent to a word is not part of the word.
- Step 1, separate every mark: He ’ d say ” I promise you ! ” , but then disappear .
- Step 2, optionally expand the contraction: He would say …, a further normalization decision and a lossy one.
- Match the Unicode general category P with
\p{P}, as inregex.sub(r'(\p{P})', r' \1 ', string). This needs the third-partyregexpackage: Python’s built-inredoes not support\p{...}.- A hand-written class like
[.,!?]is the English-first assumption again: it misses 。, ،, । and ።.What is punctuation normalization, why is it hard, and where does it arise?
The same function is served by different characters across traditions, for example quotation marks: “…” (English), «…» (French, Russian), 「…」 (Japanese). There is no simple solution: the mappings must be manually encoded. It arises in quotes, end-of-sentence markers, question marks and single quotes.
Link to originalSummarise the Unicode stack in one sentence and list the main recap points.
A reader sees a grapheme, represented by one or more codepoints, each encoded in 1 to 4 bytes.
- Unicode can encode over a million characters, so writing-system-specific code pages are unnecessary.
- Normalization (NFC, NFD, NFKC, NFKD) can unify different encodings of the same text.
- UTF-8 is the most common form of Unicode in documents and NLP tooling.
- Any string breaking or length computation, byte- or codepoint-level, must be consistent across documents.
L03 Morphology and Word Formation
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Name the four morphological types, give example languages for each, and illustrate each with one analysed word or sentence. What distinguishes agglutinative from fusional?
- Isolating (Chinese, Vietnamese): words do not change form; grammar comes from word order and particles. Chinese 我 昨天 看 了 一 本 书 “I read a book yesterday”: 看 never changes, past comes from the particle 了 (PFV) and 昨天 “yesterday”.
- Agglutinative (Turkish, Finnish): morphemes strung together, one morpheme, one meaning, visible boundaries. Turkish ev-ler-iniz-den = house + plural + your + from, “from your houses”.
- Fusional (Russian, Spanish, Arabic): one affix encodes several features that cannot be separated. Russian knig-ami “with books”: -ami = instrumental and plural at once. Spanish habl-é: person, number, tense, aspect in one vowel.
- Polysynthetic (Inuktitut, Mohawk): verb, subject, object and modifiers (negation, tense, mood, instrument, location) packed into one word. Qangatasuukkuvimmuuriaqalaaqtunga “I will have to go to the airport”.
- The axis: isolating and polysynthetic are the two ends of a scale of how much meaning one word carries. Agglutinative and fusional sit in the middle and differ on a separate question: can the packed meaning be cut apart?
Losing marks: defining agglutinative as “more morphemes”. It means separable morphemes. Fusional languages can be just as dense; defining the classes by density misclassifies Spanish. Also: claiming -ami marks feminine gender. It is the instrumental plural for all three genders.
Why do fixed word-level vocabularies break for morphologically rich languages and for multilingual models, and what is the practical fix?
- Neural weight matrices are fixed size; the output layer is a matrix plus a softmax over at every position, so drives memory.
- Realistic sizes: English about 200K words, Russian about 1M, because Russian marks case on every noun so each lemma has many more surface forms.
- Inflectional features multiply: forms per lemma = product of slot sizes. Productivity (doomscrolling, doomscrolled, doomscrolls, doomscroller) means no vocabulary can be closed, so OOV never goes away.
- Two horns: keep all forms and explodes (embedding matrix and softmax with it); or use a vocabulary sized for English and get massive OOV rates, with every rare form mapped to
<unk>and given equal probability.- Word indices hide relations (move 2863 vs moved 87542) and give nothing for unseen words (hammered). Morphology could build meaning for unseen words; context alone cannot.
- Multilingual: a shared vocabulary leaves English words whole while Finnish gets fragmented into meaningless pieces, so sentences cost different numbers of tokens: a bias built into the tokenizer before training.
- Fix: morphological analysers are language dependent and need per-language expertise, so the practical compromise is subword tokenization: statistically useful pieces, language-agnostic, learned from data, zero OOV by falling back to characters.
Run Forward and Backward Maximum Matching on 研究生活 with vocabulary {研究, 研究生, 生活, 生, 活}. Why do the results differ, how does bidirectional matching resolve it, and what are the method's limits?
- FMM: at , 研究生活 is not in , 研究生 is (3 chars), emit it. At , 活 is in . Result 研究生 | 活.
- BMM: at , the longest suffix in is 生活 (2 chars), prepend it. At , 研究 is in . Result 研究 | 生活.
- Why: same string, same dictionary, only the scan direction changed. The string is genuinely ambiguous and a greedy rule lets the ambiguity leak through. Neither algorithm is buggy.
- Bidirectional: if FMM and BMM agree, accept with high confidence. If not, tie-break: fewer total words, or fewer single-character words. Here both give 2 words (tie), and BMM wins on single-character words (0 against FMM’s 1 for 活).
- Limits: needs a precompiled per-language vocabulary; new words (names, neologisms, typos) can never be matched; cannot improve from data; OOV words are silently chopped into single characters; genuine ambiguity cannot be fixed by any direction heuristic.
Reformulate Chinese word segmentation as sequence labelling and explain how an HMM with Viterbi solves it. Give the recursion, the complexity, and what the approach still cannot do.
- BMES tags: B (first char of multi-char word), M (interior of 3+ char word), E (last char of multi-char word), S (single-char word). Boundary after every E and S.
- HMM: hidden tag sequence generates the observed characters. Transitions , emissions . Decode .
- Training: MLE counts from a segmented corpus: , .
- Viterbi: , with a backpointer to the winning ; backtrace from the best final cell.
- Complexity: , i.e. for 4 tags, against brute-force tag sequences.
- Gains over maximum matching: no fixed dictionary, ambiguity resolved by probabilistic evidence, trainable from any segmented corpus.
- Still cannot: first-order Markov assumption (higher orders cost parameters); no longer-range context (motivates CRFs); unsupervised Baum-Welch (EM) training degrades quality. Neural sequence labellers (LSTMs, Transformers) are used in practice.
Distinguish inflection, derivation and compounding with examples, and explain why productivity of these processes matters for NLP.
- Inflection: keeps the part of speech, changes grammatical features (tense, number). see, saw, seen; book, books. Another form of the same dictionary entry.
- Derivation: changes POS and/or meaning, giving a new dictionary entry. happy → happiness (noun), happy → happily (adverb); perfect → imperfect (meaning only).
- One-line test: does the part of speech survive? see / saw inflection, happy / happiness derivation.
- Compounding: several words into one. German Tisch + Bein = Tischbein “table leg”; sicher + gehen = sichergehen “make sure”. German writes it as one token where English uses two.
- Productivity: the processes apply iteratively and to new combinations. doomscroll is a new compound already carrying inflection (-ing, -ed, -s) and derivation (-er). No fixed vocabulary can be closed over a productive process, which is the formal reason OOV never disappears.
What is non-concatenative morphology, and why can no string-cutting segmenter (from maximum matching to byte-pair encoding) fully handle it? Use German, Arabic and Inuktitut.
- Concatenative: word = concatenation of morphemes; string splitting can recover them.
- Non-concatenative: the change happens inside the root. Apophony: German Buch → Bücher, Haus → Häuser (vowel change plus suffix); English foot / feet, sing / sang / sung. No cut separates “book” from “plural” in Bücher.
- Circumfix (related problem): German ge-seh-en: ge- and -en are one morpheme in two pieces, so neither stripping the prefix nor the suffix alone works.
- Arabic: consonantal root k-t-b “write” poured into a pattern ya- C1 C2 -u- C3 -u = yaktubu “he writes”. Morphemes are interleaved. Short vowels are not written, so يكتب cannot distinguish yaktubu from yuktabu “it is written”: part of the information is absent from the string.
- Inuktitut: phonological rules at boundaries rewrite morphemes (-k-mut-uq surfaces as -mmuu-), so the word cannot be split back into its morphemes by string matching.
- Every segmentation method cuts a string into contiguous pieces, so subword segmentation is a solution to the concatenative case only.
Word segmentation and subword segmentation are often confused. Define each, give a language where each is needed, and say how they are solved.
- Word segmentation (tokenization): establishing word boundaries in scripts that do not mark them (Chinese, Japanese, Thai, Burmese). Must happen before POS tagging, parsing or anything else. Solved by maximum matching (FMM, BMM, bidirectional) or statistical BMES labelling (HMM plus Viterbi, CRFs, neural taggers). Benchmarked in the SIGHAN bake-offs.
- Subword segmentation: splitting words that are already well delimited into units smaller than the word, because morphologically rich languages (Finnish) blow up fixed vocabularies. Solved statistically, e.g. byte-pair encoding.
- Both decide where to cut a character string; they differ in what the pieces are meant to be: whole words against units smaller than words.
Define vocabulary, lexicon and dictionary, and say which one an NLP system usually has.
- Vocabulary: the inventory of words in a language, a corpus, or known to a person. A list.
- Lexicon: a store of the meanings and functions of words, beyond the word forms.
- Dictionary: a written version (published or digital) of a lexicon.
- An NLP system almost always has a vocabulary and rarely a lexicon.
What open design questions arise when building a vocabulary, and why are they all really questions about morphology?
Questions: all words seen in a corpus? language-specific? include typos? include very rare words? all variants (book, books, booking, booked) or just book? limited in size?
They are all one trade-off: giving related forms separate slots costs vocabulary size and hides that they are related, while refusing slots to rare forms and typos means you need some other way to represent them.
Name the four classical lexical relations with an example each. Why are hand-built resources like WordNet a poor basis for meaning, and what is the alternative?
- Synonyms (same meaning): purchase :: acquire. Hyponyms (is-a): car :: vehicle. Meronyms (part-whole): wheel :: car. Antonyms (opposites): small :: large.
- Hand-built resources are expensive (every relation entered by a human), incomplete (never cover a real corpus) and language-specific (cost multiplies per language).
- Alternative: learn relations automatically and quantitatively, e.g. , which is what word embeddings provide.
Define word embeddings by their four properties.
Distributed representations of words as vectors that are:
- low dimensional (e.g. 512, against of hundreds of thousands)
- dense (no zeros, unlike one-hot or count vectors)
- continuous:
- learned by performing a prediction task, rather than counted or hand-assigned.
State the CBOW task. How does it differ from n-gram language modelling, and how does Skip-Gram relate to it?
- CBOW (part of Word2Vec, Mikolov et al.): given position , the words to the left and words to the right , predict . Example:
the man X the road.- It resembles n-gram LM with = LM order and , but an LM sees only the left context. CBOW sees both sides, which is fine because its goal is good embeddings and prediction is only the means.
- Skip-Gram is the mirror image: predict each context word from the centre word.
What are the architectural choices in CBOW, and why is it called "bag of words"?
- Feed-forward network; focus on learning the embeddings, not prediction quality; simple network so capacity goes into the representation; embedding/projection layer moved closer to the output; typically with .
- Context word vectors are summed in the projection layer, and the sum predicts the centre word.
- Summation is commutative, so word order is destroyed:
the man X the roadandroad the X man thegive the same projection. The context is a bag. That information loss buys a very cheap model that trains on enormous corpora.What is a vocabulary short-list, what are typical sizes, and what is its main disadvantage?
- Keep only the most frequent words and map every rare word to a single token
<unk>. Typical sizes: 10K, 50K, 100K, sometimes 200K.- Motivation: weight matrices are fixed size, and the output layer ( plus softmax over at every position) dominates memory.
- Disadvantage: all rare words get equal probability in a given context, since they are the same token. The model cannot prefer hammered over doomscrolled in
She ___ the nail. The whole tail collapses into one symbol.Why is the English vocabulary about 200K words and the Russian about 1M?
Russian speakers do not know five times more concepts. Russian marks case on every noun, so each lemma surfaces in many more distinct forms, and each form needs its own slot in a word-level vocabulary.
Why is distributional (contextual) information not enough to represent words, and what does morphology add?
- Contexts pull related words together (desk and table), but word indices carry no structure: move (2863) and moved (87542), or transports and transported, must each have their relation learned from scratch from data.
- Unseen or rare words (hammered) have no contexts at all, so context cannot help.
- Claim: word-internal structure can induce meaning for unseen words. Knowing hammered = hammer + -ed, a model with an embedding for hammer and knowledge of -ed could construct a representation. Context and word-internal structure are complementary; word-level indexing uses only the first.
Why is word segmentation hard and why does it matter for the rest of the pipeline? Give the standard ambiguous example.
- It must happen before POS tagging, parsing and everything else, so errors propagate; it is a task in its own right (SIGHAN bake-offs).
- The same string often has several valid segmentations: 研究生活动 can be 研究生 | 活动 “graduate-student activities”, or cut as 研究 | 生活 | 动 “research on life activities”. Whether the second is a natural reading is not settled (it strands 动 as a single character), but it is a lexically possible cut.
- Even in whitespace languages tokenization has exceptions:
Mr. Smithmust not becomeMr . Smith, since the period belongs to the abbreviation.How many segmentations does a string of characters have? Compute it for 4 and for 20 characters.
Each of the internal gaps is independently a boundary or not:
- 4 characters (研究生活): .
- 20 characters: .
Enumeration is impossible at sentence length, so practical methods are greedy (maximum matching) or dynamic programs (Viterbi).
Give Forward Maximum Matching as numbered steps, explain the
or l == 1clause, and state its complexity.Input: string , vocabulary , max word length .
- Set , output empty.
- While : for from down to 1, take .
- If or : append , set , break.
- Return output.
Counting down makes the first hit the longest match.
or l == 1is the escape hatch: a single character is emitted even if not in , so the algorithm always terminates. It is also where OOV words get silently chopped into characters. Complexity: dictionary lookups.How does Backward Maximum Matching differ from Forward Maximum Matching?
Structurally identical, with three changes:
- Start at the end: , loop while .
- Take the substring ending at : , with from down to 1.
- Prepend to the output (so it stays in reading order) and set .
Define the BMES tag set and the rule for turning tags into a segmentation.
- B (Begin): first character of a multi-character word.
- M (Middle): interior character of a word with 3 or more characters.
- E (End): last character of a multi-character word.
- S (Single): a character that is a complete one-character word.
Insert a word boundary after every E and every S. The problem becomes ordinary sequence labelling.
Write the BMES tags for the FMM output 研究生 | 活 and the BMM output 研究 | 生活. What segmentation does 研/B 究/E 生/S 活/S give?
- 研究生 | 活 → 研/B 究/M 生/E 活/S.
- 研究 | 生活 → 研/B 究/E 生/B 活/E.
- 研/B 究/E 生/S 活/S → 研究 | 生 | 活, a third answer for the same four characters.
The point of labelling is that the right one can be learned from data instead of legislated by a scan direction.
What are the two kinds of HMM parameters for word boundary detection? Give an example of each and note the direction of the emission probability.
- Transition (tag bigrams): B is very likely followed by M or E, and can never be followed by another B.
- Emission : probability of the character given the tag. may be high because 究 usually closes 研究.
- Direction matters: emission is character given tag, the reverse of tag given character.
Toy calculation: estimate HMM parameters by MLE from a segmented corpus of two sentences, 研究 | 生活 and 研究生 | 活. Give , , , and .
Tags: 研究 | 生活 = B E B E; 研究生 | 活 = B M E S.
- (both sentences start with B).
- B occurs 3 times, each followed by a tag: B→E twice, B→M once. , .
- B emits 研 twice and 生 once: .
- E occurs 3 times (究, 活, 生), emits 活 once: .
Formulas: and .
Which BMES transitions are structurally possible, and why does that help a first-order HMM?
- B → M or E only. M → M or E only. E → B or S only. S → B or S only.
- Impossible (probability 0): B→B, B→S, M→B, M→S, E→M, E→E, S→M, S→E.
- Reason: a word opens with B, continues with zero or more M, closes with E; so B and M must be followed by something inside the same word, E and S by something that starts a new word.
- Half the transition matrix is effectively zero, and the model learns this from data without being told, which is why even a first-order model does well.
Give the Viterbi algorithm for BMES tagging as numbered steps and say what backpointers are for.
- Initialise , all other .
- For each position and each tag : .
- Store a backpointer to the that achieved the max.
- Backtrace from the highest-scoring final tag to recover the full sequence.
is the probability of the best path ending at position with tag . Backpointers avoid recomputing whole paths: only the winning predecessor is stored, and the pointers are walked backwards.
State the complexity of Viterbi for BMES tagging and compare it with brute force for a 4-character string.
: at each of positions, each of tags is compared against each of predecessors. With that is , linear in sentence length.
For 4 characters: predecessor comparisons, against complete tag sequences by brute force. The gap grows exponentially with .
What are the advantages of an HMM over FMM and BMM for word boundary detection?
- No fixed dictionary needed: it works over characters, so an unseen word is just a character sequence with a plausible tag path.
- Ambiguity resolved by probabilistic evidence instead of direction-based tie-breaking heuristics.
- Trainable directly from any segmented corpus, even from another domain.
What are the shortcomings of HMMs for word boundary detection, and what replaces them?
- First-order Markov assumption: a tag depends only on the previous tag and the current character. Higher orders are possible but parameters grow as .
- No longer-range context (the word two positions back, the surrounding phrase). Motivates CRFs, which condition on arbitrary features of the whole input.
- Unsupervised training with Baum-Welch (an instance of EM) is possible but degrades quality.
- In practice, neural sequence labellers (LSTMs, Transformers) are used.
What is the morphological status of a word whose part of speech is unchanged but whose meaning is negated, e.g. function → dysfunction or perfect → imperfect?
Derivation. Derivation changes the POS and/or the meaning, so a negating prefix that keeps the POS is still derivational: it yields a new dictionary entry. (Course material spells the example disfunction; the standard spelling is dysfunction, from Greek dys-. The morphological point is unaffected.)
Give the five forms of an English verb with see and call. How many distinct strings does each paradigm have, and what does that say about English?
Present, simple past, past participle, present participle, 3rd person singular:
- see, saw, seen, seeing, sees: 5 slots, 5 distinct forms.
- call, called, called, calling, calls: 5 slots, 4 distinct forms (past = past participle). Likewise send (sent, sent).
English inflection is impoverished, which is why English-shaped assumptions transfer badly to other languages.
What does English noun inflection mark, and what is its crucial limitation?
- Number: book, books; house, houses.
- Case on pronouns only: he, him (accusative), his (possessive), them (accusative plural), their (possessive plural).
- English nouns do not mark case at all, so an English-trained intuition has no slot for it.
Give three inflectional features other languages mark that English does not (or marks only marginally), with examples.
- Case on all nouns: Russian (more cases than English, marked on every noun), e.g. knigami.
- Mood in German: kommt (indicative), käme (subjunctive), komm (imperative).
- Aspect in Russian, Czech, Polish: Russian sdelat’ (perfective, completed) against delat’ (imperfective, ongoing). English has no inflectional marking for it.
Why do inflectional features multiply rather than add the number of word forms? Work through the numbers.
Features are independent slots, so forms per lemma = product of slot sizes.
- English noun: 2 numbers, no case: forms.
- Add a six-way case system: .
- Add possessive marking with six persons: forms from one lemma, before any derivation.
This is the arithmetic behind English 200K against Russian 1M.
Give English derivational patterns for nominalization, adjectivization and negation, with examples.
- Nominalization: verb + -ation (derivation); verb + -er (killer, baker); adjective + -ness (happiness).
- Adjectivization: verb + -able (accountable, reasonable); noun + -al (parental, colonial, official).
- Negation (meaning changes, POS kept): un- (unseen, unheard), mis- (misjudge, misappropriation), dis-/dys- (dysfunctional), im- (implausible), in- (indifferent).
Why does a naive "strip the suffix and look up the stem" analyser fail on English derivation? Give two phenomena.
- Spelling changes at the boundary: colony + -al = colonial (not colonyal); office + -al = official. The morpheme is regular but the surface string is not.
- Conditioned variants of one morpheme: im- before labials (implausible), in- elsewhere (indifferent). A model treating them as unrelated strings misses that they do the same job.
Define a morpheme and decompose disproportionally, saying what each piece contributes.
The smallest meaning-carrying part of a word.
dis + proportion + al + ly = prefix + stem + suffix + suffix: negation, core meaning, POS change to adjective, POS change to adverb.
Distinguish root from affix and free from bound morphemes. Why does the free/bound distinction matter for segmentation?
- Root (free): determines the basic meaning and can stand alone.
- Affix (bound): attaches to a root to change meaning or grammatical function, cannot stand alone.
- Free morphemes are words on their own (proportion, book, see); bound ones are not (dis-, -al, -ly, -ed).
- Bound morphemes are exactly the pieces that never appear as standalone tokens in a corpus, so a word-level system never sees them on their own.
Name the four kinds of affix by position, with an example of each. What is the caveat about the English infix example?
- Prefix: front; in English mostly one per word (dis- in disproportionally).
- Suffix: end; can stack (-al, -ly).
- Infix: inserted inside the word; more common in other languages. English example passerby → passersby.
- Circumfix: front and back simultaneously: German ge-seh-en, past participle of sehen. ge- and -en are one morpheme in two pieces, so it cannot be analysed by stripping a prefix or suffix independently.
Caveat: passersby is really plural marking on the head passer inside a compound. English has almost no true infixes; languages like Tagalog have productive infixation of a genuine morpheme inside the root.
Isolating languages: describe the type and analyse the Chinese and Vietnamese examples. What do they mean for a tokenizer?
- Words do not change form; grammatical relations come from word order; particles add modification. Chinese, Vietnamese.
- Chinese 我 昨天 看 了 一 本 书 “I read a book yesterday”: 看 kàn never changes; past comes from particle 了 (perfective, PFV) and 昨天 “yesterday”; SVO order carries the relations; 本 běn is an obligatory classifier (CL) between numeral and noun.
- Vietnamese Tôi đã đi học “I went to study”: đi and học do not change; đã is a past (PST) particle.
- Every morpheme is already a separate token: the easy case for word-level vocabularies, the hard case for anything expecting tense to be recoverable from the verb.
Analyse Turkish evlerinizden and Finnish taloissamme. What makes the agglutinative type learnable, and where are the boundaries not clean?
- ev-ler-iniz-den = house + plural + your + from, “from your houses”: one word for a four-word English phrase, each suffix one job.
- talo-i-ssa-mme = house + plural + inessive + our, “in our houses”. Inessive = locative “in”; other locative cases include elative “out of”, illative “into”, allative “to”.
- Learnable because suffixes are separable and reusable across nouns: a segmenter that finds -ssa once can apply it everywhere.
- Not clean at the character level: kala “fish” + -i- + -ssa = kaloissa, with stem vowel a becoming o. Boundaries are clean at the level of morphemes, not always of characters.
Analyse Russian knigami and Spanish hablé. Why can a string-based segmenter not recover their features, and what is wrong with glossing -ami as feminine?
- knig-ami “with books”: -ami fuses case (instrumental) and number (plural). No substring of -ami means “plural”, so no string segmenter can recover the features even in principle.
- habl-é “I spoke”: -é fuses person (1st), number (singular), tense (preterite, past) and aspect (completed). Four features, one vowel.
- Gender: Russian neutralises gender in the plural; -ami is used for all three (stolami masculine, oknami neuter, knigami feminine). An ending that really fuses case, number and gender is the instrumental singular: feminine knig-oy against masculine stol-om.
Explain Arabic root-and-pattern morphology with yaktubu, and why Arabic is the worst case for a naive tokenizer.
- Root = three consonants k-t-b “write”; pattern = template ya- C1 C2 -u- C3 -u; result ya-k-t-u-b-u = yaktubu “he writes”.
- The pattern fuses person (3rd), gender (male) and aspect (not completed).
- Morphemes are interleaved, so no set of cuts recovers them (non-concatenative).
- Short vowels are not written: in يكتب the prefix ي is written but the vowels are not, so the string cannot distinguish yaktubu “he writes” from yuktabu “it is written”, and the mood ending -u is invisible.
- A subword tokenizer learns consonant clusters and cannot represent the pattern as a unit.
Analyse the Inuktitut word Qangatasuukkuvimmuuriaqalaaqtunga and say what polysynthesis implies for vocabulary size.
- qangata- fly; -suukkuvik (-suu- habitually + -kkuvik place) = airport; -mut to (allative); -uq- go to; -riaqaq- have to; -laaq- future; -tunga 1st person singular. “I will have to go to the airport”.
- One word = a whole English sentence; it contains its own internal derivation (fly + habitual place = airport).
- Phonological rules rewrite boundaries (-k-mut-uq → -mmuu-), so string matching cannot split it.
- Word types number closer to sentences than to morphemes, so no vocabulary size (not 200K, not 1M) covers the language.
Is a morphologically rich language more complex than English? Explain with Turkish.
No. Rich morphology encodes what morphologically poor languages encode through word order and separate words. Turkish evlerinizden and English “from your houses” carry the same information, one in suffixes, the other in separate words and their arrangement. Only one of them suits a whitespace tokenizer.
Link to originalWhy are morphological analysers not the general solution to the vocabulary problem, and what does subword tokenization guarantee and not guarantee?
- Analysers are highly language dependent: one per language, built with per-language linguistic expertise (the WordNet cost problem again). Modern approaches are data-driven.
- Subword tokenization: find statistically useful pieces, language-agnostically, from data, without seeking linguistically correct morphemes.
- Guarantees: a fixed vocabulary with zero OOV, by falling back to characters.
- Does not guarantee: alignment with morphemes. It does not recover knig + ami or k-t-b + pattern; alignment is “somewhat, in concatenative languages”.
L04 Subword Segmentation
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Compare BPE and Unigram LM as subword segmentation methods: mechanism, segmentation quality and downstream effect.
Must hit three layers:
- Mechanism. BPE (Sennrich et al., 2016) is bottom-up: start from characters and greedily merge the most frequent adjacent pair until the merge budget is spent. A merge is never undone, and each word has exactly one segmentation. Unigram LM (Kudo, 2018) is top-down: start from a huge seed vocabulary of substrings, fit a unigram LM by EM, and prune the tokens whose removal costs the least likelihood until the target size is reached. It scores whole segmentations, , takes the global best by dynamic programming, and can return n-best or sampled segmentations.
- Segmentation quality (Bostrom and Durrett, 2020, same data and same vocabulary size): F1 against CELEX2 English morphology is 19.3% (BPE) against 30.3% (Unigram LM); against MeCab Japanese word segmentation, 73.8% against 77.2%. Unigram LM recovers suffixes (
ly,ed,s); BPE leaves word-initial single capitals and frequency artefacts such asn|an|ote|chn|ology.- Downstream (identical pretraining and model, only the tokenizer differs): Unigram LM wins the English tasks by roughly one point (SQuAD 1.1 EM 80.6 to 81.8) and Japanese TyDi QA by 12.3 EM and 12.3 F1 (EM 41.4 to 53.7).
- Vocabulary use: Unigram LM produces longer segments on average and uses its vocabulary more effectively; BPE spends the bottom of its vocabulary on near-useless tokens.
Losing marks: calling SentencePiece the algorithm (it is the toolkit that implements both); reading the low absolute English F1 as “tokenizers fail at English” when the meaningful number is the ratio between the methods.
Why does neural machine translation need subword segmentation? Argue from the weaknesses of word-level, character-level and rule-based (FST) alternatives.
- The vocabulary is one of the main bottlenecks of training an NMT system. How well a word is represented depends on its frequency and on the number of different contexts it occurs in, both properties of the training corpus.
- Most new surface forms are word formations: inflection (
tall,taller,tallest) and compounding (Donaudampfschifffahrtsgesellschaft). A word-level vocabulary treatstallandtallestas unrelated integers.- Word-level: 300K to 500K entries and still OOVs, huge embedding and softmax matrices, every unseen word collapses to
<UNK>(“Kyrgyzstan” in a test sentence becomes an unspecified place).- Character-level: near-zero OOV, but each symbol carries almost no meaning and sequences are long and slow.
- Hand-built FSTs: accurate, but extremely laborious, totally language dependent and give no preference among ambiguous analyses, so they do not scale to a hundred languages.
- Subwords (typically 30K to 128K): learned from frequencies with no annotation, vocabulary size tunable, near-zero OOV, moderate sequence length. The price is that pieces are not clean morphemes.
Train BPE on the corpus
the tall man is taller than the tallest manfor 10 merges, breaking ties by first occurrence. Give the merges and comment on the result.
- Word types:
the2,tall1,man2,is1,taller1,than1,tallest1. Split each into characters and append</w>.- Initial counts: six pairs tie at 3:
t+h(2 fromthe, 1 fromthan),t+a,a+l,l+l,a+n,n+</w>.t+his encountered first, so it wins.- Merges in order:
t+h(3),t+a(3),ta+l(3),tal+l(3),a+n(3),an+</w>(3),th+e(2),the+</w>(2),m+an</w>(2),tall+e(2).- Final corpus:
the</w>,tall </w>,man</w>,i s </w>,talle r </w>,th an</w>,the</w>,talle s t </w>,man</w>.- Comment: the frequent words
theandmanbecame single tokens, which is desired. Buttallerandtallestended astalle + randtalle + s + twhere the morphological split istall + erandtall + est. The greedytall + emerge was locally the most frequent move and destroyed the morpheme boundary, and BPE can never undo a merge.Losing marks: counting
thantowardst+a(inthanthetis followed byh).Segment
catsby bottom-up DP with p(c)=.02, p(a)=.04, p(t)=.03, p(s)=.05, p(ca)=.002, p(at)=.0015, p(ts)=.0004, p(cat)=.003, p(ats)=.00005, p(cats)=.0001.
- Length 1: is the character probability: 0.02, 0.04, 0.03, 0.05.
- Length 2:
castays whole, 0.002 (split gives 0.0008);atstays whole, 0.0015 (split 0.0012);tssplits, , so .- Length 3:
catstays whole, 0.003 (splits give 0.00003 and 0.00006);atsis best asat|s, , beating unsplit 0.00005 anda|ts0.00006, so .- Length 4: unsplit
cats0.0001; : ; : ; : . Winner , .- Read back: splits into and ; is unset, so the result is
cat | swith .- Points to make: the rare whole word
cats(0.0001) loses to two common pieces, which BPE cannot do once it has mergedcats. The morphologically correct split appears without any morphology in the model, because pluralsis very frequent. Combining the best scores of two halves is valid only because the unigram model treats segments as independent.Explain fertility and why one subword vocabulary serves languages unequally. Include byte-level BPE in your answer.
- Fertility = subword tokens / words, the average number of segments per word. 1.0 means every word is one token. English: 1.343 (BPE) and 1.318 (Unigram LM).
- Drivers: merge count or target vocabulary size (the one factor you control), average surface word length, morphological richness. Isolating Latin-script languages are close to 1 for common words; agglutinative languages (Finnish, Turkish, Hungarian) have many more, individually rarer word forms, which get split further.
- Consequences: (1) the model must reassemble meaning from many small parts; (2) no tokenization scheme dominates across all languages and scripts, and a multilingual vocabulary is a compromise for all of them, hence the push to 128K; (3) commercial APIs price by token, so high-fertility languages pay more for the same content; (4) fixed context windows fill up faster.
- Byte-level BPE: the 256-byte base alphabet guarantees zero OOV, but UTF-8 uses 1 byte for ASCII, 2 for Latin with diacritics, Cyrillic and Greek, 3 for most CJK, Devanagari and Thai, and 4 for many emoji. Those scripts spend merges just getting from bytes back to characters and end up with shorter subsegments. The cause is UTF-8 encoding length, which has nothing to do with the language itself.
- Methodology: when comparing across languages, measure and report fertility per language, because sequence length is a confound.
Walk through Kudo's (2018) Unigram LM vocabulary construction algorithm and correct the pruning rule in the version printed as Algorithm 2 in MNLP.
- Seed with all substrings that occur more than once in and do not cross word boundaries (millions of candidates: this is the top-down start).
- While : fit a unigram LM to by EM (segmentation is latent: the E-step computes expected subword counts over all segmentations, the M-step renormalises them into ).
- For each token : , where is the LM without and strings that used are re-segmented.
- Remove tokens, a hyperparameter, then refit. Pruning is gradual because each is computed assuming all other tokens are still present.
- Fit the final unigram LM and return and .
The error: the printed version removes the tokens with the highest . Removing a token cannot raise the likelihood, so and a large marks a valuable token: taken literally the printed rule prunes the most useful pieces first, and it is wrong. Correct rule: remove the tokens with the smallest (equivalently, keep the top-scoring pieces), which is Kudo’s intent and what SentencePiece implements. Also, Kudo never prunes single-character tokens, so every string stays segmentable; the printed lines would allow a character to be removed.
What two properties of the training data determine how well a word is represented in NMT, and which word-formation processes produce most new surface forms?
Its frequency in the training data and the number of different contexts it occurs in. Both belong to the corpus, so any fixed word list is a bet that future text looks like the training text. Many word occurrences are the result of word formation: inflection (
tall,taller,tallest) and compounding (Donaudampfschifffahrtsgesellschaft). A word-level vocabulary that has seentalla thousand times andtallesttwice treats them as unrelated.Define a subword, and compare word-level, character-level and subword vocabularies on size, OOV rate and sequence length.
A subword is either a character n-gram or a morphologically meaningful unit (language dependent). Data-driven methods aim at the first and are judged against the second.
- Word-level: 300K to 500K and still not enough; high OOV, unbounded on new text; short sequences; huge embedding and softmax matrices, no sharing between related forms, unseen words become
<UNK>.- Character-level: tens to a few hundred symbols; essentially zero OOV; very long sequences; symbols carry almost no meaning.
- Subword: typically 30K to 128K; near-zero OOV; moderate length; the vocabulary size is a tunable dial learned from data.
What is a finite-state transducer (FST) morphological analyzer, and what are its strengths and three weaknesses?
A hand-built machine of states and transitions that consume a string and emit a string, written
input:output(e.g.a:x). The input is a morpheme, the output an analysis such as “1st person, noun, singular”.
- Strengths: very powerful, efficient, handles cases and exceptions with great accuracy.
- Weaknesses: (1) extremely laborious to produce; (2) totally language dependent; (3) no preference among ambiguous analyses, just an unranked set (probabilistic FSTs exist, but the plain formalism ranks nothing).
Language dependence alone rules it out for multilingual NLP: one expert per language does not scale to a hundred languages.
What is Morfessor, what principle is it based on, and what two costs does that principle trade off?
The first data-driven morphological analyzer, Creutz and Lagus (2002), based on Minimum Description Length (MDL). MDL asks (1) how much it costs the model to generate the data and (2) how much the model itself costs, and minimises the sum. A whole-word vocabulary makes the data cheap and the model enormous; single characters make the model tiny and the data expensive. The optimum lies where reusable pieces exist, roughly where morphemes are, because morphemes are the pieces a language reuses.
Current data-driven segmenters make three simplifications relative to morphological analysis. Name them and the two practical properties gained in exchange.
Simplifications:
- No labelling of the data: only a split, no morphological analysis.
- No character substitutions or deletions:
run/ranandcity/citiesget no special treatment.- Strictly concatenative: a word is exactly the concatenation of its segments.
Gains: no data annotation or linguistic knowledge required; they can be fitted to a desired vocabulary size, which sets the size of the embedding and output layers.
Name the three frequency-based subword methods with their reference, direction and key property.
- Byte Pair Encoding, Sennrich et al. (2016): bottom-up (merge); greedy and deterministic.
- WordPiece, Schuster et al. (2012): bottom-up (merge); chooses merges by likelihood gain where BPE uses raw count.
- Unigram LM (in SentencePiece), Kudo (2018): top-down (prune); probabilistic, supports multiple segmentations.
All three split words using frequencies alone, with no linguistic knowledge.
Give the BPE training algorithm as numbered steps. When does it stop, and what is its only hyperparameter?
- Split words into characters, keeping word boundaries, and collect word frequencies.
- Consider all neighbouring symbol pairs and collect their frequencies.
- Merge the most frequent pair and repeat from step 2.
Stop when the maximum number of merge operations is reached. The number of merges is the only hyperparameter and controls the final vocabulary size.
What does the BPE output
The preside@@ nt of France arr@@ ived in K@@ yrg@@ yz@@ stan .illustrate?
@@marks a piece that continues into the next token. Frequent words (The,of,France,in) survive whole.presidentis cut atpreside|nt, which is no morpheme boundary;arrivedatarr|ived, close to one but on the wrong side. The unseen wordKyrgyzstanbecomes four pieces and is therefore representable at all. The trade: clean morphology is given up in exchange for being able to write down any word.In Sennrich et al.'s BPE reference code (Algorithm 1), what does
vocabmap, what doget_statsandmerge_vocabdo, and why is training cheap?
vocabmaps a space-separated symbol sequence to a word-type frequency:'l o w </w>': 5means the typelowoccurred 5 times.get_statscounts every adjacent pair weighted by the word’s frequency (a pair inside a word seen 5000 times counts 5000 times).merge_vocabrewritesa basabeverywhere; the regex(?<!\S)...(?!\S)(not preceded or followed by a non-space) ensures it matches whole symbols only, never part of a longer symbol.- Training runs over the type vocabulary instead of the whole corpus, which is why it is cheap.
What is the end-of-word marker
</w>for in BPE, and what does it not do?It lets a piece be word-final or word-internal: word-final
estinhighestand word-initialestinestablishbecome different symbols. That makes it possible to see where one word ends and the next begins after segmentation, so detokenization works. It does not stop merges across words: each word type is a separate key invocab, so no pair ever spans two words in the first place.How does Sennrich et al.'s BPE code break ties between equally frequent pairs, and why does that matter for reproducibility?
max(pairs, key=pairs.get)returns the first key attaining the maximum, and dicts iterate in insertion order, so the pair encountered first while scanning the vocabulary wins. Other implementations break ties lexicographically or by the frequency of the constituent symbols. So two tokenizers trained on the same corpus with the same merge count can produce different vocabularies. Pin the tokenizer library and its version; “BPE, 32k merges” does not identify a tokenizer.Trace BPE for 10 merges on {
l o w </w>: 5,l o w e r </w>: 2,n e w e s t </w>: 6,w i d e s t </w>: 3}, with counts.
e+s9 (newest 6 + widest 3; three-way tie withs+tandt+</w>,e+sis seen first)es+t9est+</w>9l+o7 (low 5 + lower 2)lo+w7n+e6ne+w6new+est</w>6low+</w>5w+i3Final:
low</w>5,low e r </w>2,newest</w>6,wi d est</w>3.loweris still four tokens because ten merges are not enough to reacher</w>: the vocabulary-size dial in miniature.How is a trained BPE model applied to a new word, and why is BPE deterministic?
Split the input into characters (with
</w>appended) and apply the merge operations in the order they were learned:
- For each merge in learned order,
- scan the symbols left to right and replace every adjacent with .
- Return the remaining symbols.
The learned model is an ordered list of merge rules, and applying them out of order gives a different segmentation. Practical implementations keep a rank table, repeatedly apply the lowest-ranked pair present (equivalent and faster), and cache results per word type. The same word always gets the same segmentation, whatever sentence it is in.
What are the three major advantages of BPE reported by Sennrich et al., and for which language pairs do they hold?
- Significantly reduces vocabulary size (less memory, better speed).
- Better translation quality.
- Significantly reduces the number of out-of-vocabulary items.
They hold for both high- and low-resource language pairs.
In Sennrich et al.'s English to German results, how do word-level and subword systems compare, and what do WDict and BPE-J90k show?
- WUnk and WDict carry 300 000 source and 500 000 target entries and score 20.6 and 22.0 BLEU (single system). Subword systems with 60k to 90k entries match or beat them: C2-50k 22.8, BPE-60k 21.5, BPE-J90k 22.8 (8-way ensembles: 25.3, 24.5, 24.7 against WDict’s 24.2).
- WUnk maps unknown words to
<UNK>; WDict uses a back-off dictionary to copy or translate them, worth 1.4 BLEU on its own, which shows how much damage the unknown-word problem does.- BPE-J90k is joint BPE, trained on the union of both languages so the same string is segmented identically on both sides, which helps the model copy names across.
What vocabulary range is typical for BPE today, and why is the upper end tied to multilinguality?
30K to 128K merge operations. A vocabulary that must cover many scripts needs more room before any single language gets a decent share of it, so the 128K end is a multilingual concession.
State the WordPiece merge criterion next to BPE's, define the terms, and explain how it changes which pairs get merged.
WordPiece merges the pair that most increases the likelihood of the training data under a unigram language model, which reduces to this ratio. is the frequency of the adjacent pair; and are the frequencies of each symbol alone. The denominator divides out pairs whose parts are common anyway (BPE happily merges
t + h), so WordPiece prefers pairs that are surprisingly frequent together, a pointwise-mutual-information-style criterion. This pushes it slightly closer to morpheme-like pieces. Otherwise it is the same bottom-up merge procedure as BPE. It is the tokenizer of BERT.How do WordPiece and BPE mark word structure differently, and why does it matter in practice?
WordPiece marks continuation:
playingbecomesplay,##ing, where##means “attaches to the left”. BPE as in Sennrich et al. marks word end with</w>. The information content is the same but the strings differ, and mixing the two conventions is a classic way to break a pipeline.What are BPE's two structural limitations, and what does each one prevent?
- Greedy: once a sequence is merged it stays merged. No lookahead and no undo, so an early locally optimal merge can block a globally better segmentation (
tall + edestroyingtall|er).- Deterministic: a word is always segmented the same way. You cannot sample alternative segmentations, so segmentation cannot serve as training-time regularisation, and you cannot hedge at inference.
Unigram LM (Kudo, 2018) addresses both.
State the four properties of Unigram LM (Kudo, 2018), and say how it relates to SentencePiece.
- Considers all possible segmentations of a word.
- Optimises for the best global segmentation.
- Supports multiple segmentations (sampling and n-best lists).
- Top-down: starts from a large existing vocabulary and shrinks it to the desired size, the mirror image of BPE, which merges upward from characters.
It is part of Google’s SentencePiece toolkit, whose other algorithm is BPE.
Give the Unigram LM probability of a segmentation and the rule for choosing the best one, defining every term.
- : one segmentation of the string into subwords, assumed independent
- : unigram probability of subword
- : the subword vocabulary
- : the set of all segmentations of string
BPE picks a segmentation by replaying local decisions; Unigram LM scores whole segmentations and takes the best.
In Unigram LM vocabulary construction, what does the token loss measure, which tokens should be pruned, and what is wrong with Algorithm 2 as printed in MNLP?
, with the LM without token : how much corpus likelihood drops if is removed and every string using it is re-segmented. A token always replaceable by a cheap split has a tiny ; one nothing else can express has a large . Correct rule: remove the tokens with the smallest (keep the top-scoring pieces), as Kudo (2018) intends and SentencePiece implements. The printed line 12 says remove those with the highest , which would prune the most valuable pieces first: that version is wrong and must not be implemented as printed. Kudo also never prunes single-character tokens, so every string stays segmentable and pruning cannot create OOV items; the printed algorithm omits this.
In Unigram LM vocabulary construction, what does fitting the LM involve, and why is pruning done gradually?
- Fitting is EM: the segmentation of each string is a latent variable. The E-step computes expected counts of each subword over all segmentations of each string; the M-step renormalises them into . It is refitted every round, which makes the loop expensive.
- Gradual pruning: remove at most tokens per round, , then refit, so remaining pieces can absorb the work of deleted ones. Removing straight down to in one step would be much worse, since each was computed assuming all other tokens are still present.
- The seed vocabulary (all substrings occurring more than once, not crossing words) is typically millions of candidates.
How many segmentations does an -character string have, why, and what is the brute-force procedure? What two ingredients does scoring need?
: there are positions between characters and each is independently a cut or not. Brute force: (1) generate all segmentations, (2) score each, (3) select the best. A 15-character German compound has segmentations per word per occurrence, so this is unusable. Scoring needs scores for segments (units not further split) and a way to combine scores of sub-segmentations. In the unigram model these are and multiplication, and the product over independent parts is what makes dynamic programming applicable.
Cut a sequence of length 4 to maximise value with segment prices 1, 5, 8, 9 for lengths 1 to 4. What is the best cut, and what does it show about segmentation?
There are cuts. Uncut: 9. or : 9. : 10. Three pieces ( in any order): 7. Four pieces: 4. Best is with value 10. Neither extreme wins: the optimum is an interior split, and finding it requires comparing whole configurations instead of making one local decision. Subword segmentation is the same problem with unigram probabilities as prices and multiplication in place of addition.
Define dynamic programming, name its two strategies, and compare them.
Solve each sub-problem only once and store its solution: extra memory buys computation time, a time-memory trade-off.
- Top-down with memoization: write the recursion naturally and cache each result the first time it is computed.
- Bottom-up: fill a table in an order that guarantees every sub-result is ready before it is needed (for segmentation, by increasing span length).
Both have the same asymptotic run time. Bottom-up avoids recursion depth limits and has better cache behaviour; top-down computes only the sub-problems it actually needs.
Give the bottom-up DP segmentation algorithm as numbered steps, and say what and hold.
- Let be the length of word and its characters; create tables and .
- For : .
- For span length , for , set and:
- initialise , the no-split option (zero, or in log space, if the substring is not in );
- for : ; if , set and .
- Return and .
is the score of the best segmentation of ; is the split point that achieved it, unset when leaving the span whole is best. Filling by increasing length ensures both halves are final. Optimising halves separately is valid only because the unigram model is independent across segments.
How is the segmentation read back out of the DP split table ?
Recursive procedure Segments:
- If is unset, output as one subword.
- Otherwise let , call Segments, then Segments.
Start with Segments.
What is the complexity of span-table DP segmentation, and how do production implementations do better?
cells with work each gives time and space in word length : fine for words, not for sentences, one reason segmentation is done per word. Because segments are independent, one best score per end position suffices: which is . Capping subword length at gives the Viterbi lattice in , which production implementations use; SentencePiece runs it over whole sentences since it does no pre-tokenization.
Why must DP segmentation under a unigram model be done in log space, and what changes in the algorithm?
Multiplying many probabilities underflows to zero in float32; every candidate then ties at zero and the argmax is meaningless. Replace
*with+, replace with , initialise unknown substrings to , and keep>as the comparison, since log is monotone. Nothing else changes.How can you obtain more than the single best segmentation, and what is that used for?
Viterbi returns only the most probable segmentation. For more:
- Yen’s and Eppstein’s algorithms give n-best segmentations (general k-shortest-path algorithms; the segmentation lattice is a DAG);
- or disallow certain sub-solutions: re-run the search with the winning split forbidden.
Use: subword regularization. At training time, sample a segmentation from the n-best list, so the model sees
tallerastall|ersometimes andtalle|rother times and cannot overfit to one segmentation. It is free data augmentation, unavailable to plain BPE, which has no distribution to sample from; BPE-dropout later retrofitted a similar trick by randomly skipping merges.How does SentencePiece's input handling differ from standard BPE implementations, and what are the two consequences?
Standard BPE implementations assume word boundary detection: a whitespace or language-specific tokenizer has already split the input into words. SentencePiece treats the input as a raw stream of Unicode characters, including spaces, with the space escaped as a visible symbol
▁(U+2581). Consequences: (1) no language-specific pre-tokenizer is needed, removing the last per-language engineering step and the thing that makes Japanese and Chinese awkward; (2) segmentation becomes lossless.What is the desegmentation problem, why is it worse for languages like Japanese, and how does lossless segmentation solve it?
Hello world.tokenized as[Hello] [world] [.]could be desegmented asHello world .,Helloworld.orHello world.: nothing records a space beforeworldand none before., so detokenizers guess with hand-written per-language rules. Japanese (こんにちは世界。) puts no spaces between words, so a rule that inserts spaces breaks Japanese and one that never does breaks English. Fix: a special whitespace character▁:Hello▁world.becomes[Hello] [▁wor] [ld] [.]. Desegmentation is joining the tokens and replacing▁with a space: no rules, no language knowledge, exactly invertible. Pieces may straddle a human word boundary because the boundary is encoded in the characters.Contrast BPE and Unigram LM segmentations of English words and numbers from Bostrom and Durrett (2020).
- Unigram LM:
▁fur ious ly,▁tri cycle s,▁nano technology,▁corrupt ed,▁Complete ly,▁pre post er ous,▁suggestion s,▁1848.- BPE:
▁fur iously,▁t ric y cles,▁n an ote chn ology,▁cor rupted,▁Comple t ely,▁prep ost erous,▁184 8.Unigram LM finds real morpheme boundaries; BPE produces frequency artefacts. In
nanotechnology, BPE has spent its early merges on generic frequent clusters and has no symbol fornano, while Unigram LM chose its vocabulary by usefulness. Numbers: BPE splits1848as184+8, so the model must rebuild the year from a meaningless shared prefix. Check what a tokenizer does to numbers, dates and code.How do BPE and Unigram LM segment 磁性は様々に分類がなされている。 ("Magnetism is classified in various ways"), and why does the difference matter?
- BPE:
磁 | 性は | 様々 | に分類 | がなされている | 。- Unigram LM:
磁 | 性 | は | 様々 | に | 分類 | がなされている | 。BPE glues the topic particle は onto 性 and the particle に onto 分類. Absorbing grammatical function words into content words is worse than an arbitrary split inside a word, because it destroys a unit the model needs as a unit. Unigram LM separates them correctly.
What do token length and frequency-rank distributions show about BPE and Unigram LM vocabularies of the same size?
- English lengths: identical at length 1 (both keep the full character alphabet); BPE has more vocabulary entries of length 3 to 5, Unigram LM more from length 7 upward.
- Conclusions: Unigram LM produces longer segments on average and uses its vocabulary space more effectively, with more tokens of moderate frequency.
- Frequency against rank: the curves coincide for roughly the first 15 000 ranks; then BPE’s collapses while Unigram LM’s holds near for a couple of thousand more ranks. BPE spends the bottom of its vocabulary on near-useless tokens, wasted embedding parameters.
- Japanese: distribution peaks at length 2 and dies out by about 9 (a Japanese character carries far more information than a Latin one); Unigram LM again produces longer segments on average.
Which tokens does BPE over-produce relative to Unigram LM in English and vice versa, and how do their tokens per word compare?
- BPE: word-initial single capitals (
▁H,▁L,▁M,▁T,▁B,▁P,▁C,▁K,▁D,▁R), debris left when a capitalised word cannot be merged into anything.- Unigram LM: suffixes and punctuation (
s,.,,,ed,d,ing,e,ly,t,▁a): morphology.- Tokens per word type: 4.721 (BPE) against 4.633. Tokens per word: 1.343 against 1.318, about a 2% difference in sequence length, in Unigram LM’s favour but small.
Give BPE and Unigram LM precision, recall and F1 against CELEX2 (English) and MeCab (Japanese), and interpret them correctly.
- English (CELEX2): BPE P 38.6%, R 12.9%, F1 19.3%; Unigram LM 62.2%, 20.1%, 30.3%.
- Japanese (MeCab): BPE 78.6%, 69.5%, 73.8%; Unigram LM 82.2%, 72.8%, 77.2%.
Low English recall is expected: neither method tries to find morphemes, and most English words are frequent enough to stay whole (
walkedas one token misseswalk|ed). The meaningful figure is the ratio: about 1.6x on English F1, about 1.05x on Japanese. The references measure different things (morpheme boundaries inside spaced words against word boundaries in unspaced text), and the downstream gap runs the other way (about one point in English, 12.3 on Japanese TyDi QA), so a larger gold-standard gap does not predict a larger task gap.What did Bostrom and Durrett find downstream when only the tokenizer differed, and what is the multilingual moral?
- English: Unigram LM wins consistently by about one point: SQuAD 1.1 EM 81.8 against 80.6 (F1 89.3 against 88.2), MNLI matched 82.8 against 81.4, CoNLL NER test F1 90.4 against 90.2.
- Japanese TyDi QA: EM 53.7 against 41.4, F1 54.4 against 42.1, a gap of 12.3 on both.
- The BERT_BASE row was trained on different data and is no controlled comparison (which is why it wins MNLI and NER); it only shows the models are in a sensible range.
Moral: the tokenizer is a hyperparameter invisible in English benchmarks and dominant outside them. A one-point English gain from a modelling change with an unreported tokenizer may be a tokenizer effect, so multilingual evaluation must control for segmentation.
Why is OOV handling a motivation for subword segmentation, and how can BPE and Unigram LM still produce OOVs?
Handling low-frequency words, including zero-frequency (OOV) ones, is one of the main motivations for subword segmentation. BPE and Unigram LM have very low OOV rates, but not zero:
- typos can sometimes still cause OOVs;
- unknown characters result in OOVs. This is the real problem: a character vocabulary holds only the characters seen in training, so an emoji, a rare CJK character or an unseen script yields a genuine unknown symbol with nothing to fall back to.
What is byte-level BPE, why does it guarantee a zero OOV rate, and what does it cost in the multilingual case?
It runs BPE on raw UTF-8 bytes instead of Unicode characters. The base alphabet is fixed at 256 symbols, so any input in any script, including unseen scripts, can be encoded byte by byte and then merged where patterns emerge. It is used by GPT-2 and most decoder-only models since, and its base alphabet is smaller than a Unicode-character alphabet, freeing slots for merges. Cost: non-Latin, non-ASCII scripts need more bytes per character (1 for ASCII; 2 for Latin with diacritics, Cyrillic, Greek; 3 for most CJK, Devanagari, Thai; 4 for many emoji). Extra merges are spent getting from bytes back to characters, giving shorter subsegments for those scripts: the same nominal vocabulary size is a weaker tokenizer for, say, Hindi than for English.
Define fertility with its formula, and name the three factors that influence it for a language.
The average number of segments a word is split into; 1.0 means every word is one token, and higher is worse for the model. Factors:
- the maximum number of merge operations (BPE) or target vocabulary size (Unigram LM), the one factor you control;
- the average length of surface words in the language;
- morphological richness: how many affix combinations are possible.
Isolating Latin-script languages typically see fertility close to 1 for common words; agglutinative languages such as Finnish, Turkish or Hungarian get split much further.
Name the four consequences of high fertility for a language.
- Modelling: a word chopped into many small, less meaningful parts forces the model to work harder to reassemble its meaning.
- No universal scheme: no single tokenization dominates across all languages and scripts; a multilingual vocabulary is a compromise for all of them, which is why multilingual models push to 128K.
- Money: commercial LLM APIs price by token count, so high-fertility languages pay more for the same content (English against Telugu).
- Capacity: fixed context windows fill up faster, so a 128k-token context holds less document in a high-fertility language.
Together these show the cost of English-centric NLP: pricing, context budget and per-token compute are all worse elsewhere, tracing back to merge tables learned from mostly English corpora.
Why is "we used SentencePiece" an under-specified description of a tokenizer?
SentencePiece is Google’s toolkit, and it implements both BPE and Unigram LM, so the phrase does not say which algorithm was used. What is specific to SentencePiece is the input handling: a raw Unicode stream including spaces, whitespace escaped as
▁, no language-specific pre-tokenizer, lossless desegmentation. Report the tokenizer, the algorithm and the vocabulary size.Link to originalWhat methodological rules follow for experiments that compare tokenizers or compare results across languages?
- Report the tokenizer, algorithm and vocabulary size as experimental settings.
- Comparing across languages: measure and report fertility per language, because sequence length confounds anything that depends on it.
- Comparing tokenizers: hold the vocabulary size fixed; 32k BPE against 64k Unigram LM measures the vocabulary size.
- Inspect actual segmentations of real inputs before looking at metrics, to catch numbers split into digits and particles glued to nouns.
L05 Static Embeddings
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Does unsupervised bilingual lexicon induction work? Answer with the mechanism, the headline result and the conditions under which it fails, with numbers.
Must hit:
- Mechanism. The shared space hypothesis says separately trained embedding spaces are approximately isomorphic, so one linear (orthogonal) map should align them. MUSE (Conneau et al., 2018) bootstraps a rough adversarially (a linear generator against an MLP discriminator, with re-orthogonalisation), picks the checkpoint with the unsupervised CSLS criterion, then refines with Procrustes () on a synthetic dictionary of mutual CSLS nearest neighbours, repeated.
- Headline result. Fully unsupervised matches or beats supervised Procrustes-CSLS on close, well-resourced pairs (P@1): en-es 81.7 against 81.4, es-en 83.3 against 82.9, en-fr 82.3 against 81.1, en-de 74.0 against 73.5.
- Where it falls short. Distant pairs: en-ru 44.0 against 51.7, en-zh 32.5 against 42.7.
- Where it collapses (Søgaard et al., 2018). Mixed- or double-marking, case-rich languages: EN-ET 0.00, EN-FI 0.09, EN-EL 0.07. EN-FI stays at 0.0 even when retrained on 1.7 billion Finnish words, so data size is not the explanation. Mismatched domains: 0.0 to 0.13 in every cross-domain en-es cell. Mismatched algorithms: Spanish CBOW against English skipgram gives 0.00 to 0.13.
- A free baseline beats it. A seed dictionary of identically spelled words beats adversarial training on every pair involving English (EN-ES 82.62 against 81.89, EN-ET 31.45 against 0.00).
- Verdict. It works when the languages are typologically similar, the corpora share a domain and both sides use the same embedding algorithm. Those are the conditions it was evaluated under, and they are the opposite of the low-resource setting that motivates it.
Losing marks: giving only the Conneau et al. headline numbers, or blaming the failures on data size alone.
Derive the negative sampling objective from a binary classification set-up, then derive its gradient with respect to the input vector and interpret it.
- Why. A full softmax needs a score for all words per training example. Replace “which word is the context?” with “is this (word, context) pair real or noise?”: 1 positive pair from the data, negative pairs from a noise distribution.
- Likelihood. , with the observed pairs, the noise pairs, and , where (column of ) and (row of ).
- Sigmoid identity. , so the objective becomes , and after logs
- Per-example loss. .
- Gradient. With :
- Interpretation. The positive coefficient is negative, so a descent step moves towards ; each negative coefficient is positive, so moves away from . Both shrink to zero once the classifier gets the pair right, so the biggest corrections come from negatives the model mistakes for real contexts. Cost per example drops from dot products to .
Compare Word2Vec (CBOW and Skip-gram) with GloVe: what each optimises, what information each uses, how each handles frequency imbalance, and what the final embedding is.
- CBOW: predict the centre word from the average of context embeddings, , , cross entropy, SGD.
- Skip-gram: predict each of the context words from the centre word, , context words assumed independent given ; in practice trained with negative sampling.
- GloVe (Pennington et al., 2014): weighted least-squares regression on log co-occurrence counts, .
- Information used: Word2Vec sees individual local context windows, treating each as an independent event, so it keeps “rediscovering” the same association. GloVe counts once over the whole corpus (global statistics) and then fits gradient updates to those counts.
- Frequency imbalance: Word2Vec subsamples frequent words (and draws negatives from ); GloVe uses the explicit weighting with , .
- Interpretability: Word2Vec is indirect (it learns to predict contexts); GloVe directly fits log counts.
- Final embedding: Word2Vec typically , or ; GloVe . Both sum the two vector sets.
- Context: prediction-based methods tend to outperform count-based ones; count-based ones use global information directly. GloVe is the hybrid.
Why does constraining a cross-lingual mapping to be orthogonal help? Give the objective, the two problems it fixes with a derivation for each, its closed-form solution, and the evidence.
- Objective (Xing et al., 2015). subject to , with all embeddings normalised to unit length. An orthogonal only rotates and reflects: no scaling or shearing.
- Problem 1, overfitting. An unconstrained can stretch, shear and rotate, so it can warp the space to fit the seed pairs and fail to generalise to the words outside the dictionary (the majority). Orthogonal preserves geometry: , so the source space moves rigidly, which is exactly what the shared space hypothesis says should suffice.
- Problem 2, objective mismatch. Training minimises Euclidean distance, retrieval uses cosine. For unit vectors and orthogonal : , so minimising distance is maximising cosine.
- Solution. , : one SVD of a matrix, exact, no learning rate.
- Evidence (en-es, P@1). Unconstrained (Mikolov et al., 2013) falls from 30.43% at 300 dimensions to 20.69% at 700, because has free parameters for the same dictionary. Orthogonal is better everywhere and rises slightly, 38.99% to 41.04%.
Losing marks: claiming Euclidean training and cosine retrieval always agree (they coincide only for unit vectors under an orthogonal map), or calling the unconstrained least-squares map “Procrustes” (in the results tables Procrustes means the orthogonal SVD solution).
Explain the hubness problem in cross-lingual word retrieval and how CSLS corrects it, including which term of the formula does the work.
- Hubness: in high-dimensional spaces a few vectors become the nearest neighbour of many points regardless of real similarity. It is a general property of high-dimensional geometry and worsens with dimension. A hub sits near the centre of the cloud and is moderately close to everything.
- Effect on translation: one target hub becomes the nearest-neighbour “translation” of many unrelated source words: nn(cat) = nn(car) = nn(house) = thing. Nearest neighbour is asymmetric:
thingcan be the neighbour ofcatwithoutcatbeing the neighbour ofthing.- CSLS (Conneau et al., 2018): , with , the mean cosine of to its nearest target vectors, the mean cosine of to its nearest mapped source vectors.
- Which term works: penalises a candidate that is close to lots of things (a hub has high ). is constant across candidates for a fixed , so it cannot change which wins; it matters when scores are compared across source words, e.g. ranking pairs to build a dictionary.
- Effect: significantly better retrieval with no parameter tuning beyond (Conneau et al. use ). Over plain NN on the same Procrustes map: en-es 77.4 to 81.4, fr-en 76.1 to 82.4.
Describe the full MUSE pipeline of Conneau et al. (2018) for aligning two embedding spaces without any bilingual signal, and justify each design choice.
- Adversarial step. Generator = the mapping (a single linear matrix); discriminator = an MLP (2 hidden layers of 2048, ReLU, input dropout) that tells mapped source vectors from real target vectors . They train in alternation: discriminator step with frozen, generator step with labels flipped and the discriminator frozen.
- Linear generator, deliberately. A deep generator could match the target distribution while scrambling which source word lands where. Only a constrained, near-rigid map forces point-level correspondence to come with the distribution match.
- Label smoothing . Targets stop the discriminator becoming overconfident, which would give the generator vanishing gradients.
- Only the 50k most frequent words are fed to the discriminator: rare words have poorly estimated embeddings, so frequent words give a cleaner signal.
- Re-orthogonalisation after each step, with , keeps near orthogonal so it cannot drift into a warping general linear map.
- Model selection without a dictionary. The adversarial loss is unreliable. Instead: for the 10k most frequent source words, find each one’s CSLS nearest target, average the cosines, keep the checkpoint with the highest value.
- Refinement. Build a synthetic dictionary of high-confidence mutual CSLS nearest neighbours among frequent words, re-solve by SVD (Procrustes), repeat. Translate with CSLS.
- Contribution of each part (en-es P@1): adversarial alone 69.8, + CSLS 75.7, + refinement 79.1, both 81.7.
What is fastText, how does it differ from Skip-gram with negative sampling, and what does the evidence of Bojanowski et al. (2017) say about where it helps and where it does not?
- Problem it addresses: Word2Vec and GloVe give one opaque vector per word type, so
run,runs,running,runnershare no structure; morphologically rich languages (Finnish, Turkish, Russian) split each lemma into many rare forms; OOV words get no vector.- Model: the centre word vector is the sum of its character n-gram vectors, , where holds all n-grams of
<w>for plus the whole word<w>.<and>mark word boundaries so prefixes and suffixes are distinct n-grams.- Same as SGNS: the loss , with context vectors still per word.
- Different: every n-gram of gets the same gradient (since ); the output is the n-gram table , and any word, seen or not, gets a vector by summing its n-grams.
- Word similarity: sisg (fastText) best or tied-best on 9 of 10 datasets; biggest gains in morphologically rich languages (Russian HJ 59 to 66 over sg) and on English Rare Words (43 to 47). Building OOV vectors from n-grams (sisg against sisg-) adds up to 6 points. Only loss: English WS353 (71 against cbow’s 73), frequent words where whole-word vectors are already good.
- Analogies: gains are syntactic (Czech 52.8 to 77.8, German +11.9, Italian +11.2, English +4.8 over sg). Semantic analogies do not improve and sometimes drop (German 66.5 to 62.3), because capital-country relations have nothing to do with character overlap.
How can the degree of isomorphism between two embedding spaces be measured without a dictionary, what did it show for English and Finnish, and what can it not show?
- Why a measure is needed: isomorphism is a strict true/false criterion; Søgaard et al. (2018) want a degree.
- Eigenvector similarity: build a nearest-neighbour graph per language (adjacency matrix ), take the degree matrix , form the Laplacian , keep the largest of its eigenvalues, and compute .
- Reading it: larger Laplacian eigenvalues mean denser connectivity; larger means more structurally different graphs. Eigenvalues are invariant to relabelling the nodes, so no alignment is needed.
- Results: EN-ES 2.07 is the lowest; the three pairs where adversarial alignment fails completely have the largest values (EN-ET 6.61, EN-FI 7.33, EN-EL 5.01). tracks adversarial success.
- Finnish explanation: a lemma spreads across dozens of inflected forms that cluster tightly by meaning overlap, while English gives a flatter, more evenly connected graph. Different connectivity means high , and no rotation can fix it, since a rotation preserves the neighbour graph and so the eigenvalues.
- Limit: near-isomorphic graphs have low , but a low does not imply near-isomorphism (non-isomorphic cospectral graphs exist). can rule isomorphism out, never in.
Losing marks: concluding that a low proves the spaces are isomorphic.
How should word embeddings be evaluated? Cover the evaluation axes, the word similarity and analogy tasks, and the limits of 2D visualisations.
- Two independent axes: intrinsic (how good the embeddings are by themselves) against extrinsic (usefulness in downstream tasks such as summarisation, MT, IR); qualitative (inspect selected examples, neighbours, plots) against quantitative (an overall task-dependent score). A 2D plot is intrinsic and qualitative; a similarity correlation is intrinsic and quantitative; BLEU after plugging embeddings into MT is extrinsic and quantitative.
- Word similarity: rank word pairs by cosine, compare with human judgements by Spearman correlation. WS-353 annotates relatedness, SimLex-999 similarity;
coffee/cupis related but not similar, so the two benchmarks can rank the same embeddings differently. Both mix kinds of similarity.- Analogies: a : b :: c : X, answer the word closest by cosine to (input words excluded), scored by accuracy; semantic and syntactic sets.
- Visualisation: project to 2D with PCA (linear, top-variance directions) or t-SNE (non-linear, keeps local neighbours, gives up global distances). Useful for showing sense clusters or a consistent gender direction.
- Caveats: non-linear projections distort distances; t-SNE hyperparameters change cluster sizes, distances and even apparent clusters (Wattenberg 2016). Plots complement quantitative and extrinsic evaluation and do not substitute for it.
Derive the closed-form solution of the unconstrained linear mapping , state when it exists, and explain why this mapping gets worse as the embedding dimension grows.
- Matrix form. Stack the dictionary pairs as columns: is , is . Loss .
- Gradient. (per pair: ).
- Set to zero. , so .
- Existence. () must be invertible, which needs at least linearly independent source vectors, i.e. pairs.
- Closed form against SGD. Closed form is exact; SGD (used by Mikolov et al., 2013) is approximate, scales better with large dimensions and dictionaries, and never forms or inverts .
- Dimension. has free parameters for the same seed dictionary, so more dimensions means more room to overfit: en-es P@1 drops from 30.43% (300 dimensions) to 25.76% (500) to 20.69% (700). The unconstrained map can stretch and shear to fit the anchors and generalises poorly.
State the distributional hypothesis, who formulated it, and the two families of methods built on it.
“You shall know a word by the company it keeps” (Firth, 1957). Represent each word by the contexts it occurs in; words in similar contexts get similar vectors and are taken to be similar in meaning.
- Count based (distributional semantics): count context words directly.
- Prediction based (neural embeddings): train a network to predict words from context or the reverse, and keep its weights.
GloVe is the hybrid of the two.
Along which dimensions should word representations let us compute similarity?
- Meaning, which itself has several dimensions: synonymy (
car/automobile), topical relatedness (car/road), and others.- Morphology:
run,runs,runningshould be recognisably related.Give the count-based recipe for building context vectors, step by step.
- Define a vocabulary of context words.
- Define a vocabulary of target words (may equal ).
- Define a window of size (or use the whole sentence or document as context).
- For each occurrence of , count how often each occurs within words to the left or right.
- Store the counts in a vector , with the frequency of context word in the context of . The vector has dimensions.
With window and leash, walk, run, owner, pet, bark, which words count as context for
dogin "he took the dog for a walk", and why aretookandtheignored?Only walk (3 to the right) counts.
took,the,forandaare within 3 tokens too, but they are not in , so they are ignored. A word only contributes if it is both inside the window and in the context vocabulary.Give the cosine similarity formula, what each part means, its range for count vectors, and why vector length is deliberately ignored.
Numerator: dot product; denominator: product of Euclidean lengths. It is the cosine of the angle: 1 for the same direction, 0 for orthogonal; with non-negative counts it lies in . Length is ignored because a frequent word has large counts everywhere and a rare word small ones; only the proportions of contexts should decide similarity.
Worked example: context words
runsandlegs, with dog , cat , car . Compute the three cosines.
- (about )
- (about )
- (about )
The two animals point almost the same way;
carpoints elsewhere.Worked example: over contexts (leash, walk, run, owner, pet, bark), dog and cat . Compute their cosine.
- Dot product: .
- ; .
- .
In a count table over (leash, walk, run, owner, pet, bark), bark and car come out at cosine 0.775, higher than dog and car (0.617). Why, and what goes wrong for a word like
lightwith an all-zero row?
- bark and car share only
owner: dot product , norms and , cosine . With such short vectors one shared context word dominates: the signal-versus-noise problem of count vectors.lighthas none of its contexts in , so its vector is zero and its cosine with anything is , undefined. A word whose contexts were not included has no representation at all.What are the properties of count-based context vectors, and their advantages and disadvantages?
Properties: high-dimensional (one dimension per context word, tens or hundreds of thousands), discrete (integer counts), sparse (almost all zeros). Advantages: an unsupervised way to induce word similarity; interpretable dimensions ( is the frequency of an actual word). Disadvantages: they stay relatively sparse despite tricks like stemming; there is no explicit criterion to separate signal from noise in the context: some words occur by chance, and it is unclear which context words are the important discriminators (
theco-occurs with everything).Give the idf formula as a weighting for context words, explain its terms, and compute it for a word in 10 and a word in 900 out of 1000 contexts.
= number of contexts, = number of contexts containing . A word in nearly every context gets and is effectively removed. With : and (natural log), so the rare, discriminative word counts about 44 times as much.
Are projections the same as embeddings? Give the criterion for a good word embedding and explain the difference between embeddings as by-products and representation learning.
A projection is the lookup layer of a neural model such as the probabilistic neural language model (PNLM): it maps each one-hot word to a dense vector learned with the rest of the network. Good embedding: if and only if and mean the same thing and show the same syntactic behaviour.
- By-products: in most models the main objective is a task (next-word prediction, classification); embeddings are learned only because they help.
- Representation learning: learning good embeddings is the main objective. Word2Vec is built this way: the prediction task is a pretext and the network is discarded except for the embedding matrices.
How do prediction-based word embeddings differ from count-based context vectors? Give four properties.
Same intuition (a word is characterised by its contexts), but embeddings are:
- low dimensional ( to , against ),
- dense (no zeros),
- continuous (),
- learned by performing a prediction task.
Their dimensions are not interpretable, unlike count vectors.
State the CBOW task, how it relates to n-gram language modelling, and its design choices.
Task: given the words left and words right of position , predict the word at (the man X the road). Relation to n-gram LM: an n-gram LM is the special case with = LM order and (left context only). Because CBOW also looks right, it is useless as a language model and only good for learning representations. Design choices (feed-forward network): focus on learning the embeddings; simpler than the PNLM, with no hidden non-linear layer; the embedding layer is brought closer to the output so its gradient is not diluted; typically with .
Write the CBOW model with every term and matrix shape, and give its loss and training method.
- : one-hot vector (length ) of the word at position ; : sum of the context one-hots (a count vector).
- : input (projection) matrix; : hidden layer of size , the embedding dimension.
- : output matrix, not necessarily shared ().
- : predicted distribution over the centre word.
No non-linearities, one hidden layer. Loss: cross entropy. Training: SGD.
In CBOW, what does multiplying by a one-hot vector do, what does the cross-entropy loss reduce to, why is it called "continuous bag of words", and what is the computational bottleneck?
- selects column of , so is the average of the context words’ columns: lookups, no real matrix multiplication.
- With a one-hot target, cross entropy is , the negative log-probability of the correct word (as in the PNLM).
- Bag of words: the context vectors are summed, so word order is lost. Continuous: the bag is a sum of dense vectors.
- Bottleneck: the softmax normalises over all words, so each example touches all of , cost . Negative sampling removes this.
In Word2Vec, where do the embeddings live, and which matrix is used as "the" embedding?
- Column of (): word as an input (context) word.
- Row of (): word as an output (predicted) word.
Typically is used, or , which combines both into one matrix whose row is word ‘s embedding (the transpose makes the shapes match).
Write the Skip-gram model, its independence assumption, and explain why its output layers produce the same distribution.
: one-hot of the centre word; is simply column of , no averaging. Context words are assumed independent given the input word. All output layers use the same and the same , so they produce the same ; only the target each is scored against differs. Loss (cross entropy, SGD). Equivalently, independent (input word, context word) pairs per position, which is how negative sampling treats it.
Compare CBOW and Skip-gram on input, output, hidden layer and predictions per position.
- CBOW: input context words; output the centre word; hidden layer = average of the context embeddings; 1 prediction per position.
- Skip-gram: input 1 centre word; output each of the context words; hidden layer = the centre word’s embedding; predictions per position.
What problem does negative sampling solve, and what task replaces the original prediction task?
Both CBOW and Skip-gram benefit from large data, but the full softmax needs a score for every word in for every training example. Negative sampling turns it into a binary classification: distinguish words that do and do not occur in the context of the input word, using 1 positive example (the word that actually occurred) and negatives drawn from a noise distribution. Cost per example drops from dot products to . It belongs to the contrastive learning family: pull the true pair together, push sampled pairs apart.
In the negative sampling objective, what are , , , and , and what is ?
- : observed (word, context) pairs from the data; : pairs drawn from a noise distribution.
- : the pair came from the data; : noise.
- : the parameters, the matrices and .
- is column of (input embedding); is row of (output embedding).
- with , and .
The dot product is a compatibility score pushed up for real pairs and down for random ones; words sharing many contexts get pushed towards the same output vectors and so towards each other.
What are the practical Word2Vec settings for negative sampling and context size?
- Skip-gram: add 5 to 20 negative samples per observed .
- Draw negatives from the unigram distribution scaled to , to bias towards rarer words.
- Context size typically 2 to 5.
- The more data, the smaller the context and the negative sample set can be.
Worked example: three words have unigram probabilities 0.9, 0.09 and 0.01. Compute the negative sampling distribution after raising to the power 0.75, and say why this matters.
- , , ; sum .
- Renormalised: , , .
The rare word’s probability almost triples (1% to 2.8%) and the frequent word’s drops (90% to 82.5%). Without it nearly every negative would be
the,oforand, and the model would learn little about the rest of the vocabulary.Give the Skip-gram with negative sampling training loop as steps.
- Initialise (input, columns ) and (output, rows ) with small random values; set the noise distribution .
- For each position with word , and each context position , : take the positive pair and draw negatives from .
- Positive: error ; accumulate into the gradient for ; update .
- Each negative: ; accumulate ; update .
- Update with the accumulated gradient.
- Return (or ).
Worked example: one positive pair scores , two negatives score and . Compute the negative sampling loss and the gradient with respect to .
- Positive: , loss .
- Negative 1: , loss .
- Negative 2: , loss .
- Total .
- Gradient: .
The largest correction comes from , the negative the model currently mistakes for a real context; the positive is already mostly right.
Worked example: at random initialisation all scores are 0, with one positive and two negatives. What are the negative sampling loss and the gradient with respect to ?
for every pair, so each term contributes and . Gradient: : equal pull towards the true context and push away from each negative.
Contrast count-based and prediction-based learning of word representations, including global matrix factorisation, and say what GloVe takes from each.
- Count-based: directly, or as global matrix factorisation (latent semantic analysis: factorise the co-occurrence matrix with SVD). Uses global co-occurrence information directly, but tends to perform worse.
- Prediction-based (CBOW, Skip-gram): local context windows. Tends to outperform count-based methods, but treats each window as an independent event, repeatedly rediscovering the same association.
- GloVe (Pennington et al., 2014) combines global co-occurrence statistics with local prediction-style updating: count once over the corpus, then train vectors by gradient updates to fit those counts.
Define , and in GloVe.
- : number of times word occurs in the context of word (the co-occurrence matrix).
- : number of times any word occurs in the context of word .
- : probability of word occurring in the context of word .
Write the GloVe objective and explain every term.
- : vector for word as target; : vector for context word .
- , : scalar biases.
- : log co-occurrence count, the regression target.
- : weighting function.
It is weighted least-squares regression: the dot product plus biases should predict the log co-occurrence count, so words with similar co-occurrence rows get similar vectors.
Why does GloVe regress on log counts? Use the ice and steam example.
If , then , a log ratio of co-occurrence probabilities. Ratios are what separate meaning:
iceandsteamboth co-occur withwater, but the ratio is large forsolidand small forgas. Vector differences then encode such ratios, which is also why analogy arithmetic works.Give the GloVe weighting function with its typical parameters, its purpose, and compute and .
Purpose: de-emphasise very rare (noisy) and very frequent (uninformative) co-occurrences. It rises smoothly from zero for rare pairs and caps at 1 so frequent pairs cannot take over. ; ; ; for .
Why is the GloVe objective well defined for word pairs that never co-occur, and why does that make GloVe cheap?
when , but , so those terms vanish. The sum effectively runs only over the non-zero cells of , and since is sparse, training is cheap.
Worked example: in GloVe, , , and . Compute the weight, error and loss term (natural log).
- Weight: .
- Error: .
- Loss term: .
The prediction is too low, so the update increases the dot product and biases for this pair.
Give the GloVe training algorithm as steps, and explain the distance weighting, AdaGrad and the final embedding.
- One pass over the corpus builds : for each centre word and context word in the window, += weight, often (1 for adjacent words, at distance 3), so need not be an integer.
- Initialise randomly.
- For several epochs, for each pair in shuffled order: weight , error , loss weight error; update all four by AdaGrad.
- Final embedding .
After step 1 the corpus is never read again (the “global” part). AdaGrad is SGD with per-parameter learning rates that shrink for parameters with large past gradients, useful since frequent words get many updates and rare words few. With a symmetric window is symmetric, so the two vector sets differ only by random initialisation, and summing averages out that noise.
What is subsampling of frequent words in Word2Vec?
Randomly discarding occurrences of very frequent words before training, with each occurrence of dropped with a probability that grows with ‘s frequency, so that words like
thedo not dominate the training pairs. It is Word2Vec’s way of handling frequency imbalance, where GloVe uses .Define the Spearman correlation used in word similarity evaluation, and compute it for four pairs with human ranks 1, 2, 3, 4 and model ranks 2, 1, 3, 4.
Pearson correlation between the ranks of the two lists, so only the ordering of pairs matters and the scale of the cosine is irrelevant. With no ties: with the rank difference for pair and the number of pairs. Here , , so .
What do WS-353 and SimLex-999 each annotate, and what do the WS-353 scores for media/radio and bread/butter show?
WS-353 annotates relatedness; SimLex-999 annotates similarity. Both mix kinds of similarity (synonyms, topical, unrelated). WS-353 scores (0 to 10): media/radio 7.42 is higher than television/radio 6.77, and bread/butter gets 6.19 although bread and butter are not the same kind of thing. These are relatedness judgements.
coffee/cupis the same case: highly related, not similar, rewarded by WS-353 and penalised by SimLex-999.Describe the analogy task: the question form, how the answer is computed, how it is scored, and the two kinds of analogy.
a is to b as c is to X (Paris : France :: Berlin : X). Compute (here ) and return the vocabulary word closest by cosine, excluding the three input words (the vector usually stays closest to one of them). Scored by accuracy: is the top answer exactly right. Classic case: . Semantic analogies (capital-country) and syntactic ones (acquired : acquire :: tried : try).
Worked example: , , . Compute the analogy vector and decide between candidates queen and prince .
, .
Answer: queen.
How do PCA and t-SNE differ as ways of visualising embeddings, and why is visualisation hard in the first place?
Embeddings have many dimensions (128, 512, 1024…), so they must be projected to 2D.
- PCA: linear projection onto the two directions of largest variance.
- t-SNE: non-linear; keeps each point’s nearest neighbours near it in 2D and gives up preserving global distances. A full-vocabulary t-SNE map shows local cluster structure but no readable global layout.
What do 2D projections of the neighbourhood of
powerand of male/female word pairs show about embedding spaces?
- power: different senses separate into regions: an electricity cluster (voltage, battery, solar), an energy/systems cluster (motor, fuel, supply), a control/operation cluster, and an abstract ability/authority cluster (strength, authority, ability). Capitalised
Powersits isolated frompower: the vocabulary is case-sensitive and the two forms get different vectors.- male/female pairs (brother/sister, uncle/aunt, king/queen, duke/duchess…): lines joining each pair all point roughly the same way, a consistent gender direction. This is the geometric fact behind and the property cross-lingual alignment relies on.
Why should 2D embedding visualisations not be taken at face value?
- Non-linear projections group points that are close in high-dimensional space, so distances in the plot are not distances in the space.
- t-SNE hyperparameters have a substantial impact (Wattenberg 2016): the same data can show different cluster sizes, distances, or apparent clusters that are not there.
They complement quantitative and extrinsic evaluation and do not replace it.
What three problems come from giving each word type one opaque vector, as Word2Vec and GloVe do?
- Related forms (
run,runs,running,runner) get potentially unrelated vectors with no shared structure.- Morphologically rich languages (Finnish, Turkish, Russian) suffer most: huge word-form families, each form a separate, rarer word with a worse-estimated vector.
- No fallback for out-of-vocabulary words: an unseen word has no vector.
Name the two options for making word embeddings morphologically sensitive.
- Sub-segment words (with BPE or a morphological analyser) and learn an embedding for each sub-segment.
- Sliding character window, fastText (Facebook AI, Bojanowski et al., 2017): represent a word by all its character n-grams.
Why does fastText wrap words in
<and>, and why is the whole word included in ?Boundary markers make prefixes and suffixes distinct n-grams:
<rucan only be word-initial,ns>only word-final, and the trigramherinside<where>differs from the word<her>. The whole word<w>is added so frequent words still get a word-specific vector component on top of their n-grams.Give the fastText training algorithm as steps.
- Inputs: corpus, window , n-gram range (default 3 to 6), dimension , negatives .
- For each vocabulary word, = all character n-grams of
<w>in that range, plus<w>; initialise n-gram vectors and context vectors randomly.- For each (centre , context ) pair: sample negatives from ; compose ; loss .
- Update every , , with the same gradient ; update and every .
- Return the n-gram table . There is no word table: any word’s vector is computed on demand by summing its n-grams.
Worked example: list the fastText character n-grams of
runsfor , count them, and say whatrunsshares withrunning.Wrapped:
<runs>(6 characters).
- :
<ru,run,uns,ns>- :
<run,runs,uns>- :
<runs,runs>- :
<runs>(also the whole-word token).
<running>has 23 n-grams (22 of length 3 to 6, plus the 9-character whole word) and shares<ru,<runandrunwith<runs>, so the two vectors share three summands and are pulled together automatically.Worked example: how many elements does have for the word
catwith ?Wrapped:
<cat>(5 characters). :<ca,cat,at>(3); :<cat,cat>(2); :<cat>(1); : none. The whole word<cat>is already the 5-gram, so .How does fastText handle an unseen word such as
runnings, and how does it cope with the huge number of distinct n-grams?The OOV word still has n-grams like
<run,runn,ning,ings>, most of them seen in training, so summing their vectors gives a sensible vector where Word2Vec would have nothing. In practice the n-grams are hashed into a fixed number of buckets (about 2 million), with one shared vector per bucket.In the fastText evaluation of Bojanowski et al. (2017), what are sg, cbow, sisg and sisg-, and what does sisg against sisg- measure?
- sg, cbow: the Word2Vec baselines.
- sisg (subword information skip-gram): fastText.
- sisg-: fastText where words absent from the training vocabulary get a null vector instead of an n-gram vector.
sisg against sisg- isolates the value of building OOV vectors from n-grams: up to 6 points of Spearman × 100 (German GUR350 64 to 70, Russian HJ 60 to 66).
What are the cross-lingual embedding goal, its main uses, the three training set-ups, and the hope behind alignment?
Goal: put translation-equivalent words close together, e.g. , as monolingual spaces already do for synonyms. Uses: bilingual lexicon induction (dictionary by nearest-neighbour search) and transferring models across languages. Set-ups:
- Train each language separately, then align the spaces.
- Train all languages together (benefiting from shared words like names and numbers), then align the regions.
- Train on data aligned at the word or sentence level.
Either way some alignment is needed: a mapping that brings translation equivalents close. Hope: large monolingual corpora plus only a small bilingual signal, since monolingual text is plentiful and parallel text is not.
State the shared space hypothesis and what follows if it is true.
- Languages trained separately still encode similar relational structure: king/queen in English relates like rey/reina in Spanish.
- There is an approximate isomorphism between the spaces: the overall shape of the point cloud is similar, while coordinates and orientations differ (random initialisation makes axes meaningless).
- If true (it is a hypothesis), a single geometric transformation, a general linear map, can align the spaces and bring translation pairs close together.
What does the 2D comparison of English and Spanish numbers and animals show?
Projections of one to five and horse, cow, pig, dog, cat in English and their Spanish translations (uno…cinco, caballo, vaca, cerdo, perro, gato): the Spanish plot is stretched, but the arrangement is the same.
onelies far right of the other numbers,fourat the top,twoat the bottom;catisolated at bottom left,horseandcowtogether at top left,dogto the right. Two independently trained spaces, one shape.State the structure-preservation condition for a mapping , apply it to the king/queen analogy, and say why it motivates a linear map.
, with vector addition and subtraction. So and the English gender offset maps to the Spanish one (rey, reina, hombre, mujer may be rotated and scaled differently, with the same relational structure). For linear this holds exactly, ; the reflects that the spaces are only approximately isomorphic. Strictly, a structure-preserving map is a homomorphism; an isomorphism must also be invertible.
Name the three types of cross-lingual alignment method with the bilingual signal each uses.
- Supervised: a seed bilingual dictionary as anchors, typically thousands of translations; solve for the mapping directly.
- Semi-supervised: a very small seed dictionary (hundreds); bootstrap new translations by statistical confidence, then re-solve.
- Unsupervised: no bilingual signal; uses adversarial training.
Where can seed dictionaries come from, and what is the problem with each source?
- Human-compiled machine-readable dictionaries. They list lemmas while embedding vocabularies are full of inflected forms (
hablamos), contain many senses per entry, cover low-resource languages poorly and lack multi-word expressions.- Learned from data: needs a sentence-level parallel corpus, then word alignment of each pair; alignments are often not one-to-one, so the dictionary is noisy.
- Words spelled the same in both languages (names, numbers, borrowings): false friends (Dutch
badmeans bath), only works for languages sharing a script, skewed to names and numbers. Despite this, identical-word dictionaries work surprisingly well.In a word alignment between "The Secretary of State visits The Netherlands" and "De minister van buitenlandse zaken brengt een bezoek aan Nederland", which alignments are not one-to-one, and why does it matter?
buitenlandse zaken(foreign affairs): both words align toState.brengt een bezoek aan(“brings a visit to”): four Dutch words align tovisits.Nederlandaligns to the two wordsThe Netherlands.A word aligner over a parallel corpus produces such links and the most frequent link per word gives a dictionary; the many-to-one cases make that dictionary noisy at the word level.
State the general alignment problem with every term, and the three basic questions it raises.
: paired vectors, a source embedding and the target embedding of its translation; : a matrix applied to every source vector; Euclidean norm. Questions: what should be (linear)? what constraints should it satisfy (orthogonality)? where do the pairs come from (a seed dictionary, or nowhere)?
What does "Procrustes" refer to in cross-lingual mapping, and which method belongs to which paper?
Without a constraint on , is ordinary multivariate least squares: Mikolov et al. (2013), the “translation matrix”. The Procrustes problem proper, and the “Procrustes” rows of the Conneau et al. (2018) results, is the orthogonally constrained version solved by SVD: Xing et al. (2015) onwards. The term is sometimes applied loosely to the unconstrained version.
What is a seed dictionary in practice, what are its limits, and what is the real goal of learning a mapping from it?
A list of (source word, target word) pairs, often the top few thousand most frequent words translated with an existing dictionary; automatic translation can be used with caution (errors become wrong anchors). Limited to single words, since multi-word expressions have no embeddings. Each pair is one anchor point, e.g. (cat, gato). The goal is a mapping that generalises beyond the anchors: 5000 pairs are useless on their own, while a mapping learned from them that translates the other 195 000 words is the point.
State the linear transformation hypothesis and justify why the map should be linear.
Strong assumption: a single linear transformation suffices to map one language’s space onto another’s. Why linear: king/queen and man/woman analogies are already linear regularities within one space, and the assumption is that the same regularities hold across spaces; a linear map preserves them exactly. It also reduces alignment to classic regression: find the matrix that best maps one paired point set onto the other.
Give the least-squares mapping algorithm (Mikolov et al., 2013) as steps, its gradient, and how a word is translated at inference.
- For each dictionary pair, = source embedding, = target embedding; assemble () and ().
- Initialise randomly.
- Until convergence, per mini-batch: loss , , with gradient .
- Return .
Inference: compute and return the target word whose embedding is nearest by cosine.
Worked example: in one dimension, the dictionary pairs are and . Find the least-squares mapping and its loss.
In 1D, . Residuals: and ; loss . Check: gradient .
In Mikolov-style word translation results (En to Sp, Sp to En, En to Cz, Cz to En), how does the translation matrix compare with edit-distance and co-occurrence baselines, and what does combining it with edit distance do?
- Edit distance translates to the most similarly spelled word (only works via cognates); word co-occurrence uses count vectors; the translation matrix is the learned linear .
- En to Sp P@1: translation matrix 33% against 13% (edit distance) and 19% (co-occurrence).
- Adding edit distance to the matrix gains about 10 points for Spanish (33 to 43) but only 2 for Czech (27 to 29), since Spanish shares far more cognates with English.
- Czech, the more distant language, is harder for every method.
How does linear-mapping translation accuracy depend on monolingual data size and on word frequency?
- Data size: P@1 rises from about 9 at training words to about 53 at (P@5 from 19 to 75), roughly linear in the log of data size and flattening beyond about . Better monolingual embeddings make the mapping work better.
- Frequency: accuracy drops as test words get rarer, from about 53 in the 5–7K frequency-rank bin to about 40 in the 17–19K bin, because rarer words have worse embeddings and less reliable mapped positions.
What do the errors of a linear Spanish-to-English mapping (imperio, millas, hablamos, protegida, determinante) reveal?
imperiogives dictatorship, imperialism, tyranny: topically related, none isempire.determinantegives crucial, key, important: the running-text sense, against the dictionary’sdeterminant.millasranks kilometers ahead of miles: the units occur in identical contexts.protegidaranks wetland first, presumably from contexts like zona protegida.hablamos(we talk) gives talking, talked, talk: morphology does not line up one-to-one, so the dictionary’stalkis only rank 3.Embedding nearness captures relatedness, while a dictionary demands exact equivalence.
Give the four limitations of an unconstrained linear mapping .
- It can stretch, shear and rotate the space, with nothing preserving the source language’s internal geometry (distances, angles).
- This allows overfitting to the seed dictionary: words near the seed pairs map well, the majority outside it do not.
- Training minimises Euclidean distance while inference retrieves by cosine; the two are not equivalent unless vectors are unit-normalised.
- So additional constraints on the form of are needed.
State the orthogonal mapping objective of Xing et al. (2015), its normalisation, and what an orthogonal can and cannot do.
with all embeddings unit length, . An orthogonal allows only rotations and reflections, no scaling or shearing, so it preserves angles and distances: . It can be solved exactly with SVD. After normalisation all points lie on the unit sphere and the map is literally a rotation (or reflection) of that sphere.
Derive that for unit vectors and orthogonal , minimising Euclidean distance equals maximising cosine, and compute when .
using and for unit vectors. Distance is a decreasing function of cosine. With : . (Cosine 0 gives 2; cosine gives 4.)
Give the closed-form solution of orthogonal Procrustes and derive it.
, , with , () holding the pairs as columns.
- .
- For orthogonal , , so minimising means maximising .
- is orthogonal, so and , with equality at .
- Hence , so .
One SVD of a matrix, no iteration, no learning rate.
Worked example: two dictionary pairs with , , , . What does orthogonal Procrustes return?
, so , which is already orthogonal: an SVD is , , . , the 90° rotation. Check: , , loss 0, and .
Compare unconstrained (Mikolov et al., 2013) and orthogonal (Xing et al., 2015) en-es mappings across embedding dimensions.
P@1: at 300/300 dimensions 30.43% against 38.99%; 500/500 25.76% against 39.91%; 700/700 20.69% against 41.04%; 800 (EN)/200 (ES) 35.36% against 40.06%.
- The unconstrained map gets worse as dimension grows ( free parameters for the same dictionary, more room to overfit); its best setting is the asymmetric 800/200.
- The orthogonal map is better everywhere and improves slightly with dimension.
Why can the orthogonality constraint not hold as written when English has 800 dimensions and Spanish 200?
is then non-square ( for English to Spanish), with rank at most 200, so () cannot be the identity. The constraint can only hold as (orthonormal rows) or with roles reversed. How Xing et al. handled this is not settled in the course material; the constraint as written applies cleanly only to square cases.
Worked example: translating
catwith . Candidate A has , ; candidate B has , . Which wins under nearest neighbour and which under CSLS?
- Nearest neighbour: A (0.70 > 0.62).
- CSLS: A ; B . B wins.
A is a hub: close to many mapped source words on average, so its high cancels its slightly higher cosine.
Give the definition of in CSLS and say what is.
the mean cosine between the mapped source vector and its nearest target vectors . is the mean cosine between target and its nearest neighbours in the mapped source space. The translation of is the maximising .
With supervised Procrustes on fastText embeddings, how do NN, ISF and CSLS retrieval compare, and how does accuracy vary by language (Conneau et al., 2018)?
- NN: plain cosine nearest neighbour. ISF: inverted softmax, another hubness correction. CSLS as defined.
- CSLS beats NN on every pair: en-es 77.4 to 81.4 (+4.0), fr-en 76.1 to 82.4 (+6.3), en-de 68.4 to 73.5 (+5.1), ru-en 58.2 to 63.7 (+5.5).
- Accuracy follows language distance from English: Spanish and French around 80, German low 70s, Russian 47 to 64, Chinese 30 to 43, Esperanto 20 to 29.
Why is unsupervised cross-lingual alignment wanted, and what is the chicken-and-egg problem?
Supervised methods need thousands of translation pairs: for many low-resource pairs even a small dictionary may not exist, and dictionaries may not cover inflected forms, which are the forms most often seen in data. So: align two point clouds, hypothesised to be approximately isomorphic, from geometry alone. Chicken and egg: a mapping is needed to find translation pairs, and translation pairs to learn the mapping. Conneau et al. (2018) break it with adversarial training, which bootstraps a rough alignment from distributional structure, then feeds translation pairs to the supervised machinery.
In adversarial cross-lingual alignment, what are the generator, the discriminator, the real and fake data, and why would fooling the discriminator align translations?
- Generator: the mapping , sending source embeddings into the target space; aim: make mapped vectors indistinguishable from real target vectors.
- Discriminator: a classifier trained simultaneously to tell mapped source vectors from real target vectors.
- Real data: target embeddings ; fake data: mapped source vectors .
If the discriminator cannot tell them apart, the mapped cloud has the target cloud’s shape and position; under the shared space hypothesis, the only rotation achieving this puts translations on top of each other.
Give the discriminator architecture and sampling choices in MUSE, with the reason for each.
- MLP with 2 hidden layers of 2048 units, ReLU, dropout on the input layer.
- Two-class output with label smoothing : targets instead of , preventing pathological overconfidence (a discriminator outputting exactly 0 or 1 gives the generator vanishing gradients).
- Source and target vectors drawn from the 50k most frequent words of each language: rare words have poorly estimated embeddings, so frequent words give a cleaner signal.
Write the discriminator and generator losses of adversarial alignment and explain each.
: probability that is a mapped source (fake) vector.
- : binary cross entropy with true labels, updated with frozen.
- : same cross entropy with labels flipped, updated with frozen; the term does not depend on and gives no gradient.
- Gradient for : , with , backpropagated through the frozen discriminator.
- Initialise as the identity (or a random orthogonal matrix).
Give the re-orthogonalisation update in MUSE, show that orthogonal matrices are fixed points, and compute its effect on singular values 1.2 and 0.8.
If is orthogonal, and the update gives . Per singular value: , pushing towards 1.
- :
- :
- : unchanged.
Without it, gradient steps would let drift into a general linear map with the warping and overfitting problems of unconstrained mappings.
Why is model selection hard in unsupervised alignment, and what criterion does MUSE use instead?
The adversarial loss is not a reliable indicator: the discriminator loss can look fine while the mapping is poor, and there is no validation dictionary. Criterion:
- Take the 10k most frequent source words.
- With the current , find each word’s CSLS nearest neighbour in the target space.
- Average the cosine similarities of these induced pairs; keep the checkpoint with the highest value.
Rationale: a poor mapping puts mapped words far from any real target word, so even best matches have low cosine. It is a proxy: high cosine does not prove the pairs are translations.
Give the MUSE refinement (self-learning) loop as steps, and say why it counts as semi-supervised.
- Start from the rough from adversarial training (global orientation roughly right, too imprecise for good retrieval).
- For each frequent source word , find ; keep only if is also the best source for (mutual nearest neighbours), and only high-confidence pairs.
- Re-solve exactly: , .
- Repeat 2 and 3.
Steps 2 to 4 are the supervised Procrustes method run on a dictionary the system built itself; starting from a small human seed dictionary instead of adversarial training gives the semi-supervised method. Mutuality filters out hub matches.
Who criticised the unsupervised alignment results of Conneau et al. (2018), and along which five lines?
Søgaard, Ruder and Vulić (2018): the choice of languages, the role of morphology, data size, domain, and the embedding algorithm.
How did Søgaard et al. (2018) change the language sample relative to Conneau et al. (2018)?
Conneau et al.’s languages (EN, FR, DE, ZH, RU, ES) are mostly dependent-marking, have few or no cases (German 4, Russian 6–7, the rest none) and none is agglutinative. Søgaard et al. added Estonian and Finnish (mixed marking, agglutinative, 10+ cases), Greek (double marking, fusional, 3 cases), Hungarian (dependent, agglutinative, 10+), Polish (dependent, fusional, 6–7) and Turkish (dependent, agglutinative, 6–7).
Define head-marking, dependent-marking, double-marking and zero-marking, and identify head and dependents in a possessive phrase and a clause.
Which word in a grammatical relationship carries the marker?
- Dependent-marking: the dependent. Head-marking: the head. Double-marking: both. Zero-marking: neither (word order alone).
- Possessive phrase: head = possessed noun, dependent = possessor.
- Clause: head = verb, dependents = subject and object.
Explain the marking in German "das Haus des Mannes" and "der Mann sieht den Hund", and contrast English.
- das Haus des Mannes: head
Haus, dependentMann; genitive marking on the possessor des Mannes (dependent-marking).- der Mann sieht den Hund: head
sieht, dependentsMann,Hund; nominative and accusative case on subject and object (dependent-marking). The verb also agrees with the subject (third person singular), so there is some head-marking, but dependent marking predominates.- English marks the same relations mostly by word order (the man sees the dog against the dog sees the man) and a few prepositions (of the man).
What did Søgaard et al. (2018) find when comparing adversarial alignment with identical-word supervision (P@1)?
- Mixed or double marking pairs fail completely with adversarial: EN-ET 0.00, EN-FI 0.09, EN-EL 0.07; identical-word supervision gets 31.45, 28.01, 42.96.
- Identical beats adversarial on every pair with English: EN-ES 82.62 against 81.89, EN-HU 46.56 against 45.06, EN-PL 52.63 against 46.83, EN-TR 39.22 against 32.71.
- ET-FI, two structurally similar languages, works adversarially (29.62, against 24.35 identical): the one case where adversarial wins.
- Identical-word supervision needs no human dictionary, so it is a fair free baseline.
French is listed as mixed-marking, yet English-French unsupervised alignment works well. What does that say about the explanation for the failures?
Marking type is not the whole story. The failing languages are mixed or double marking and case-rich (Estonian, Finnish: 10+ cases, agglutinative; Greek: 3 cases, fusional). The real factor is how many distinct surface forms a lemma has, to which marking type contributes. Hungarian, dependent-marking but agglutinative with 10+ cases, sits in between (45.06).
According to Søgaard et al. (2018), why does unsupervised alignment fail for mixed- and double-marking languages?
- Grammatical information is distributed across more word forms.
- A lemma is realised in many more distinct surface forms, each with its own embedding and each of lower frequency.
- This makes the mapping between a language that marks relations by word order and one that marks them by morphology more complex: non-isomorphic.
English
housecorresponds to Finnishtalo,talon,taloa,talossa,talosta,taloon,talolla…: no rotation maps one point onto a dozen.Give the eigenvector similarity algorithm as steps, and the correct form of the Laplacian.
- For each language, compute each word’s nearest neighbours and record them in an adjacency matrix (undirected nearest-neighbour graph).
- Degree matrix : diagonal, each entry the number of neighbours of that node.
- Laplacian (eigenvalues all ). Writing flips every eigenvalue’s sign, so “keep the largest” would pick the values nearest zero.
- Keep the largest of the eigenvalues of .
- over the kept eigenvalues.
Worked example: compute between a 4-node star and a 4-node path, keeping all Laplacian eigenvalues.
- Star (one hub, three leaves): eigenvalues .
- Path: , , , .
Two stars with nodes numbered in any order give . The star’s top eigenvalue (4) is its hub, one densely connected node, like tight clusters of Finnish inflected forms against a flatter English graph.
Worked example: compute between a triangle (3 nodes, all connected) and a 3-node path, keeping all Laplacian eigenvalues.
- Triangle: has eigenvalues .
- Path: eigenvalues .
. The triangle is more densely connected, which shows up as a larger second eigenvalue.
What does a high eigenvector-similarity between two languages mean, and why can no rotation fix it?
- Words in language A cluster in a way structurally unlike language B.
- The isomorphism assumption between A and B does not hold.
- A rotation preserves the nearest-neighbour graph exactly, so it cannot change the Laplacian eigenvalues.
Is the failure of English-Finnish unsupervised alignment a data size problem? Give the evidence.
No. Finnish Wikipedia is 12M words against 363M for Spanish, but retraining on the Finnish WaC corpus (1.7 billion words) left English-Finnish P@1 at 0.0. Meanwhile Estonian-Finnish, both mixed-marking agglutinative languages, reaches 29.62. Pairing structurally dissimilar languages is the problem.
How does unsupervised alignment accuracy differ by part of speech for en-es, en-hu and en-fi, and why?
- en-hu: nouns 26.87 and verbs 25.44 are worst; adjectives 53.28, adverbs 51.57, other 53.40 are about twice as good. Nouns (case) and verbs (person, number, tense) inflect most in Hungarian: the form explosion per lemma shows up per POS.
- en-fi: 0.00 for every POS.
- en-es: verbs weakest (66.05 against nouns 80.94, adjectives 85.53), since Spanish verbs are its most heavily inflected words.
How sensitive is bilingual lexicon induction to domain, for identical-word supervision and for adversarial alignment (en-es, EuroParl, Wikipedia, EMEA)?
Monolingual corpora of 1.1M sentences each from EuroParl (political), Wikipedia (general) and EMEA (medical).
- Same domain both sides: both work, 41 to 64 P@1 (identical: EP 64.09, Wiki 46.52, EMEA 49.24; adversarial: 61.01, 41.38, 49.43).
- Different domains: identical-word supervision degrades but survives (25.17 and 25.48 between EP and Wiki; 4.84 to 9.63 involving medical). Adversarial collapses to 0.0 to 0.13 in every cross-domain cell.
This matters because a low-resource language rarely lets you pick a monolingual corpus matching the English domain, and a domain mismatch alone breaks even English-Spanish.
How sensitive is unsupervised alignment to the embedding algorithm and its hyperparameters?
English fixed at fastText skipgram, window 2, n-grams 3–6; Spanish varied.
- Variations within skipgram (window 10, n-grams 2–7, or both): 81.89 down to 80.15 at worst, a loss of at most 1.74 points.
- Spanish CBOW against English skipgram: 0.00 to 0.13, even with identical hyperparameters.
The two algorithms produce spaces with different geometry on the same kind of data, so the approximate isomorphism unsupervised alignment depends on is partly an artefact of using the same algorithm on both sides.
Link to originalWhat does "static" mean for Word2Vec, GloVe and fastText embeddings, and what is the standard example of its limitation?
One vector per word type, whatever the sentence.
bankhas the same vector in river bank and bank account. Contextual embedding models (an LSTM or, in most current models, a Transformer encoder such as BERT) remove this restriction.
L06 Contextual Embeddings
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Why can mBERT transfer a task from one language to another when nothing in its training is cross-lingual? Argue from the experimental evidence.
- What mBERT lacks. Same architecture and loss as BERT (MLM + NSP) on 104 Wikipedias; no parallel data, no language-ID embedding, no language-specific parameters, no alignment loss, NSP pairs always within one language. Languages share only the parameters and the subword vocabulary.
- Vocabulary overlap is a minor cause. Pires et al. (2019): mBERT transfers across scripts (Hindi to Urdu POS 85.9, English to Bulgarian 87.1), and its zero-shot NER F1 is roughly flat in entity-wordpiece overlap, while English BERT’s F1 is near 0 at low overlap and rises with it. K et al. (2020): fake English removes all subword overlap and costs only 0.5 to 1.4 XNLI points.
- Structure matters. Transfer is best within the same word-order type (SVO to SVO 81.55, SVO to SOV 66.52) and rises with the number of shared WALS features. Permuting all words during pre-training drops XNLI by 8.4 (Spanish), 16.5 (Hindi) and 12.1 (Russian), yet transfer stays well over chance.
- Depth matters. At roughly constant parameter count, the gap between fake-English and Russian XNLI shrinks from 21.6 (1 layer) to 11.3 (24 layers): deeper networks learn more language-independent representations.
- Bottom line: no single factor explains transfer, or the lack of it.
- The limit. When premise and hypothesis are in different languages, accuracy falls under both monolingual settings, so the shared space is not truly language-neutral. Rajaee and Monz hypothesise that transfer of heuristics (such as premise-hypothesis word overlap) contributes to cross-lingual generalisation; mixed-language inputs remove that shortcut.
Losing marks: calling shared wordpieces the main cause; claiming mBERT saw parallel data or an alignment objective.
Explain why static word embeddings are inadequate and how contextual embeddings fix this. What design tension do contextual models face, and how can it be managed?
- Static embeddings, whether a by-product of a task (embedding layer of a classifier or MT system) or trained directly (Word2Vec with negative sampling, limited window), give one vector per word type: for the vehicle, the verb “to practise” and “train of thought”. Three senses, two parts of speech, one vector.
- In the by-product route the network did see the whole sentence, but only the input embedding layer was kept, so the context computed in deeper layers was thrown away.
- Contextual embeddings keep it: a deep network reads the whole sentence and every occurrence gets its own vector (a hidden state in that word’s column).
- Tension: integrate enough context to distinguish occurrences of the same word, without capturing so much that the word’s own contribution becomes unclear (every position becomes “the meaning of the sentence”).
- Managed by: (1) choosing the right layer(s): lower layers are closer to the word, higher closer to the sentence; (2) choosing the right training task(s): per-token outputs keep each top state about its word, a single pooled output pushes the top layer toward the sentence; (3) tightening connections between layers, e.g. residual connections, which carry the word’s own information upward alongside the added context.
Describe how BERT is pre-trained and fine-tuned: input representation, both pre-training objectives, the role of
[CLS], and how task heads attach.
- Model: encoder part of the Transformer, bidirectional (every position attends left and right). Devlin et al., NAACL 2019.
- Input:
[CLS] A [SEP] B [SEP], with (token + segment A/B + position embedding). WordPiece tokens.- MLM: select 15% of tokens; replace 80% with
[MASK], 10% with a random token, leave 10% unchanged; loss on the selected positions only, predicting the original token.- NSP: 50% B really follows A (IsNext), 50% B is random (NotNext); classified from , the final-layer output at
[CLS].[CLS]: never masked, not tied to a word, represents the whole sequence, trained through a classification loss.- Fine-tuning: small randomly initialised head; connect it to the token outputs for tagging or span extraction, to for classification; train on the task with or without updating BERT. Batch 16 or 32, 2 to 4 epochs.
- Self-supervised pre-training can use huge amounts of task-irrelevant data; supervised fine-tuning then works with little task data.
Losing marks: saying the loss covers all tokens, or only the positions showing
[MASK]; saying BERT-base has 16 attention heads (it has 12).Describe how mBERT is trained and what it deliberately lacks. Explain its language sampling scheme and the problems a single shared multilingual vocabulary causes.
- Released on GitHub in 2018 with no paper. Exact BERT architecture and loss (MLM + NSP), on the concatenated Wikipedias of 104 languages.
- Absent: parallel data (at least intentionally), language-ID embedding, language-specific parameters such as adapters, any alignment loss. NSP pairs are always in one language; no language mixing within a sequence (a batch can mix languages).
- Sampling: Wikipedia sizes are extremely skewed (English alone exceeds dozens of small languages combined), so languages are drawn with . : proportional, big languages dominate; : uniform, tiny languages are memorised; : up-weights low-resource and down-weights high-resource languages.
- Vocabulary: one shared WordPiece vocabulary of 120k (English BERT 30k), which is why mBERT has about 178M parameters against 110M. CJK characters are split into single characters before WordPiece. The sampling smoothing is also applied to the counts used to build the vocabulary.
- Problems: even smoothed, the vocabulary favours high-resource languages and Latin script; low-resource and morphologically rich languages get much higher fertility (more pieces per word); the uncased variant lowercases and strips accents, which damages languages that depend on diacritics.
Calculate: English has 1000 units of training text and Swahili 10. Using , give both sampling probabilities for , and , and interpret each.
- : , . Proportional to size; English dominates.
- : , , so and . Swahili gets about 3.9 times its raw share; English still gets 96%.
- : . Swahili’s share grows about 50-fold over its raw share, and each unit of Swahili text is drawn 100 times as often as each unit of English ( against ): the model would memorise the small language.
- General rule: the ratio between two languages goes from to , because with is concave and compresses large values more. Here 100:1 becomes about 25:1 at .
Compare XLM-R with mBERT: what changed in training, and how do the two compare within a language and across languages on XNLI and cross-lingual QA?
- XLM-R (Conneau et al. 2020) is a RoBERTa-style extension of mBERT: MLM only (NSP dropped: contributes little, sometimes hurts), dynamic masking, CC-100 (2.5TB filtered CommonCrawl, 100 languages, two orders of magnitude more data than Wikipedia), 250k Unigram LM vocabulary via SentencePiece (better fertility, bigger embedding matrix).
- Within a language XLM-R is better on both tasks: XNLI 74.2 against 65.7 (+8.5), XSQuAD F1 72.0 against 64.4 (+7.6).
- Across languages on XNLI it is also better: 64.8 against 54.5 (+10.3); gap within to across 9.4 against 11.2.
- Across languages on QA it is worse: 36.8 against 44.2 F1, a within-across gap of 35.2 against 20.2. Thai: within QA rises from 40.0 to 66.5, across barely moves (26.1 to 29.6).
- XLM-R’s QA matrix is “own language, or English question”: strong diagonal and English-question column (58.2 to 75.0), little in between. mBERT degrades more smoothly.
- Conclusion: more data and a bigger vocabulary made each language better on its own without making the model better at relating two languages inside one input.
What happens when the two inputs of one task are in different languages? Give the evidence for B-BERT, mBERT and XLM-R, and a possible explanation.
- Standard zero-shot transfer is English fine-tuning, then monolingual testing in another language. Cross-lingual inference within a task means e.g. an English premise with a Spanish hypothesis.
- K et al. (2020), B-BERT on XNLI: fake English pairs 78.5 to 79.3, target-language pairs 59.6 to 70.9, mixed pairs 45.7 to 61.1, under both monolingual conditions (Hindi 45.7, 13.9 points under its own zero-shot score). A fake-English hypothesis beats a fake-English premise.
- Rajaee and Monz (2024), mBERT on XNLI, 15 languages: within 65.7 against across 54.5 on average. A Swahili hypothesis is near chance (40 to 42.5).
- XLM-R: better within and across on XNLI, but across-language QA falls to 36.8 F1 against mBERT’s 44.2.
- Interpretation: the model can do the same task in another language but is poor at relating two languages inside one input, so the space is not language-neutral.
- Heuristics hypothesis: transfer of heuristics can contribute to cross-lingual generalisation. In SNLI full word overlap means entailment 94.7% of the time; a model can learn “high overlap means entailment”, which works in any single language but fails when premise and hypothesis are in different languages, where overlap is near zero. (How exactly this explains the drop is an interpretation, not an established result.)
What do layer-wise analyses reveal about where mBERT holds language-neutral information? Use translation retrieval and parameter freezing, and link to the choice of layer.
- Translation retrieval (Pires et al.): per layer, mean-pool hidden states (excluding
[CLS],[SEP]), shift English vectors by the mean EN-to-DE difference, retrieve the nearest German sentence by . Accuracy follows an inverted U: 20 to 33% at layer 1, peak 71 to 76% at layers 6 to 8, falling in the last layers.- Reading: low layers are close to language-specific surface tokens; middle layers are the most language-neutral (one constant offset maps one language onto the other); top layers are shaped by predicting words in a specific language.
- Freezing (Wu and Dredze 2019): fine-tune on English with the lowest layers frozen. Feature-based use without fine-tuning is always worst (loses 1.9 to 14.8 average points). Freezing the embeddings up to layer 3 or 6 helps or is neutral (best at layer 0/3 for NER, 3 for POS, 6 for MLDoc and XNLI), since English-only fine-tuning would pull lower layers toward English. Freezing up to layer 9 hurts (NER 67.3 against 74.3): upper layers must adapt to the task.
- Both measure the “choose the right layer” strategy for balancing word and context information.
Explain how subword segmentation interacts with masked language modelling and with multilingual models: WordPiece, whole-word masking, and the vocabulary choices of mBERT and XLM-R.
- WordPiece (BERT, mBERT): bottom-up merges by ;
##marks non-initial pieces; inference is greedy longest match; a word with an unmatchable character becomes[UNK]as a whole.- Masking problem: pieces carrying morphology (tense, number, case) are easy to predict from their visible stem, so per-token masking wastes many masked positions on trivially recoverable suffixes.
- Whole-word masking: mask all pieces of a word together at the same overall rate (about 15%); improves BERT-Large on SQuAD 1.1 (uncased F1 91.0 to 92.8) and MultiNLI (86.05 to 87.07).
- mBERT: 120k shared WordPiece vocabulary with smoothed counts, CJK split into characters, still biased toward high-resource languages and Latin script, higher fertility for low-resource and morphologically rich languages; the uncased variant strips accents.
- XLM-R: 250k Unigram LM vocabulary via SentencePiece, fewer pieces per word, at the cost of a larger embedding matrix.
- Wordpiece overlap itself contributes little to transfer (fake English costs 0.5 to 1.4 XNLI points).
What are the two ways of obtaining static word embeddings, and how do they differ in model and context used?
- As a by-product of an actual task: the model can be arbitrarily complex and deep and uses the whole sentence; e.g. the embedding layer of a text classifier or an MT system.
- Directly as the training objective: a simple model (Word2Vec, skip-gram or CBOW, trained with negative sampling) using a limited context window.
- Either way the result is a table with one row per word type, the same vector in every sentence.
In a stacked left-to-right RNN, how does the representation in the column of one word (e.g. train) change from the input up to the top layer, and how does the output configuration affect it?
- Input: the static embedding . First hidden layer: already mixes in the words to its left. Top layer: three rounds of mixing.
- With an output on every position (tagging, next-word prediction) the top state is trained to stay useful for a prediction about its own word.
- With one pooled vector fed by all top states (sentence encoder for a classifier or decoder) no position is supervised alone, so the top layer is pushed to be about the sentence and the word’s own contribution can be diluted.
- Any hidden state in the column is a candidate contextual embedding; which is useful depends on how the network was trained.
State the pre-training phase of the contextual representation workflow, with its notation.
- Model: , a model that can model wider contexts (LSTM, Transformer, CNN), with parameters ; for contextual.
- Task: a general task with very large amounts of training data, e.g. word prediction in language modelling.
- Train on , updating .
State the fine-tuning phase of the contextual representation workflow. What is the difference between feature extraction and full fine-tuning?
- Choose a real task (e.g. QA, POS tagging) and a task model (head) .
- Combine into , whose parameters are the union of both.
- Train it on the real task , updating and maybe also . (Training on again would just be more pre-training; a version stating here is a known typo.)
- Updating only : the pre-trained model is a frozen feature extractor. Updating both: full fine-tuning, BERT’s default.
Name the three Transformer-based architecture families, what each does, and a prominent example of each.
- Encoder-decoder: encoder builds a representation of the input, decoder generates the output the task needs; the original Transformer (Vaswani et al. 2017); summarisation, QA, MT.
- Encoder-only: predicts its own input scrambled by a noise function, in its purest form an auto-encoder; BERT.
- Decoder-only: classical next-word-prediction language model; all current LLMs are decoder-only.
How can an encoder-decoder task be framed in a decoder-only model?
Concatenate input and output into one sequence (e.g.
translate to German: <source> <target>) and train a next-word predictor on it. The encoder’s job is absorbed into the decoder’s processing of the prefix.What does BERT stand for, what part of the Transformer is it, what makes it "bidirectional", and why is its pre-training called self-supervised?
- Bidirectional Encoder Representations from Transformers (Devlin et al., NAACL 2019).
- The encoder part of the Transformer; every position attends to every other position, left and right (unlike a left-to-right RNN).
- Pre-training is missing-word prediction: the labels are the words themselves, so huge amounts of data can be used and the data need not be relevant to any downstream task. Fine-tuning on the real task is supervised and works with small amounts of data.
In the standard BERT pre-training and fine-tuning diagram, what are , and , and what input pair is used for MNLI, NER and SQuAD?
- : input embeddings; : final-layer contextual outputs; : final-layer output at
[CLS].- Pre-training: feeds the NSP classifier, the at masked positions feed the MLM predictor. The same pre-trained parameters initialise every fine-tuning model.
- MNLI: premise and hypothesis. NER: a single sentence. SQuAD: question and paragraph, output a start/end span over the paragraph tokens.
How is BERT's input vector at position formed? Illustrate with my dog is cute / he likes playing.
- Token embedding + segment embedding ( or , telling the model which sentence the token belongs to; the first
[SEP]belongs to A) + positional embedding- Sequence:
[CLS] my dog is cute [SEP] he likes play ##ing [SEP], positions 0 to 10. playing is split by WordPiece intoplay ##ing,##marking a word-internal piece.State BERT's masking rule exactly, including where the loss is computed.
- Select tokens with probability (special tokens
[CLS],[SEP]are not masked).- A selected token is replaced 80% of the time by
[MASK], 10% by a random word, 10% left as the original word.- The loss is computed only at the selected positions, whatever they show (mask, random word or original), and the target is the original token.
Why does BERT not simply replace every selected token by
[MASK]?
[MASK]never occurs at fine-tuning time. A model that only ever sees[MASK]at prediction positions could learn good representations only where[MASK]appears. Random and unchanged tokens mean it cannot tell which positions will be scored, so it must build a good representation of every input token. (This rationale is Devlin et al.’s; the 80/10/10 split itself is the core fact.)Write BERT's MLM loss for one sequence and say what each term means.
- : the selected (scored) positions; : final-layer output of BERT on the corrupted input at position ; : output projection to the vocabulary; : the original token at .
What is the
[CLS]class embedding, and what are its four defining properties?
- Not tied to a specific word and never masked.
- Not meant to learn context-specific word representations.
- Captures a general representation of the entire input sequence.
- Its loss is computed with respect to a classification task (next sentence prediction in pre-training, the task label in fine-tuning).
How is a next sentence prediction training example built, and from which output is it classified?
- Take a document and a segment A from it.
- 50%: B is the segment that follows A (IsNext); 50%: B is from a random document (NotNext).
- Input
[CLS] A [SEP] B [SEP]; loss , with the final-layer output at[CLS].Give BERT's pre-training configuration: batch, sequence content, corpus and tokenisation.
- Batch of 256 sequences of 512 tokens, quoted as “128,000 tokens per batch” (exactly ).
- Sentences A and B are in practice segments of running text, much longer than single sentences.
- BookCorpus (800M tokens) plus English Wikipedia (2.5B tokens).
- WordPiece tokenisation.
Give the sizes of BERT-base and BERT-large: parameters, layers, hidden size, attention heads, per-head dimension.
- BERT-base: 110M parameters, 12 layers, hidden 768, 12 heads, so dimensions per head.
- BERT-large: 340M parameters, 24 layers, hidden 1024, 16 heads, per head.
- A figure of 16 heads for BERT-base is an error.
List the steps of fine-tuning BERT and the typical hyperparameters. Why is fine-tuning kept short?
- Choose a small-ish task-specific model, randomly initialised.
- Connect it to the top-layer token outputs for tasks needing word representations (tagging, span extraction).
- Connect it to the top-layer
[CLS]output for classification (sentiment, entailment).- Train on the task data, with or without updating BERT’s parameters.
Batch size 16 or 32, 2 to 4 epochs. The encoder starts out already good, and long fine-tuning on small data overfits and erodes what pre-training learned.
What is mBERT? Give its origin, architecture, data and how NSP and language mixing are handled.
- Released on GitHub in 2018, with no paper.
- Exactly the BERT architecture and loss (MLM + NSP): no cross-lingual loss, nothing language-specific.
- Trained on the concatenation of Wikipedia in 104 languages, sampled with a smoothing parameter.
- NSP sentence pairs are always within one language; no mixing of languages inside a sequence, though a batch can contain several languages.
- One shared WordPiece vocabulary for all languages.
List what mBERT's training deliberately does not contain, and say what the languages therefore share.
- No parallel data (at least not intentionally).
- No language-ID embedding.
- No language-specific parameters (e.g. adapters).
- No alignment loss rewarding translation equivalents for being close.
- NSP pairs only within one language.
Nothing tells the model that Hund and dog mean the same thing. Languages share only the parameters and the subword vocabulary, so any cross-lingual ability must come from these.
Why does mBERT need weighted language sampling? Give the formula and define each term.
Wikipedia sizes vary enormously; English alone is larger than dozens of smaller languages combined.
- : probability of drawing training data from language .
- : size of language ‘s resource (characters, tokens, etc.).
- : controls how strongly raw sizes are respected.
- The sum over all languages normalises the probabilities.
What happens with language sampling at , and , and why does shrink the gap between languages?
- : proportional to raw size; English and a few big Wikipedias dominate.
- : uniform (); tiny Wikipedias are oversampled so much that the model mostly memorises them.
- : up-weights low-resource and down-weights high-resource languages, a compromise.
- with is concave, compressing large values more than small ones: a size ratio becomes , e.g. 100:1 becomes about 25:1 at 0.7.
How big is mBERT's vocabulary compared to English BERT's, and how does this explain the parameter counts?
- mBERT: one shared vocabulary of 120k entries for 104 languages; English BERT: 30k.
- mBERT has about 178M parameters against BERT-base’s 110M with the same architecture; the difference is due to the vocabulary. Each extra entry adds a 768-dimensional embedding row: .
Describe four special properties or problems of mBERT's shared vocabulary.
- CJK: Chinese characters, Japanese kanji and Korean hanja are split into individual characters before WordPiece.
- Smoothed counts: the language-sampling smoothing also applies to the frequencies used to build the vocabulary, giving small languages more entries.
- Remaining bias: even so, the vocabulary favours high-resource languages and Latin script; low-resource and morphologically rich languages are split into more pieces per word (higher fertility).
- Uncased variant: lowercases and strips accents, damaging languages that depend on diacritics.
What is WordPiece, who proposed it, how does it relate to BPE historically and in direction, and what is its pair score?
- Schuster and Nakajima (2012); very similar to BPE, and predates BPE as an NLP segmentation method (Sennrich et al. 2016), though BPE as a compression algorithm is older.
- Bottom-up like BPE (start from characters, merge); Unigram LM is top-down.
- A pair scores high when it occurs together more often than the frequencies of its parts suggest (association), where BPE merges the raw most frequent pair.
Give the steps of WordPiece training.
- Assume word boundaries; collect word types with frequencies.
- Split each word into characters, prefixing every non-initial character with
##(hug becomesh ##u ##g); the vocabulary starts as all such symbols.- While the vocabulary is smaller than the target size : score every adjacent pair with .
- Merge every occurrence of the best pair into one symbol (the
##of the second part is dropped) and add it to the vocabulary.Give the steps of WordPiece inference for one word. What happens to a word containing an unknown character?
- Start at the first character.
- Try the longest remaining substring first, shortening from the right until a candidate is in the vocabulary; non-initial candidates carry
##.- Append the match and continue after it.
- If no candidate matches at some point, the whole word becomes
[UNK].E.g. playing becomes
play ##ing, tallest becomestall ##est. Inference does not replay merges, so it can segment differently from replaying the training merges.Worked WordPiece step: types hug 10, pug 5, pun 12, bun 4, hugs 5. Which pair is merged first, and which would BPE merge?
- Symbol counts:
h15,##u36,##g20,p17,##n16,b4,##s5.- Pairs:
##u ##g20,p ##u17,##u ##n16,h ##u15,b ##u4: all score (co-occurring with the very common##uis unsurprising).##g ##s: , the highest, because##sonly ever follows##g. WordPiece merges##gs, the rarest pair.- BPE would merge
##u ##g, the most frequent pair (20).Compare BPE, WordPiece and Unigram LM on direction, merge/keep criterion, inference, output and unknown handling, and name typical models for each.
- BPE: bottom-up; raw pair frequency; replay merges in order; deterministic; character/byte fallback; GPT family, RoBERTa, original XLM.
- WordPiece: bottom-up; association score (likelihood gain); greedy longest match; deterministic; whole word becomes
[UNK](the harshest); BERT, mBERT, DistilBERT.- Unigram LM: top-down pruning; loss in corpus likelihood under EM; Viterbi or sampling; deterministic or sampled; character fallback; T5, XLM-R, ALBERT (via SentencePiece).
Why are some masked tokens or subword pieces easier to predict than others, and what does this do to MLM?
- Content words (nouns, adjectives, verbs) are harder to predict than function words (determiners, prepositions).
- The same holds for BPE pieces: masking the stem (
[MASK] ed) leaves a hard question (which verb?), while masking the suffix after a visible stem (view@@ [MASK]) is nearly certain.- Pieces carrying mostly morphological information (tense, number, case) are easier than pieces carrying root/stem information.
- Consequence: with per-token random masking many masked positions are trivially recoverable suffixes and continuation pieces, so the model gets loss signal without learning much and MLM becomes effectively easier.
What is whole-word masking, how is it implemented, and what results did it give for BERT-Large?
- Always mask all subword tokens of the same word together; the masking rate is unchanged (about 15% of tokens).
view@@ edbecomes[MASK] [MASK].- Implementation: group tokens into words, shuffle the words, add whole words to the masked set until the token budget of about 15% is reached, then apply 80/10/10 and score only those positions.
- BERT-Large uncased: SQuAD 1.1 F1/EM 91.0/84.3 to 92.8/86.7, MultiNLI 86.05 to 87.07. Cased: 91.5/84.8 to 92.9/86.7, MultiNLI 86.09 to 86.46. Improves every column; the model can no longer complete a word from its own visible pieces.
What is zero-shot cross-lingual transfer, and how is it read off a fine-tune by evaluate table (Pires et al. 2019)?
Fine-tune on task data in language A only, evaluate on the same task in language B with no task data in B. In the table (rows fine-tuning language, columns evaluation language) the diagonal is ordinary in-language performance and every off-diagonal cell is zero-shot. Pires et al. (2019) studied mBERT this way over 16 languages.
What did Pires et al. find for mBERT zero-shot NER and POS between English, German, Dutch, Spanish and Italian?
- Performance is decent off the diagonal, showing generalisation beyond the fine-tuning language.
- NER: English-trained reaches 77.36 F1 on Dutch against 89.86 in-language (about 86%).
- POS transfers better in absolute terms (English to German 89.40 against 93.99), and closely related pairs best (Spanish to Italian 93.71, Italian to Spanish 91.28).
- Transfer is asymmetric: English to Dutch NER 77.36, Dutch to English 65.46.
How did Pires et al. test whether mBERT's transfer depends on shared script, and what did they find?
- POS between languages with different scripts, which share almost no wordpieces.
- Hindi (Devanagari) to Urdu (Perso-Arabic) 85.9, Urdu to Hindi 91.1; English (Latin) to Bulgarian (Cyrillic) 87.1. mBERT generalises well across scripts.
- Scenarios involving Japanese drop sharply (English to Japanese 49.4, Bulgarian to Japanese 51.6, Japanese to English 57.4), possibly due to larger typological differences.
Give the entity-wordpiece overlap measure of Pires et al. and define its terms.
A Jaccard similarity over wordpieces occurring in labelled named entities only. : entity wordpieces of the fine-tuning data in language ; : entity wordpieces of the evaluation data in language . 0 for disjoint sets, 1 for identical.
What does the scatter plot of zero-shot NER F1 against entity-wordpiece overlap show for mBERT and English BERT, and what follows?
- English BERT: near 0 F1 under about 10% overlap, rising roughly linearly to 40 to 70 F1 at 25 to 38% overlap. It transfers only through shared surface strings.
- mBERT: overlap only 0 to about 27%, F1 between roughly 40 and 82 across the whole range, already 40 to 70 near 0 overlap: essentially flat.
- mBERT’s transfer is not string matching; it has a representation shared across languages beneath the level of surface wordpieces.
What is WALS? Give its size, coverage, structure and use in NLP.
- The World Atlas of Language Structures, a typological database: a large catalogue of how languages are structured (phonological, grammatical, lexical properties), compiled from reference grammars by 55 authors.
- 192 features, each with a small set of values (e.g. Tone: no tones, simple, complex).
- 2,662 languages, but sparse: no language has every feature, no feature covers every language; English has the most (159), still missing 33.
- NLP use: a measure of structural similarity between languages (word order, number of cases), with languages represented as feature vectors.
What did Pires et al. find about typological similarity and mBERT's zero-shot POS transfer?
- Subject/verb/object order: SVO to SVO 81.55, SVO to SOV 66.52 (a 15-point gap for the same fine-tuning data); SOV to SVO 63.98, SOV to SOV 64.22.
- Adjective/noun order: smaller effect; AN to AN 73.29, AN to NA 70.94; NA to NA 79.64.
- Shared WALS features: accuracy rises with the number of common features for both models (mBERT about 58 at 1 feature to 77 at 6; English BERT about 31 to 50), mBERT 27 to 40 points higher, with wide error bars.
- Transfer is best between structurally similar languages.
Describe the translation-retrieval method Pires et al. use to test whether mBERT's sentence representations align across languages.
- Sample translation pairs from WMT16.
- For each layer , represent a sentence by the mean of its hidden activations, excluding
[CLS]and[SEP].- Compute one mean offset for the language pair:
- “Translate” each English sentence as .
- Find the nearest German sentence by distance (cosine is not used); a hit is the true translation. Accuracy = hits / .
What layer-wise pattern does mBERT translation retrieval show for EN-DE, EN-RU and UR-HI, and why?
- All three follow an inverted U: 20 to 33% at layer 1, peak 71 to 76% in layers 6 to 8, then falling.
- EN-RU (different scripts) starts lowest (about 20%) but catches up by layer 8; UR-HI peaks earliest (layer 6) and falls fastest; EN-DE peaks at about 76% at layer 8.
- Lower layers are close to language-specific surface tokens; middle layers are the most language-neutral, so one constant shift maps one language onto the other; top layers are shaped by predicting words in a given language and become language-specific again.
What did Wu and Dredze (2019) set out to test, and with which five tasks?
The desired property: a common, aligned embedding space, where training on language A for task X improves performance on language B for task X with no task data in B (zero-shot), English being the fine-tuning language. Tasks: document classification (MLDoc, document level), entailment (NLI/XNLI, sentence pairs), NER and POS (tokens), dependency parsing (token pairs). 8, 15, 5, 15 and 31 languages respectively.
Describe the MLDoc task and mBERT's results on it (Wu and Dredze).
- Document classification into four classes: CCAT (Corporate/Industrial), ECAT (Economics), GCAT (Government/Social), MCAT (Markets). Only the first two sentences are used, due to memory constraints.
- In-language mBERT averages 91.2, beating Schwenk and Li (2018, 89.5) everywhere except German.
- Zero-shot mBERT averages 74.5, just under Artetxe and Schwenk’s 74.9, which was trained with parallel text.
- Zero-shot costs a lot: 91.2 to 74.5 on average; Japanese drops from 88.4 to 56.5.
Define natural language inference and describe SNLI and XNLI, including their sizes.
- NLI: given a premise and a hypothesis, the premise entails, contradicts or is neutral to the hypothesis: 3-way classification.
- SNLI (Bowman et al. 2015): 500k+ pairs, English only, commonly used for training; gold label is the majority of five annotator judgements, which can disagree in all three directions.
- XNLI: crowd-sourced 5,000 test and 2,500 dev pairs, translated into 14 languages (15 in total, 112.5k annotated pairs); pairing any premise language with any hypothesis language gives more than 1.5M combinations ().
How does mBERT do on XNLI under pseudo-supervision and zero-shot transfer, and what is the fair comparison (Wu and Dredze)?
- Pseudo-supervision (machine-translate the English training set into each target language): 71.6 average, better than zero-shot 66.3.
- Zero-shot mBERT (66.3) beats the older X-LSTM (65.6) but loses to systems using parallel data or bilingual signal: Artetxe and Schwenk 70.2, Lample and Conneau (2019) MLM+TLM 75.1.
- Fair comparison: Lample and Conneau (2019) MLM only, also trained without cross-lingual signal: 71.5 against mBERT’s 66.3.
- Worst languages are low-resource and distant: Swahili 50.4, Thai 55.8, Urdu 58.0, Hindi 60.0, Turkish 61.6.
How does zero-shot mBERT compare with dedicated systems on NER and POS (Wu and Dredze)?
- NER: zero-shot mBERT averages 74.03 F1 (nl, es, de) against 67.13 for the cross-lingual system of Xie et al. (2018), about 7 points better. Chinese collapses to 51.90 against 93.17 in-language.
- POS: zero-shot mBERT averages 84.3, under Kim et al. (2017) with only 320 target sentences (89.9) or 1280 (93.3). A small amount of target-language data is worth more than mBERT’s cross-lingual knowledge. Weakest: Persian 72.8, Dutch 75.9.
Describe the parameter-freezing experiment of Wu and Dredze and the meaning of "Feat" and "Lay ".
- Fine-tune mBERT on English while keeping the lowest layers frozen, , where 0 is the embedding layer; “Lay ” means layers up to are frozen. Evaluate zero-shot on other languages for MLDoc, XNLI, NER and POS.
- “Feat”: the feature-based setting, where mBERT is not fine-tuned at all and only the task head is trained.
What were the findings of Wu and Dredze's layer-freezing experiment?
- Feat is always worst: loses 1.9 (POS), 3.5 (NER), 4.0 (NLI) and 14.8 (MLDoc) average points against the best setting. Fine-tuning is necessary.
- Freezing lower layers (embeddings up to layer 3 or 6) helps or is neutral: best at Lay 0/3 for NER (74.3), Lay 3 for POS (85.2), Lay 6 for MLDoc (77.4) and XNLI (67.1). Frozen lower layers keep their multilingual representation instead of being pulled toward English.
- Freezing up to layer 9 hurts, most for NER (67.3 against 74.3): upper layers need to adapt to the task.
- The effect is largest for MLDoc and small for POS.
Give Wu and Dredze's observed-wordpiece percentages and and define each term.
- : wordpiece types in English training data; : types in language ‘s test data; : test types seen in English training; : frequency of in the test set.
- : % of test types seen; : % of test tokens whose wordpiece was seen.
How strongly does wordpiece overlap with English training data correlate with zero-shot performance per task (Wu and Dredze), and what is the caveat?
- NER: almost perfect, type , token (the most lexical task, entities often copied verbatim; only 5 languages).
- XNLI: no type-level correlation, (token 0.36, not significant): sentence-level inference does not depend on shared wordpieces.
- MLDoc, POS, parsing: in between, 0.5 to 0.8; low-overlap languages score lowest but with wide spread.
- Caveat: correlation is not cause. Languages with high overlap with English are also typologically closer to English, which needs a separate experiment (fake English) to disentangle.
What is fake English (K et al. 2020), how is it made, and why does it isolate the effect of subword overlap?
- K et al. train their own bilingual BERT (B-BERT) on (fake) English plus one other language, so factors can be switched off one at a time.
- Fake English shifts every English Unicode codepoint by a large constant, so no character overlaps with the other language: a bijective mapping.
- It is English in every respect except its characters: same words, grammar, word order and frequencies. Shared vocabulary with the other language becomes exactly zero, so any drop in transfer measures the contribution of shared wordpieces.
What were the fake-English XNLI results of K et al. (2020), and what is the conclusion?
- en-es 72.3 against enfake-es 70.9 (contribution 1.4); en-hi 60.1 against 59.6 (0.5); en-ru 66.4 against 65.7 (0.7).
- B-BERT on English and fake English, fine-tuned on fake English: 78.0 on fake English, 77.5 on real English (0.5).
- Removing all subword overlap costs only 0.5 to 1.4 points: shared wordpieces are not what makes multilingual BERT multilingual.
How did K et al. (2020) test the role of word order, and what did they find?
- Randomly permute a fraction of words (0, 0.25, 0.5, 1.0) during pre-training only; no permutation in fine-tuning. Fine-tune on English, evaluate XNLI in the target language.
- Spanish: 70.9 to 62.5 at full permutation (drop 8.4); Hindi: 59.6 to 43.1 (16.5); Russian: 65.7 to 53.6 (12.1).
- Significant drop, but transfer remains reasonable, well over the 33.3% chance level. Hindi, SOV and furthest from English word order, suffers most.
What did K et al. find when the premise and hypothesis of XNLI were in different languages (B-BERT)?
- Fake English on both sides: 78.5 to 79.3; target language on both sides: 59.6 to 70.9 (the zero-shot results).
- Mixed: enfake-target 57.9 (es), 45.7 (hi), 51.1 (ru); target-enfake 61.1, 55.6, 57.9.
- Mixed pairs score under both monolingual conditions; Hindi falls 13.9 points under its own zero-shot score. A fake-English hypothesis is consistently better than a fake-English premise.
- A truly language-neutral space would make mixed pairs as easy as monolingual ones; they are not.
What did K et al. (2020) find when varying B-BERT's depth, and why does it matter for transfer?
- Depth varied from 1 to 24 layers with the parameter count held roughly constant (132.78M to 139.33M), 12 attention heads.
- Fake-English (in-language) XNLI saturates fast: 66.6 at depth 1, 76.9 at 4, about 79 from 6 on.
- Zero-shot Russian keeps rising: 45.0 (1), 63.1 (6), 67.6 (24).
- The gap shrinks from 21.6 to 11.3. Depth has the largest impact of the architecture factors and buys cross-lingual transfer more than in-language performance: deeper networks learn more language-independent representations.
What do "within" and "across" mean in Rajaee and Monz's (2024) XNLI evaluation of mBERT, and what are the averages?
- Within for language : premise and hypothesis both in (diagonal of the premise by hypothesis matrix).
- Across for : the mean of the cells where exactly one side is in and the other in a different language. The definition is reconstructed from the accuracy matrix rather than stated explicitly:
- mBERT: within 65.7, across 54.5 on average (11.2 lower); Swahili across 45.6.
State the heuristics hypothesis (Rajaee and Monz) and the SNLI word-overlap evidence behind it.
- Hypothesis: transfer of heuristics can contribute to cross-lingual generalisation.
- Rajaee et al. (2022) binned SNLI pairs by premise-hypothesis word overlap: full overlap is entailment 94.7% of the time (17,364 against 963); [0.8, 1.0) 58.2%; falling to 13.9% at (0, 0.2). So high overlap is a strong cue for entailment.
- A model fine-tuned on such data can learn “high overlap means entailment”; that shortcut works within any single language but cannot fire when premise and hypothesis are in different languages. How much of the cross-lingual drop this explains is an interpretation, not a settled result.
Name four patterns in mBERT's full XNLI premise by hypothesis accuracy matrix.
- The diagonal is the maximum of every column: for any hypothesis language a same-language premise is best. Not along rows: for German, Urdu, Hindi, Swahili and Thai premises an English hypothesis beats the in-language one (sw-en 55.0 against sw-sw 50.3).
- An English hypothesis helps: English column averages 64.9, English row 57.7.
- A Swahili or Thai hypothesis is near-hopeless: Swahili column 40 to 42.5, Thai 43.8 to 46.7.
- Related languages pair well: ru-bg 62.0, bg-ru 64.9, es-fr 68.3, fr-es 69.3, hi-ur 55.6, ur-hi 53.5.
List the four changes from mBERT to XLM-R and the reason for each.
- Objective: MLM only, NSP dropped, because NSP contributes little and sometimes hurts.
- Masking: dynamic instead of static.
- Data: CC-100, 2.5TB of filtered CommonCrawl in 100 languages, two orders of magnitude more than Wikipedia.
- Vocabulary: 250k entries (double mBERT’s 120k), Unigram LM via SentencePiece, for better fertility (fewer pieces per word, especially for languages starved in mBERT’s vocabulary), at the cost of a bigger embedding matrix.
- XLM-R is by Conneau et al. (2020), a RoBERTa-style extension of mBERT.
Distinguish static from dynamic masking. Why is dynamic masking better?
- Static (BERT, mBERT): masking is done once at the data level during pre-processing; every epoch reuses the same masked version.
- Dynamic (RoBERTa, XLM-R): a fresh mask is drawn each time a sequence is fed to the model.
- With static masking a token not selected in pre-processing is never a prediction target however many epochs run; dynamic masking gives new training signal on every pass at no extra data cost.
Give the three GLUE-style example tasks (sentiment, Winograd, reading comprehension) and say what makes each hard.
- Sentiment: Skip the film and buy the Philip Glass soundtrack CD is negative, though it contains no negative word.
- Winograd schema: The trophy doesn’t fit into the brown suitcase because it is too large: it = the trophy; change large to small and it = the suitcase. Same syntax, so it needs world knowledge.
- Reading comprehension: question At what pressure is water heated in the Rankine cycle? over a paragraph; the answer is the span from word 46 to 47, high pressure. Nothing is generated.
How is BERT fine-tuned for SQuAD span prediction? Give the new parameters, the probabilities and the decision rule.
- Input
[CLS] question [SEP] paragraph [SEP]. The only new parameters: a start vector and an end vector .- : final-layer output at paragraph token ; : hidden size (768 for base); forbids spans that end before they start.
Give the steps of decoding a SQuAD answer span from fine-tuned BERT.
- Run BERT on
[CLS] q [SEP] p [SEP]to get final-layer outputs .- For each paragraph position compute a start score and an end score .
- Over all pairs (in practice capped at a maximum answer length) pick the pair maximising start score + end score.
- Return paragraph tokens to .
Give the average within and across scores for mBERT and XLM-R on XNLI and XSQuAD (Rajaee and Monz 2024), and the result to remember.
- mBERT XNLI: 65.7 / 54.5 (gap 11.2). XLM-R XNLI: 74.2 / 64.8 (gap 9.4).
- mBERT XSQuAD F1: 64.4 / 44.2 (gap 20.2). XLM-R XSQuAD: 72.0 / 36.8 (gap 35.2).
- XLM-R wins within on both tasks and across on XNLI, but is worse across languages on QA than mBERT (36.8 against 44.2) despite being 7.6 better within.
How do the XSQuAD context by question matrices of mBERT and XLM-R differ?
- mBERT degrades smoothly: an English question works with every context (strongest column), a Thai question works with none (18.8 to 23.4 off-diagonal), Thai context is weak for every question.
- XLM-R: strong diagonal and strong English-question column (58.2 to 75.0), almost everything else collapses (Arabic context with a non-English, non-Arabic question 14.6 to 36.7; Chinese context with Thai or Turkish question 17.6). Structure: own language, or English question.
- In XLM-R an English context with a foreign question is much weaker than a foreign context with an English question (en-ar 38.1 against ar-en 59.2).
- On XNLI, by contrast, XLM-R lifts the whole matrix (off-diagonal mean 65.1 against 54.5); the Swahili hypothesis column stays weakest (44.5 to 51.7).
Link to originalSummarise the conclusions on what drives mBERT's cross-lingual transfer, and what could be missing.
- Shared (sub)word vocabulary plays a role but is not the main cause (fake English costs 0.5 to 1.4 XNLI points; transfer works across scripts).
- Structure (grammar, word order): permuting words hurts significantly without eliminating transfer; typologically closer languages transfer better.
- Architecture: deeper networks learn more language-independent representations.
- No single factor explains transfer (or the lack of it).
- What is missing is left as an open question; the obvious candidates are what mBERT lacks: parallel data and an alignment loss pulling translation equivalents together, since neither mBERT nor XLM-R handles mixed-language inputs well.
L07 Crosslingual NLP
Flashcards
Click a question to reveal its answer, or press Study to drill the whole set. Cards marked as exam questions are meant to be answered out loud or on paper first, then checked against the points listed.
Why does multilingual pretraining (mBERT, XLM-R) not by itself make a model crosslingual, and what evidence shows that explicit crosslingual objectives fix it?
Must hit:
- The mechanism. Multilingual models train jointly on many languages with one model and one subword vocabulary, but MLM predicts a word only from same-language context. When predicting word in language , unrelated context in is of no use, so nothing in the objective links languages. Any alignment is an emergent side effect.
- Why it looked fine. Standard zero-shot benchmarks (fine-tune on English, test on another language) keep all parts of the input in one language, and mBERT does reasonably there.
- The harder test. Put premise and hypothesis (XNLI) or context and question (XSQuAD) in different languages and compare within (both parts in one language) with across (mixed).
- The numbers. QA F1 within/across: mBERT 64.4/44.2, XLM-R 72.0/36.8, InfoXLM 73.8/64.5. XLM-R is better than mBERT in every single language but worse across languages.
- The fix. InfoXLM adds TLM and XLCo on parallel data: within improves only +1.2 (XNLI) and +1.8 (QA) over XLM-R, across improves +5.5 and +27.7. The QA gap shrinks from 35.2 to 9.3.
Losing marks: treating “multilingual” and “crosslingual” as the same thing, or citing only standard zero-shot scores as proof of crosslingual ability.
You have English-only training data for NLI and need a system for 14 other languages. Compare translate-train, translate-test and zero-shot transfer, using the XLM results on XNLI.
- Translate-train: machine-translate the English training set into each language and fine-tune on the translation. XLM (MLM+TLM) average 76.7, the best of the three, but needs an MT system and a translated training set for every language.
- Translate-test: machine-translate each test example into English and apply an English model. XLM (MLM+TLM) average 74.2.
- Zero-shot crosslingual transfer: fine-tune on English only, test directly on each language. XLM (MLM) 71.5, XLM (MLM+TLM) 75.1.
- Key point: a single zero-shot model that never saw non-English NLI data beats the translate-test pipeline (75.1 against 74.2), and TLM is what gets it there (+3.6 over MLM alone).
- Zero-shot XLM (MLM+TLM) also beats mBERT and LASER in every language where they are reported.
Losing marks: mixing up which side gets translated (train set into the target language, or test set into English).
Explain translation language modeling (TLM): its input, its loss, why it forces crosslingual alignment, and how it differs from MLM and CLM.
- XLM (Conneau and Lample, 2019) has three objectives: CLM (predict the next word from a prefix), MLM (predict masked words from monolingual context, as in BERT), TLM (predict masked words in a sentence concatenated with its translation). TLM is always paired with MLM or CLM; NSP is dropped.
- Loss: TLM is MLM applied to : The loss is the same as MLM’s; only the input changes.
- Why it aligns: in “the [MASK] [MASK] blue” / “[MASK] rideaux étaient [MASK]”, the English context barely constrains curtains, but the unmasked French rideaux gives it away. The cheapest way to lower the loss is to attend across languages and learn rideaux ↔ curtains. Context in becomes useful for predicting in .
- Two input details: position embeddings restart at 0 in the second sentence (so position cannot separate the halves, and corresponding words get similar position signals); language embeddings (en, fr) mark which half is which.
- Evidence: zero-shot XNLI average 71.5 (MLM) to 75.1 (MLM+TLM).
How would you build a parallel corpus from the web, and how do LASER embeddings make sentence alignment possible?
- Sources first: natural by-products (multilingual news such as Xinhua, the UN and EU, websites of multilingual countries such as Canada and Belgium) and OPUS, a large research collection of parallel corpora.
- Crawl pipeline: (1) document alignment (find parallel documents), (2) sentence alignment inside them. Or mine sentences directly from raw crawls such as Common Crawl.
- Why sentence alignment is separate: segments do not follow paragraph boundaries (in the NHK swine fever example, two English paragraphs map into one Chinese paragraph, which must be split at a sentence boundary).
- Measuring equivalence: old methods use dictionary overlap and relative length; the modern method compares LASER sentence embeddings (Artetxe and Schwenk, 2019).
- LASER: BPE embeddings, stacked BiLSTM encoder, max pooling into one sentence vector; an LSTM decoder translates from that vector alone and is told the output language by a language ID embedding. So the vector must encode meaning and has no reason to encode the input language. Translations end up as near neighbours.
- vecalign uses LASER similarities to match sentences, with an efficient search that scales to massive data sets.
Explain neural machine translation as conditional language modeling, from the RNN encoder-decoder to the Transformer encoder-decoder.
- Seq2seq drops two assumptions of sequence labeling: that corresponds to , and that . MT needs both dropped (Hiermit hörte sie nicht auf / She did not stop with this: different lengths, reordering, one-to-many, many-to-one).
- It is conditional language modeling: , with the encoder output and the decoder’s prefix. The sequence probability is the product of these terms over .
- Data: source ids ; decoder input =
<s>+ sentence; target = sentence +</s>( shifted by one).- RNN encoder-decoder (Sutskever et al., 2014, LSTMs): . Everything about the source must pass through one fixed-size vector: an information bottleneck.
- Transformer: each decoder layer has a target context layer (masked self-attention over the target prefix), a source-target context layer (attention over all top-layer encoder outputs), and a feed-forward layer, each with a residual connection. Every decoder position in every layer can read every source position, which removes the bottleneck.
Describe parent-child transfer learning for low-resource NMT. What should be frozen, and which factors decide whether transfer pays off?
- Procedure (Zoph et al., 2016): train a parent on a high-resource pair (French→English, 300M English tokens); copy all parameters into the child (Uzbek→English, 1.8M tokens); child source words take over rows of the parent’s source embedding matrix; freeze the English embeddings; continue training with strong regularisation (dropout 0.5).
- Gains: Hausa +4.5, Turkish +5.6, Uzbek +3.7, Urdu +8.6 BLEU; the smallest corpus (Urdu) gains most.
- Freezing: train everything up to and including attention (Uz→En dev 15.0), but keep the target embeddings frozen (unfreezing them drops to 14.7, then 13.7). The general rule “freeze the decoder” overstates it: training the target RNN raises 11.8 to 14.2.
- Relatedness: a Spanish child gets 31.0 with a French parent, 29.8 with German, 16.4 with none. But French’ (scrambled vocabulary) still gains 13.3 to 20.0, so structure transfers and shared words are not the only factor.
- Shared vocabulary (Kocmi and Bojar, 2018): unrelated parents (Czech, Russian) help English→Estonian as much as related Finnish.
- Direction: transfer must flow from the larger corpus to the smaller; reversed, it helps little or hurts.
- What transfers (Aji et al., 2020): inner layers carry most of the benefit; parent embeddings alone are worse than training from scratch.
Describe BART: its architecture, its pretraining noise functions, which noise works best, and how it is fine-tuned for classification, span prediction and translation.
- Architecture: bidirectional encoder (as BERT) plus autoregressive decoder (as GPT). The encoder reads a corrupted document, the decoder reconstructs the original through cross-attention. Suited to tasks needing both, such as translation and summarisation.
- Noise functions: token masking, token deletion, text infilling (a span replaced by one
[MASK]), sentence permutation, document rotation.- Best: text infilling (SQuAD 90.8, best XSum and ConvAI2 perplexity); deletion beats masking on all generation tasks; rotation and sentence shuffling alone are poor (SQuAD 77.2 and 85.4). Infilling plus shuffling gives the best CNN/DM perplexity (5.41).
- Classification: same input to encoder and decoder; label predicted from the last decoder hidden state. Span prediction (SQuAD): label each token, predict start and end of the answer.
- MT: a randomly initialised source encoder replaces BART’s embedding layer; BART can be frozen or updated; the source vocabulary can differ from BART’s. Ro→En: baseline 36.80, Fixed BART 36.29, Tuned BART 37.96.
- Results: matches RoBERTa on SQuAD (94.6 F1 on 1.1), best on all ROUGE columns for CNN/DM and XSum.
Why does crosslingual QA expose the weakness of purely multilingual models much more than XNLI does? Use the XLM-R and InfoXLM results.
- XNLI compares two sentences; a coarse sentence-level gist in a shared space is often enough.
- Extractive QA requires finding the exact answer span, which means matching the question’s words to specific context words. With a Hindi question and Arabic context that is word-level matching between two non-English languages, never asked for by monolingual MLM, and exactly what TLM trains.
- XLM-R QA: diagonal strong (63.7 to 84.2), English-question column 58.2 to 75.0, but most other off-diagonal cells collapse (Arabic context with non-English question 14.6 to 36.7; Chinese context 16.0 to 32.0). Across average 36.8, under mBERT’s 44.2.
- InfoXLM QA: every cell at least 51.7; Arabic and Chinese contexts with non-English questions now 51.7 to 63.8.
- Gain of InfoXLM over XLM-R in the across score: +5.5 on XNLI, +27.7 on QA.
Which factors predict whether crosslingual transfer will succeed? Support each with evidence.
- Language relatedness: Nepali perplexity 157.2 alone, 140.1 with English, 115.6 with Hindi; Spanish NMT child 31.0 with a French parent against 29.8 with German.
- Being close to English: English dominates pretraining and fine-tuning data, so the English column is the brightest in every XNLI heatmap and English cells are the darkest in the WikiMatrix BLEU grid; directions between two non-English languages are mostly 5 to 25 BLEU.
- Representation of the language and its script: Swahili hypotheses leave mBERT barely over chance (40.2 to 42.5); mBERT cannot handle Thai questions in QA (18.8 to 23.4), while XLM-R has no such Thai problem.
- An explicit crosslingual training signal: TLM and XLCo on parallel data close most of the within/across gap (InfoXLM).
- Amount and direction of data in NMT transfer: transfer helps when it flows from a large parent to a small child; reversed it helps little or hurts. With a shared vocabulary and a large parent, relatedness matters less (Kocmi and Bojar).
- Which parameters transfer: inner layers carry most of the benefit (Aji et al.).
In which four senses are models such as mBERT and XLM-R "multilingual", and which one is the weakness for crosslingual ability?
- They train jointly on multiple languages.
- They use one model for all languages.
- They use one (subword) vocabulary for all languages.
- Their predictions are based on context within the same language (MLM).
The fourth is the weakness: predicting a masked German word, every visible token is German, so nothing rewards knowing its English translation.
Define crosslingual knowledge transfer and zero-shot crosslingual transfer, and say why this is the common scenario in practice.
- Crosslingual knowledge transfer: given fine-tuning data in language A for task X, how well does the model generalise to test data for task X in language B?
- Zero-shot crosslingual transfer: the fine-tuning and test languages differ and no task data in B was seen at all.
- Common because fine-tuning requires annotated data, which is scarce outside high-resource languages. It is a crosslingual capability: the model must map what it learned in A onto B.
Can crosslingual capabilities emerge without any crosslingual training signal? Give the evidence on each side.
- For: earlier zero-shot results; mBERT fine-tuned on English works to a degree on other languages.
- Against: crosslingual tasks where the input itself mixes languages (premise in one, hypothesis in another) make purely multilingual models degrade sharply.
- Verdict from the within/across results: much of the capability does not emerge without a crosslingual signal, and adding one fixes most of it.
Why does training on many languages not link them, and what is the fix?
- Training on multiple languages makes a model multilingual, but nothing explicitly links information across languages. When predicting word in language , unrelated context in gives no benefit: the languages share parameters but never share context.
- Fix: context alignment across languages. Give the model text in two languages known to correspond (parallel or comparable data), so context in actually helps predict in .
Distinguish parallel data from comparable data, with examples and degree of parallelism.
- Parallel: sentence pairs in different languages that are meaning equivalent (translations). Fully parallel by construction. Used by data-driven MT for decades.
- Comparable: sentence or document pairs on the same topic. Not fully parallel; parallelism is a sliding scale. Examples: news articles on the same event, Wikipedia articles on the same entity in different languages.
What does a Chinese-English parallel corpus excerpt with 新加坡 / Singapore highlighted in every row illustrate?
Once sentence pairs are aligned, recurring co-occurrences (新加坡 always paired with Singapore) let a model learn that the two correspond without anyone writing a dictionary. Statistical MT used this historically; TLM exploits the same signal.
Where does parallel data come from naturally, what is OPUS, and what are the two ways to crawl your own?
- Natural by-products: multilingual news (Xinhua); international organisations (UN, EU translate every official document); websites and documents from countries with two or more languages (Canada, Belgium, US).
- OPUS: a website with a very large collection of parallel corpora for research, covering many languages. The first place to look.
- Crawling: (1) find parallel documents (document aligning), then (2) parallel sentences within them (sentence aligning). Or directly align sentences from vast raw web crawls such as Common Crawl.
What does the NHK World swine fever example (English and Chinese news pages) show about parallel documents and sentences?
- The two pages are a parallel document pair: same story, same photo, published a day apart (time zone).
- Inside, segments align as parallel sentences, but not paragraph to paragraph: English paragraphs 2 and 3 both map into Chinese paragraph 2, which must be split at a sentence boundary.
- So sentence alignment is a separate step after document alignment and must allow alignments other than one paragraph to one paragraph.
How can degrees of meaning equivalence between sentences in two languages be measured? Give the old-fashioned and the recent approach.
- Old-fashioned: dictionary overlap (how many words in sentence 1 have a dictionary translation in sentence 2) and relative length distribution (translations have predictable length ratios).
- Recent: compute and compare sentence embeddings, specifically LASER (Artetxe and Schwenk, 2019): embed all sentences in one shared space, where translations are nearest neighbours.
Describe the LASER encoder, step by step.
- Input tokens are looked up in a BPE embedding table shared by all languages.
- They pass through a stack of BiLSTM layers.
- The top layer’s hidden states are max-pooled over time (element-wise maximum across positions).
- The result is one fixed-size vector, the sentence embedding.
After training on translation, only this encoder is kept.
In LASER, how does the sentence embedding reach the decoder, and what else does the decoder receive at each step?
- Through a linear map that initialises the decoder LSTM.
- By being concatenated to the input at every decoder step, together with the BPE embedding of the previous output token () and a language ID embedding saying which language to produce.
- Each output step ends in a softmax over the vocabulary.
Why does LASER's training produce language-independent sentence vectors?
The decoder sees nothing of the source except the single sentence vector, and it is told the output language by , not by the encoder. So the encoder has no reason to encode the input language and every reason to encode only the meaning, which is all the decoder needs to translate into any target language. After training the decoder is discarded, and the encoder maps all training languages into one space where translations land close together.
What is vecalign?
A sentence alignment approach that (1) uses LASER to match sentences, scoring a source and target sentence by the similarity of their LASER embeddings, and (2) uses an efficient search that scales to massive data sets.
What problem with mBERT motivated XLM, and what does XLM do about it?
- mBERT never sees an explicit translation pair during pretraining; any crosslingual alignment is an emergent side effect.
- For many language pairs at least some parallel text exists (OPUS, self-crawled corpora).
- XLM (Conneau and Lample, 2019) uses parallel data for an explicit crosslingual loss (TLM).
List XLM's three pretraining objectives (input and prediction for each), how they are combined, and where its parallel data comes from.
- CLM (causal LM): prefix in, next word out. Not crosslingual.
- MLM (masked LM): monolingual context with masked words, predict them as in BERT. Not crosslingual.
- TLM (translation LM): a sentence and its translation, predict masked words in either language. Crosslingual.
- Next sentence prediction is dropped. TLM is always combined with a monolingual objective: MLM + TLM or CLM + TLM.
- Parallel corpora from OPUS (e.g. MultiUN, OpenSubtitles, EUbookshop).
Write the CLM and MLM losses and define every term.
- is the token sequence.
- is the set of masked positions; is the sequence with those positions replaced by
[MASK].- is the Transformer’s softmax output over the shared vocabulary.
Write the TLM loss for a parallel pair and explain how it relates to MLM.
- is a sentence in language 1, its translation; , are the masked positions in each.
- Every masked word, in either language, is predicted from both (partially masked) sentences.
- TLM is literally MLM on the concatenation : only the input changes.
In XLM's MLM training example, what does the input stream look like, and what three embeddings make up each input position?
- A continuous stream of sentences separated by
[/s], cut into a fixed-length window: ”[/s] take a seat [/s] have a drink [/s] now relax and”, with take, a[/s], drink and now masked as targets. A sentence separator can itself be masked.- Each position is the sum of token, position and language embeddings; in MLM all language embeddings are the same (en).
In XLM's TLM example ( the curtains were blue / les rideaux étaient bleus), which words are masked, and what are the two input details that make crosslingual prediction work?
- Masked targets: curtains, were (English) and les, bleus (French).
- curtains is hard from “the ___ ___ blue” but easy from French rideaux; les and bleus can be read off English the and blue.
- Positions restart at 0 for the French sentence (both halves use positions 0 to 5), so position cannot tell the halves apart and roughly corresponding words get similar position signals.
- Language embeddings (en, fr) mark which half is which language, since positions no longer do.
What does XNLI measure, how is it scored, and what is the column in the XLM results?
- Crosslingual natural language inference: given a premise and a hypothesis, a three-way classification (chance 33.3%), evaluated in 15 languages.
- Metric: accuracy.
- : the plain average accuracy over the 15 languages.
Give XLM's average XNLI accuracy in each setting, and the main conclusions.
- Translate-train, XLM (MLM+TLM): 76.7 (best overall).
- Translate-test, XLM (MLM+TLM): 74.2.
- Zero-shot, XLM (MLM): 71.5; zero-shot XLM (MLM+TLM): 75.1; LASER (Artetxe and Schwenk) 70.2; Conneau et al. (2018b) 65.6. mBERT was reported for only six languages.
- Conclusions: TLM is worth +3.6 on average and every language gains; zero-shot XLM beats translate-test; translate-train is still best but needs MT for every language; zero-shot XLM beats mBERT and LASER in every reported column.
Which languages gain most and least from adding TLM to MLM in zero-shot XNLI, and what pattern does that show?
- Largest gains: vi 71.2 → 76.1 (+4.9), tr 67.8 → 72.5 (+4.7), ar 68.5 → 73.1 and zh 71.9 → 76.5 (+4.6 each).
- Smallest: English (+1.8), fr and ru (+2.2 each).
- Pattern: TLM helps most for languages distant from English in script or structure.
Worked example: compute the zero-shot XNLI average of XLM (MLM+TLM) from its 15 per-language scores.
Scores: 85.0, 78.7, 78.9, 77.8, 76.6, 77.4, 75.3, 72.5, 73.1, 76.1, 73.2, 76.5, 69.6, 68.4, 67.3.
What do XLM's Nepali language modeling results show? Give the perplexities.
Nepali perplexity (lower is better): Nepali alone 157.2; + English 140.1; + Hindi 115.6; + English + Hindi 109.3. Any second language helps, a related language helps far more (Hindi: same Devanagari script, closely related; cuts 41.6 points against English’s 17.1), and both together are best. Transfer works best between related languages.
How do XLM's word embeddings compare with MUSE and Concat on crosslingual word alignment, and what are the baselines?
- Baselines use fastText embeddings. Concat = joint training of embeddings over all languages; MUSE = the mapping-based method for crosslingual static embeddings. XLM’s word embeddings are its input embedding table.
- Cosine similarity (higher better): MUSE 0.38, Concat 0.36, XLM 0.55. L2 distance (lower better): 5.13, 4.89, 2.64. SemEval’17 (higher better): 0.65, 0.52, 0.69.
- XLM’s embeddings are better aligned on all three, though never trained to align word embeddings directly.
What is InfoXLM built on, and what are its three objectives?
InfoXLM (Chi et al., 2021) builds on XLM-R (a purely multilingual masked LM) and adds XLM’s parallel-data ideas.
- MMLM (multilingual masked LM): described as similar to MLM but using negative sampling instead of directly optimising the ground-truth loss.
- TLM, as in XLM.
- XLCo (crosslingual contrastive learning): take the
[CLS]representations of the two sentences of a parallel pair, classify whether the pair is parallel, with random sentences as negatives.What does the "negative sampling" in InfoXLM's MMLM description actually mean, according to the InfoXLM paper?
MMLM is ordinary masked LM on monolingual text in many languages, the same loss as XLM-R. The paper interprets the softmax cross-entropy over the vocabulary as a contrastive (InfoNCE) loss: the correct token is the positive, every other vocabulary entry acts as a negative. Nothing extra is sampled. (This reading comes from the paper; the short description of MMLM suggests a replaced, sampled loss.)
Write the XLCo contrastive loss and define every term.
- : a parallel pair (the positive). : the
[CLS]representation. : negatives, random sentences that are not translations of .- The fraction is a softmax over “which candidate is the translation of ?”; minimising it pulls translations together and pushes non-translations apart.
- This is the standard InfoNCE formalisation of the description; the exact formula is not given in the original source.
Contrast token-level and sentence-level crosslingual alignment, and say which of TLM, XLCo and LASER does which, and how.
- Token level: TLM. A masked word is predicted from its translation’s words.
- Sentence level: XLCo. A sentence’s whole representation and its translation’s must be closer to each other than to anything else (contrastive loss inside masked-LM pretraining).
- Sentence level: LASER. Gets sentence alignment from a translation decoder that sees only the sentence vector.
Compare XLM and InfoXLM on starting point, data, vocabulary, language embeddings, objectives and type of crosslingual signal.
- Start: XLM from scratch; InfoXLM initialised from XLM-R, then further pretrained (150K steps base, 200K large).
- Monolingual data: Wikipedia; CC-100 rebuilt (94 languages).
- Parallel data: XLM, the 15 XNLI languages; InfoXLM, 14 English-centric pairs, about 42 GB (MultiUN, IIT Bombay, OPUS, WikiMatrix).
- Vocabulary: shared BPE; XLM-R’s 250k SentencePiece.
- Language embeddings: yes; no.
- Objectives: MLM (or CLM) + TLM; MMLM + TLM + XLCo, equally weighted: .
- Signal: token level only; token and sentence level.
What does "English-centric" parallel data mean for InfoXLM, how does WikiMatrix connect to LASER, and why does InfoXLM have no language embeddings?
- English-centric: every parallel pair has English on one side (en-fr, en-de, …); there is no direct fr-de data.
- WikiMatrix is parallel sentences mined from Wikipedia with LASER (Schwenk et al., 2019), so the LASER mining pipeline feeds InfoXLM.
- InfoXLM starts from XLM-R’s weights and cannot change the architecture without losing them, so it inherits the 250k SentencePiece vocabulary and the lack of language embeddings. TLM must then tell the two halves apart from the tokens alone.
Define the within and across scores for a language in mixed-language XNLI and QA evaluation, with the formula.
- Within: score with both input parts in , the diagonal cell .
- Across: mean over all mixed combinations involving in either position (row and column without the diagonal):
- = number of languages (15 for XNLI, 11 for QA); = score with the first part (premise or context) in and the second (hypothesis or question) in .
- The definition is not stated in the source; it is recovered from the heatmaps, which it reproduces (row only or column only does not).
Worked example: mBERT's XNLI row for English (premise English, 14 other hypothesis languages) sums to 808.5 and its English column to 908.4. Compute across(en) and interpret it.
Within(en) = 81.5, so mBERT loses about 20 points once one sentence is not English. Asymmetry: column mean 64.9 against row mean 57.7, so an English hypothesis with a foreign premise is easier than the reverse.
Give the average within, across and gap for mBERT, XLM-R and InfoXLM on XNLI (accuracy) and QA (F1).
- mBERT: XNLI 65.7 / 54.5, gap 11.2; QA 64.4 / 44.2, gap 20.2.
- XLM-R: XNLI 74.2 / 64.8, gap 9.4; QA 72.0 / 36.8, gap 35.2.
- InfoXLM: XNLI 75.4 / 70.3, gap 5.1; QA 73.8 / 64.5, gap 9.3.
What is paradoxical about XLM-R against mBERT in the within/across evaluation?
XLM-R is a much better multilingual model (within +8.5 on XNLI, +7.6 on QA), but on mixed-language QA it is worse than mBERT: across 36.8 against 44.2. Being better at each language separately did not make it better at relating two languages to each other.
In the XNLI heatmaps, which hypothesis language is easiest and which hardest, and why? Give the numbers.
- English column brightest in every model (premise in any language, English hypothesis): mean 64.9 (mBERT), 74.1 (XLM-R), 77.8 (InfoXLM). English dominates pretraining and fine-tuning data, so every language is best aligned to it.
- Swahili column darkest: mBERT 40.2 to 42.5 whatever the premise (barely over chance, 33.3%), XLM-R 44.5 to 51.7, InfoXLM 56.0 to 64.9.
Name three further patterns in the XNLI premise/hypothesis heatmaps.
- Not symmetric: swapping which language holds premise and hypothesis changes the score. mBERT with a Swahili premise (row mean 49.5) does better than with a Swahili hypothesis (column mean 41.6).
- Mixing two related high-resource languages costs little: mBERT (fr premise, en hypothesis) 73.4 against fr-fr 73.5; XLM-R (fr, en) 78.8, higher than fr-fr 78.3.
- InfoXLM is uniformly brighter off the diagonal: its worst off-diagonal cell (56.0) beats every cell in XLM-R’s Swahili column (at most 51.7); its best off-diagonal cell is (es, en) 81.6.
What does mBERT's crosslingual QA heatmap show for Thai?
mBERT cannot handle Thai questions: the Thai column is 18.8 to 23.4 for every non-Thai context, and even Thai-Thai is only 40.0. XLM-R does not share the problem (Thai-Thai 66.5). A plausible reason is poor coverage of Thai script in mBERT’s WordPiece vocabulary.
Describe XLM-R's crosslingual QA heatmap and InfoXLM's, cell by cell pattern.
- XLM-R: works across languages only when the question is English (en column 58.2 to 75.0). With an Arabic context, non-English questions score 14.6 to 36.7; with Chinese, 16.0 to 32.0. Diagonal strong (63.7 to 84.2).
- InfoXLM: fills the matrix in. Every cell at least 51.7 (zh context, hi question); English context gives 70.6 to 79.3 for any question language; the Arabic and Chinese cells XLM-R left at 15 to 37 are now 51.7 to 63.8.
Give the three paradigms of machine translation with their periods, and explain the overlap.
- 1950s to 1990s: rule-based, symbolic.
- 1990s to 2016: statistical, data-driven.
- 2014 to now: neural, deep learning, data-driven.
The 2014 to 2016 overlap is deliberate: neural MT appeared in research in 2014, statistical systems stayed in production until around 2016. MT has been active AI research since the start and illustrates AI’s paradigm shifts.
How does sequence-to-sequence modeling differ from sequence labeling, and how is it written as conditional language modeling?
- It does not assume an isomorphic relationship between and , and does not assume . It models the complex mapping between and .
- : the encoder’s representation of the input. : the decoder’s representation of the output before . : the next output token.
- By the chain rule, ; training minimises over parallel pairs.
What was Sutskever et al.'s (2014) neural MT model, and how did it connect encoder and decoder?
MT as sequence to sequence with an LSTM encoder and an LSTM decoder (two-layer stacks in the figure). The encoder reads ; the decoder starts from a start symbol, emits , feeds it back, emits , and so on until
<eos>. Connection: the decoder LSTM is initialised with the last state of the encoder LSTM.Using Hiermit hörte sie nicht auf. → She did not stop with this., name the four complex mappings that make MT hard.
- Different lengths (6 tokens against 7).
- Word order differences (hörte … auf brackets the clause; sie moves from position 3 to 1).
- One-to-many: Hiermit → with this, nicht → did not.
- Many-to-one: hörte … auf → stop (the separable verb aufhören).
In sequence labeling every input has one output directly over it; in MT alignment links cross.
How are a parallel sentence pair represented as vectors , , for NMT training?
- Both are tokenized ( and tokens); each token is its index in vocabulary or .
- Foreign sentence: .
- Target as two vectors: , , last .
- (
<s>She did not stop with this .) is what the decoder reads; (She did not stop with this .</s>) is what it must predict. At step it has read the start symbol and first words and predicts word .- Strictly, with English tokens and have entries (7 tokens, 8 entries).
How does the RNN encoder-decoder connect its two halves, and what problem does that cause?
- Encoder represents as a whole; decoder reads token by token and learns to predict .
- Connection: (decoder’s initial state = encoder’s state after the last source token).
- Problem: an encoder-decoder information bottleneck. Everything about the source must fit through one fixed-size vector, however long the sentence. Attention removes it by letting every decoder step look at all encoder states.
Give the steps for training an RNN encoder-decoder on one sentence pair.
- .
- For : .
- (the bottleneck, ); loss .
- For : with the gold previous word; over ; loss loss .
- Update parameters by gradient descent on the loss.
At test time there is no gold : start from
<s>and feed each predicted word back in until</s>.Describe one decoder layer of the Transformer encoder-decoder, component by component.
- Each target token is turned into two vectors (drawn cyan and green and unlabelled; their role fits the keys and values of attention).
- Target context layer: masked self-attention; a position attends only to itself and earlier positions. Residual ”+“.
- Source-target context layer: attends over all top-layer encoder outputs (cross-attention). Residual ”+“. This replaces the RNN’s single-vector bottleneck.
- Feed-forward layer, per position. Residual ”+“.
- Output split again into vector pairs for the next layer; after layer , the output layer’s softmax over the target vocabulary.
What does NMT quality depend on, and what are two typical low-resource NMT problems?
- Quality depends on: the amount of data, the amount of variation within the data, and the relevance of the data for the actual task.
- Low-resource problems (some of these conditions unmet): domain adaptation (data in the wrong domain) and NMT for low-resource language pairs (little data at all).
How far are we from universal machine translation (Schwenk et al., 2019), and what does the BLEU grid behind the claim show?
- 86% of all language directions are of poor quality. Core problem: limited parallel training data for most directions, and current MT models do not generalise very well beyond the training data, and not at all beyond specific language directions.
- Grid (26 languages, from the WikiMatrix paper): cells involving English darkest (e.g. en with da, fr, no 41.2); non-English directions mostly 5 to 25 BLEU, except related or high-resource pairs (da and no 27.5 and 30.4); Korean and Japanese weakest (ko column 1.2 to 4.1); many cells empty for lack of mined data.
What does Zoph et al. (2016) show about how data hungry NMT is? Give the numbers.
BLEU into English, syntax-based SMT against NMT trained on child data only:
- Hausa (1.0M tokens): 23.7 / 16.8
- Turkish (1.4M): 20.4 / 11.4
- Uzbek (1.8M): 17.9 / 10.7
- Urdu (0.2M): 17.9 / 5.2
With 0.2M to 1.8M tokens NMT loses by 6.9 to 12.7 BLEU; the smallest corpus has the worst gap.
What do Koehn and Knowles's (2017) learning curves show about NMT against phrase-based MT?
- Corpus size from about to words, doubling each step.
- Neural starts catastrophically low: 1.6 BLEU at about 0.4M words, against 16.4 phrase-based and 21.8 phrase-based with a big LM.
- Neural improves faster: overtakes phrase-based between about and words (22.4 < 23.5, then 25.7 > 24.7), and phrase-based with big LM around to words. At the largest size: 31.1 against 28.6 and 30.4.
- Under tens of millions of words statistical MT wins, which is where most language pairs are.
What is transfer learning for NMT, and what four conditions make it work best?
- Learning tens of millions of parameters from small data is hard, so initialise with “reasonable” parameters. Basically fine-tuning: train on a large-data pair (parent), continue training on the low-resource pair (child).
- Works best if: (1) the target language of parent and child is identical; (2) the source languages are related; (3) decoder parameters are frozen (overstated: only the target embeddings should be); (4) child and parent share the same vocabulary.
Give the steps of parent-child transfer learning for low-resource NMT, with Zoph et al.'s concrete setup.
- Train the parent from random initialisation on the large corpus (French→English, 300M English tokens, WMT’15, 5 epochs, about 26 BLEU dev).
- Copy every parameter into the child.
- Map each child source word (Uzbek) to a row of the parent’s source embedding matrix; that row is its initial embedding.
- Freeze the target-language (English) embeddings.
- Continue training on the child corpus (Uzbek→English, 1.8M English tokens) with strong regularisation (dropout 0.5), updating only non-frozen parameters.
In parent-child NMT transfer, what is shared, what is not, and why are the target-language parameters frozen?
- Not shared: the source vocabulary and source embeddings. Child rows are initialised from parent rows, but nothing about that French word means anything for the Uzbek word.
- Shared: the target language and all other parameters (target embeddings, attention, …).
- Frozen target parameters: English is the same in parent and child, and the parent saw 300M English tokens against 1.8M, so its English embeddings are far better than the child data could make them.
Give Zoph et al.'s (2016) transfer results per child language, and what they show.
NMT / Xfer / Final / SBMT:
- Hausa 16.8 / 21.3 / 24.0 / 23.7
- Turkish 11.4 / 17.0 / 18.7 / 20.4
- Uzbek 10.7 / 14.4 / 16.8 / 17.9
- Urdu 5.2 / 13.8 / 14.5 / 17.9
Transfer alone: +4.5, +5.6, +3.7, +8.6 (Urdu nearly triples). With further improvements (Final, whose exact definition is not given in the source), NMT overtakes SBMT for Hausa and is 1.1 to 3.4 BLEU short for the others.
Give the Uzbek→English freezing experiment results and the rule they support.
Dev BLEU as parameter groups are unfrozen one by one: none 0.0; + source embeddings 7.7; + source RNN 11.8; + target RNN 14.2; + attention 15.0; + target input embeddings 14.7; + target output embeddings 13.7.
- With nothing trained the output is garbage: source embeddings must be retrained (biggest jump).
- Training helps up to and including attention.
- Unfreezing target embeddings hurts (learned from 300M tokens, overfit on 1.8M).
- Rule: freeze the target embeddings, train everything else. The general claim “decoder parameters are frozen” is contradicted by the target RNN’s +2.4.
What do Zoph et al.'s relatedness experiments, including French', show about what NMT transfer transfers?
- Spanish→English child: no parent 16.4, German parent 29.8, French parent 31.0. A related parent helps more.
- French’ = French with random vocabulary reshuffling (consistent but arbitrary word forms; same grammar and word order, no shared surface forms). French parent lifts it from 13.3 to 20.0: structure (word order, syntax, attention shape) transfers, not only shared words.
- Uzbek with a French parent (unrelated): 15.0, the smallest benefit.
What does Kocmi and Bojar (2018) find about NMT transfer with a shared parent-child vocabulary?
- Notation “enFI - enET”: parent English→Finnish, child English→Estonian.
- Every transfer from a larger parent helps, by 1.6 to 3.4 BLEU.
- Unrelated parents help as much as related ones: English→Estonian 19.74 (Finnish parent), 20.41 (Czech), 20.09 (Russian), against 17.03 child-only.
- “Only parent” (parent applied directly) is near zero except Czech/Slovak (up to 11.62), which are mutually intelligible.
- Reversed direction (smaller parent, larger child) helps little (enET - enFI +0.57) or hurts (ETen - FIen 23.95 against 24.40; SKen - CSen 28.20 against 29.61). Transfer must flow from the larger corpus to the smaller.
What do Aji et al. (2020) find about which parts of a parent NMT model carry the transfer benefit?
Children My, Id, Tr → English; embeddings and inner layers each transferred (Y) or randomly initialised (N). Reported averages:
- Y/Y 21.7 (best); N/Y (inner only) 18.3; Y/N (embeddings only) 13.7; N/N (scratch) 14.5.
- Inner layers carry most of the benefit; embeddings alone are worse than nothing (no inner layers to interpret them); embeddings on top of inner layers add more, so the parts work together.
- Burmese shows it most: 4.0 from scratch, 17.8 with full transfer.
What does Aji et al.'s comparison of a German→English and an English→German parent suggest about the "same target language" condition?
With full transfer into X→English children, the En→De parent (English on the source side) transfers about as well as the De→En parent (English on the target side): mean 21.7 against 21.8. This sits awkwardly with the rule that parent and child should share the target language.
Contrast BERT, GPT and BART on architecture and pretraining task.
- BERT: encoder only, bidirectional; masked tokens (A _ C _ E → B, D) predicted independently.
- GPT: decoder only, autoregressive; position sees only earlier positions and predicts the next token.
- BART: bidirectional encoder reads a corrupted document, autoregressive decoder reconstructs the original, attending to the encoder through cross-attention. A denoising autoencoder; covers tasks needing both (translation, summarisation).
Give BART's five noise functions with what each does to
A B C . D E .
- Token masking:
A _ C . _ E .(B, D replaced by[MASK]; predict them).- Token deletion:
A . C . E .(B, D removed; predict them and where they were).- Text infilling:
A _ . D _ E .(span B C replaced by a single[MASK]; a[MASK]also inserted between D and E where nothing was removed; predict the masked sequence).- Sentence permutation:
D E . A B C .(restore original sentence order).- Document rotation:
C . D E . A B(identify where the document really starts).What does each of token masking, token deletion and text infilling hide from the BART encoder?
- Masking marks the gap, so the model only has to fill it.
- Deletion hides where the gap is, since nothing marks it.
- Infilling hides how long the gap is, since one
[MASK]can stand for any number of tokens (including zero).Define document rotation in BART, and why is "rotation" the better name than "document permutation"?
A random position of the document is used as the first token, followed by the rest from that position, followed by the actual beginning up to the random position. The task is to predict the actual first position. “Rotation” is accurate because order is preserved cyclically; nothing is shuffled.
Summarise BART's comparison of pretraining objectives (SQuAD, MNLI, perplexity on generation tasks).
- Text infilling is the best single noise: SQuAD 90.8 F1, best XSum (6.61) and ConvAI2 (11.05) perplexity.
- Deletion beats masking on all four generation tasks (it is harder: locate the gaps).
- Document rotation and sentence shuffling alone are poor: SQuAD 77.2 and 85.4, ELI5 PPL 53.69 and 41.87; too weak a token-level signal.
- Infilling + shuffling: SQuAD 90.8, best CNN/DM perplexity 5.41.
- Plain left-to-right LM: best ELI5 (21.40), worst SQuAD (76.7), since it lacks bidirectional context.
- BERT Base keeps the best MNLI (84.3).
How is BART fine-tuned for classification and for span prediction such as SQuAD?
- Classification: the same uncorrupted input goes into encoder and decoder; the label is predicted from the last hidden time step of the decoder, which has attended to the whole input.
- Span prediction: each token is labeled, and the model learns to predict the beginning and end labels of the answer span.
How is BART fine-tuned for machine translation, and what were the Romanian→English results?
- A randomly initialised source-language encoder replaces BART’s embedding layer. It maps foreign tokens into vectors BART can treat like (noisy) English embeddings, which BART then denoises into English.
- BART’s parameters can stay frozen or be updated; the source vocabulary can differ from BART’s.
- Ro→En BLEU: baseline 36.80, Fixed BART 36.29 (worse), Tuned BART 37.96 (+1.16). Pretraining helps only when BART’s parameters may adapt.
- Same idea as parent-child transfer: keep a large English-trained model, train a small new component to connect a new source language.
How does large BART compare with BERT, XLNet and RoBERTa on SQuAD, and on CNN/DM and XSum summarisation?
- SQuAD 1.1 EM/F1: BERT 84.1/90.9, XLNet 89.0/94.5, RoBERTa 88.9/94.6, BART 88.8/94.6. SQuAD 2.0: RoBERTa 86.5/89.4, BART 86.1/89.2. Adding a decoder costs nothing on understanding tasks.
- Summarisation ROUGE-1/2/L: BART best on every column (CNN/DM 44.16/21.28/40.90; XSum 45.14/22.27/37.25). Largest lead on XSum: +3.69 R1 and +3.48 R2 over RoBERTaShare.
- Lead-3 is strong on CNN/DM (40.42 R1) and useless on XSum (16.30 R1, 1.60 R2).
Link to originalName the benchmarks and metrics used to evaluate XLM, InfoXLM, the within/across tests, NMT transfer and BART, and what each measures.
- XNLI: crosslingual NLI, premise/hypothesis three-way classification in 15 languages; accuracy.
- XSQuAD: extractive QA with context and question possibly in different languages, 11 languages; F1.
- Within/across scores: same-language against mixed-language input, per language.
- Perplexity: language modeling quality (Nepali; BART’s ELI5, XSum, ConvAI2, CNN/DM); lower is better.
- Cosine similarity, L2 distance, SemEval’17: alignment of crosslingual word embeddings.
- BLEU: MT quality. ROUGE-1/2/L: summarisation. SQuAD EM/F1, MNLI accuracy: understanding tasks.