View as table
Registers
Three in five words in the corpus come from oral literature — uwepeker, kamuy yukar and other narratives recorded from the last generations of native speakers. The modern registers are small beside that; social media and the Ainu Times are where new Ainu prose appears.
Share of word tokens. Click a segment to pin a register across the whole page.
Distinct word-form counts grow with a register's size; that view is a size signal more than a richness signal.
View as table
Dialects
Two Hokkaidō varieties — Saru and Shizunai — supply more than half of the corpus. The Sakhalin record is thin but present, from Asai Take's folktales to Piłsudski-era letters.
Sentences per dialect, as attributed in the source metadata; sentences with no dialect record are grouped at the end.
The lexicon
View as table (top 50)
Clitics keep their =
(a=, =an); spelling variants are folded for
accent marks and apostrophe-written glottal stops, and each form is shown in its commonest
spelling. The curve stops at rank 1,000.
Fingerprints
Sentence rhythm is the first measurable difference between the registers; recurring phrases are the second.
Each panel: share of the register's sentences by length in word tokens, bins of two; the gray silhouette behind every panel is the whole corpus, so shapes compare directly. The tick marks the register's mean. Length also reflects how each source was segmented — Bible verses, breath-group transcriptions and dictionary examples cut sentences differently. Click a panel to pin.
Signature phrases
The register's most frequent word sequences (raw counts over consecutive tokens, accent-insensitive) — comparable within a register only. For the words that distinguish it from the rest of the corpus, see signature words.
Signature words
Keyness scores every word by how much more often one register uses it than the rest of the corpus does. High scorers are the register's own vocabulary: what it keeps reaching for.
View as table
Grammar mix
Percent of the register's POS-tagged tokens (Universal Dependencies tags, machine-tagged). Click a column header to sort the registers by that tag.
Collections
Every source in the corpus, one row each, with the counts behind every chart above.
Words counts Latin-script tokens; mean length counts Latin and kana word tokens per sentence, as in the fingerprints above. Dialect shows the collection's most frequent attribution. Click a register chip to pin.
Methods & data
- Every sentence inherits the register of its source collection (collection→register table).
- Slices: Hokkaidō and Sakhalin cover sentences whose source metadata records a region; traditional and modern follow a per-collection era tag in the same table. The Bible translation and トピック別 アイヌ語会話辞典 sit outside the era split. Within a slice, register keywords score against the rest of that slice.
- Word-level statistics count Latin-script tokens; forms are compared with accent marks and apostrophe-written glottal stops folded, and displayed in their commonest spelling.
- Keyness: log-likelihood G² (Dunning 1993), each register against the rest of the corpus, with rates per million tokens. Words need 20 corpus occurrences, 5 in the register, and G² ≥ 10.8 (p < 0.001) to qualify; single-letter and punctuation-bearing tokens are excluded, since speaker labels and list markers dominate them.
- News, social media and the Bible translation are single-collection registers, so their keywords are also those collections' keywords.
- Function words are flagged from the tagger: closed-class UD tags and clitics.
- N-grams follow the token layer: consecutive tokens inside one sentence, clitics counted as tokens, accent-insensitive.
- Sentence length counts word tokens (Latin and kana script); histogram bins of two, overflow at 31+. Sentences with no word token stay outside the histograms.
- POS tags: Universal Dependencies, machine-tagged; of tokens carry a tag.
- Zipf series truncated at rank 1,000.
- Snapshot · data: stats.json.