Hall Kwak'wala Gloss Reviewer — Luke 4 & 8

About this tool

Word-by-word review of Rev. A. J. Hall's Kwak'wala translation of Luke (1894; digitised MissionAssist 2017). Draft glosses come from a hand-curated lexicon (built from Boas 1893 Vocabulary, Boas 1900 Sketch, Boas & Yampolsky 1947 Grammar with a Glossary of the Suffixes, and Hall's own 1888 Grammar) supported by a statistical alignment (IBM Model 1, 8 EM iterations) of Hall's full Matthew–Acts corpus (4,826 verses) against the KJV. Hall breaks phonological words at enclitic boundaries — e.g. lāh dāḵw = la=x̌da'x̌w "then-they", ḵī da = object marker + determiner — so many of his "words" are clitics of a single Kwak'wala word. Kwak'wala is predicate-initial: clauses typically open with the predicate or the auxiliary la.

Hall original vs. Hall corrected — merging

The working baseline of each tile is the Hall corrected form; the small brown line above shows Hall's original token(s) whenever the corrected form differs. Where Hall split one word into several tokens, select the first piece and use merge with next ⇒ (repeat to absorb more pieces); the pieces join into one unit with a tan border, the corrected spelling is editable, and glossing then applies to the whole unit. Unmerge restores Hall's tokenization. Merging resets that unit's review status, since its gloss needs rethinking. For a merged unit, the bulk button finds every occurrence of the same Hall token sequence in the chapter and applies the same merge + annotations. Each merge is also a retokenization rule that travels with the JSON export: fed back to Claude, the same sequences are merged across the whole corpus and the statistical alignment is re-run against the corrected tokenization, improving the draft interlinear for all remaining passages.

The lines on each word

Hall original (small, shown when it differs or via the header toggle) · Hall corrected (bold; brown when changed) · U'mista rendering · Liq'wala rendering · probable phonemic form (Boas-style, informal) · gloss. The U'mista and Liq'wala lines are derived mechanically from the reconstructed phonemic form (x̌→x̱, q→ḵ, ə→a̱, ƛ→tł for U'mista; NAPA conventions with comma-above ejectives, superscript ʷ and ʔ for Liq'wala) and inherit its uncertainty — working aids, not attested spellings. The Liq'wala line is an orthographic transliteration of Hall's Kwak'wala, not a conversion to the Liq'wala dialect's own forms. Your edits appear with a dotted underline.

Two gloss tiers

Each unit carries a morphological gloss (clitic-by-clitic, e.g. AUX=OBJ-PL=that(them)) and a word-level gloss (plain English, e.g. to them; dark slate line, toggleable in the header). Word-level defaults are derived from the morphological gloss and are rough — refining them is part of the review. On merge, both tiers get an automatic combined suggestion built from the members, fully editable before confirming.

Suggestions, bad glosses, and propagation

Below the gloss fields, suggestion chips collect the original proposal plus glosses previously confirmed for the same form. Right-click a chip to remove it as a bad gloss — it moves to a struck-through list (click there to restore). A banned gloss stops being proposed everywhere for that form; unreviewed occurrences that were showing it switch to the most recent human-confirmed gloss instead. More generally, confirming any gloss immediately propagates it as the working proposal to all unreviewed occurrences of the same form (or sequence) — shown in italic green to mark "inherited from a confirmation elsewhere, not yet reviewed here" — and into the concordance mini-interlinears. Multiple senses are preserved: per-occurrence review can always override the propagated gloss, and every distinct confirmed sense remains available as a chip. Confirming a merged unit also auto-merges matching sequences in the rest of the chapter (as unreviewed proposals; the bulk button is the way to mark them all reviewed at once).

Corpus-wide rules

Beyond the chapter, a confirmed unit can be applied to the whole Matthew–Acts corpus as a rule (sequence → corrected form + glosses). When the sequence matches ≤ 100 verses, a one-click "apply corpus-wide" button is offered next to Confirm. For higher-frequency matches — where blanket application is risky — use select occurrences… in the concordance header: every entry gets a checkbox (all checked by default), you untick the occurrences that should not receive the gloss while paging through with "show 60 more", and apply to the checked set; repeated batches extend the same rule. Verses covered by a rule show a green ✔ and render the sequence as a single merged unit with its gloss in the concordance. An active rule is summarized next to the Confirm buttons, where it can be deleted. Rules travel in the JSON export and drive the corpus-wide retokenization + re-interlinearization when fed back to Claude.

Color coding — main text

blue glosscurated, reasonably secure
amber gloss ?curated but only plausible — please scrutinize
red gloss ??curated but uncertain; also newly merged units awaiting a gloss
‹angle brackets›no curated analysis at all — raw statistical alignment suggestion only; treat with caution
green glossreviewed by a human (confirmed or edited); the tile gets a green background
italic greeninherited proposal — a human confirmed this gloss for the same form elsewhere; this occurrence is still unreviewed
tan bordera merged unit (Hall corrected ≠ Hall's tokenization)

Color coding — concordance

When "glosses" is ticked, each concordance verse is a mini-interlinear. Its small glosses are wordform-level defaults, not per-occurrence analyses — a homograph shows the same gloss in every verse even where another sense applies. Colors: blue/amber/red as above for curated entries; grey italic ‹brackets› = the form has no curated analysis, so the label is only the statistical aligner's best guess (the least trustworthy line in the tool); green = a human has confirmed a gloss for this wordform somewhere (most recent confirmation shown). Words with no gloss line at all are too rare for the aligner (typically frequency 1). For a merged unit the concordance searches the whole token sequence, highlighting it wherever it occurs in the corpus.

Reviewing

Click a word; correct the spelling if needed; edit or accept the gloss and the U'mista/Liq'wala renderings; Confirm ✓ records that single occurrence (senses are per-occurrence by design). The bulk button also fills all unreviewed occurrences of the same form (or, for merged units, the same sequence) in the chapter — any of them can still be individually re-glossed afterwards. Keyboard: move between words, change verse, Enter confirm, M merge with next, Esc deselect. Notes behave differently from the other fields: the occurrence note and the verse note save by themselves the moment you leave the field, so you can annotate a word without reviewing it, and add a note to an already-reviewed word without re-confirming. Gloss, corrected form and orthography edits are decisions, so they need Confirm. Anything typed and not yet confirmed is held while you look around the concordance, so a panel refresh never eats your text. Work autosaves in this browser (per device); use Export JSON regularly — it is the durable record, and the file to send back to Claude so the remaining chapters improve.

Gloss abbreviations

AUXauxiliary la- 'then/go' that carries most narrative clauses
DETdeterminer (Hall writes =ida 'the (subject)' split as ī da)
OBJobject marker x̌a/x̌is (Hall ḵā, ḵī)
GEN/INSTgenitive/instrumental =sa 'of/with the'
FOCfocus/relativizer yix̌- 'namely, the one who'
DISCdiscourse enclitic =ʔm(is) 'and so, just, indeed'
FUTfuture =tł(a) (Hall kla, klī, -kl)
PASSpassive =so' (Hall sū, -sau, -sawā)
PLthird-plural =da'x̌w (Hall dāḵw…)
INCHinchoative -x̌'id 'begin/become' (Hall -īd, -īda)
PURPpurposive qas 'so that' (Hall kās)
NEGnegative k'is; k'i'os 'there is none'
3.DISTthird-person distal deictic =i (Hall ī)
PROHprohibitive qəla 'don't'
SUBORDsubordinating (yix̌s 'when/as', qax̌s 'because')

Free translation lines

lit = literal back-translation following the Kwak'wala (predicate-initial) order; KJV and ASV (a close proxy for the RV 1885) are Hall's likely source texts. Tan boxes are analytical notes, including flagged difficulties.

Click a word to review it. move between words, Enter confirm, M merge with next, Esc deselect.