lang-smn
Last build: 4 days ago
Total Builds1
Last build 4 days ago
Success Rate100%
1 passed, 0 failed
Visibility public
Repository access
Recent Builds
Speller: keep a->o and á->i at the default weight
giella-core 1.16.12 fixed editdist.py so a pair listed once also sets its
reverse. The reverses a for o and á for i lose 2 top-1 rows and 1 top-5 row
(kannâd-uv -> kannat-uv loses to konnâd-uv); listed at 10 they cost what
they did. typos-default-generated.tsv (1,002 rows): 854 / 947 / 952, as
before the fix.
20m 11s
time-days-ago
15m 4s
time-days-ago
28s
time-days-ago
12m 40s
time-days-ago
17m 5s
time-days-ago
21m 33s
time-days-ago
21m 23s
time-days-ago
17m 36s
time-days-ago
13m 12s
time-days-ago
18m 42s
time-days-ago
9m 27s
time-days-ago
14m 55s
time-days-ago
Enable boundary edits at weight 20
The only runtime knob in a full config sweep that moves reachability rather
than reordering: it recovers missing-space and missing-hyphen rows
(ovdilmainâšum -> ovdil mainâšum, suomâugrâlii -> suomâ-ugrâlii) that no
letter edit reaches. Plateau 10-30, zero row regressions on either test set,
+3 combined and +8 generated top-1, about +5% average lookup latency.
word-split-weight reaches the same rows with worse ranks and adds nothing on
top, so boundary-edit alone. Also swept and left alone: beam is inert from
20 to unbounded (top-1/top-5 identical at every value), and the 3/2/3
reweight stands - the nominally better 5/0/3 costs the generated set 22
top-1, a classic acceptance-set overfit. This file ships inside the archive.
12m 2s
time-days-ago
Set maxweight to 20 for the canonical speller corpus
Fitted against the canonical corpus - corpus-smn plus corpus-smn-x-closed,
4,221,192 tokens assembled at build time - which requires giella-core 1.15.2
or later; before that, a plain build silently weighted from the in-tree file
and this value would be measured against the wrong corpus. Note smn is the
one language whose PUBLIC corpus alone is smaller than the in-tree file, so a
build without the closed corpus inverts every comparison here.
Swept 5..40 with unit refinement across 16..22. Versus the previous value of
10, on the same canonical corpus: typos.tsv +12 top-1 / +4 top-5, generated
unchanged top-1 / +2 top-5. 20 is the centre of an 18..22 plateau, the peak
of the even blend, and effectively tied for the peak under attested-weighted
blending; it is also best top-5 on all three test sets at once.
Separately from the tuning: the canonical corpus itself cost the generated
set 53 top-1 hits relative to the old in-tree corpus, and no maxweight value
recovers them - a corpus effect, not a weighting effect. Ranking is close to
exhausted on both sets (top-5 within a point of its reachability ceiling);
the remaining headroom is reachability.
12m 36s
time-days-ago
Add more real typos, extracted from the corpus files based on potential typos (Err/Orth entries) - these typos are actual ones found in corpus material
15m 9s
time-days-ago
10m 30s
time-days-ago
17m 59s
time-days-ago
16m 21s
time-days-ago
12m 32s
time-days-ago
12m 39s
time-days-ago