Skip to content

Train TiSpell on the synthetic dataset #412

Description

@tenzinyonten

Description

Train TiSpell on a stratified mix of synthetic and real annotator pairs. All three sources (BoCorpus synthetic, news synthetic, real annotator pairs) appear in train, val, and test, with scores reported per source so the real-pair numbers stay legible against the synthetic ones.

Config

openpecha/tibetan_RoBERTa_S_e3, lr 2e-5, batch 16, patience 3, max_char_length=380, w_c=2.0, dropout 0.2.

Subtasks

  • Stratify train/val/test by source and error type, split by source sentence
  • Train the mixed run
  • Compare per error category and per source against the no-op baseline
  • Evaluate on held-out synthetic errors as a diagnostic
  • Document the final mix, push model and dataset

Reviewer

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

Status
In Progress

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions