Skip to content

Explore better models & architecture #420

Description

@tenzinyonten

Description

TiSpell (RoBERTa + copy-gate) has architectural limitations when handling multi-syllable edits and complex Tibetan glyph stack corruptions. Explore, fine-tune, and benchmark alternative sequence-to-sequence and LLM architectures against TiSpell on an audited CSC evaluation set to determine the most effective production pipeline.

Subtask

  • Document candidate architectures, focusing on byte-level seq2seq (google/byt5-base), subword seq2seq (mT5), character-level models, and prompted/fine-tuned LLMs (Gemini).
  • ByT5 Fine-Tuning: Fine-tune google/byt5-base on v11_train_ready.csv.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions