Skip to content

OCR00020: Get the data split Eric used. #1

Description

@ta4tsering

Descriptions:
To compare between the different OCR architecture, we need to have a single data split of training, eval and test sets. if we have that we can easily compare between models we trained with a single test split. To do so we need to get all the distribution of the OCR training data that Eric used to train the models that he trained.

Completion Criteria:
Get the splits for

  • Lhasa Kanjur
  • Derge Tenjur
  • Lithang Kanjur
  • Google Books
  • Norbuketaka
  • Betsug
  • Durtsa

Subtasks:

  • Download the models trained from hugging face
  • parse the data distribution if available on the hugging face.
  • create data split json for OCR

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

  • Status
    Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions