Skip to content

OCR0065: Create repo on Hugging face for all the datasets we have of OCR #4

Description

@ta4tsering

Description:
So currently we have a lot of OCR data that we have annotated and all of those images are on s3 with each image as a single object and also the cvs files are all in s3. So to make the datasets easily usable and easily accessible I will creating the zip files with all the images and upload to the hugging face repo with the transcriptions and all with the data split or data distribution that eric used.

Completion Criteria:

All the Tibetan OCR data uploaded to Openpecha hugging face.

Subtasks:

  • Lhasa Kanjur
  • Lithang Kanjur
  • Derge Tenjur
  • Norbuketaka
  • Google Books
  • Betsug data
  • Durtsa data
  • update the script for the special case of google books and norbuketaka data

note:
for the Norbuketaka and Google books, we already have a hugging face repo but without the data distributions so I am using that hugging face repo to create the new hugging face repo on Openpecha hugging face with the data distribution but without the zipped image file

Card Reviewer:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

  • Status
    Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions