Description:
So currently we have a lot of OCR data that we have annotated and all of those images are on s3 with each image as a single object and also the cvs files are all in s3. So to make the datasets easily usable and easily accessible I will creating the zip files with all the images and upload to the hugging face repo with the transcriptions and all with the data split or data distribution that eric used.
Completion Criteria:
All the Tibetan OCR data uploaded to Openpecha hugging face.
Subtasks:
note:
for the Norbuketaka and Google books, we already have a hugging face repo but without the data distributions so I am using that hugging face repo to create the new hugging face repo on Openpecha hugging face with the data distribution but without the zipped image file
Card Reviewer:
Description:
So currently we have a lot of OCR data that we have annotated and all of those images are on s3 with each image as a single object and also the cvs files are all in s3. So to make the datasets easily usable and easily accessible I will creating the zip files with all the images and upload to the hugging face repo with the transcriptions and all with the data split or data distribution that eric used.
Completion Criteria:
All the Tibetan OCR data uploaded to Openpecha hugging face.
Subtasks:
note:
for the Norbuketaka and Google books, we already have a hugging face repo but without the data distributions so I am using that hugging face repo to create the new hugging face repo on Openpecha hugging face with the data distribution but without the zipped image file
Card Reviewer: