Skip to content

OCR0078: Add metadata to all the ocr training dataset hf repo. #11

Description

@10kalden

Description:

We have several OCR datasets in OpenPecha Hugging Face. We need to add metadata to all these Repo

Implementation:

  • There are several datasets for line-to-text annotation for OCR; most of them only contain the image filename and its transcript.
  • We have to extract the metadata from the relevant source and add it to the dataset

Subtask:

  • Download the dataset and check for missing metadata.
  • Refer to the annotation source or the documents to extract the metadata.
  • Add the metadata and update the HF datasets

Completion Criteria:

Add meta-data to all the OCR datasets

Card-Reviewer

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

  • Status
    Todo

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions