Skip to content

OCR0003: Benchmark data for Woodblock printed OCR images #9

Description

@tenzinyonten

Description:

Upload the Derge dataset images to AWS S3, generate metadata CSVs, and upload the CSVs to a Hugging Face repository.

Implementation:

  1. Upload Images: Upload images to the S3 bucket and generate public URLs.
  2. Create CSVs: Generate train, test, and eval CSVs with metadata.
  3. Upload to Hugging Face: Upload the CSVs to the Hugging Face repository.

Subtasks:

  • Upload images to S3.
  • Create metadata CSVs (with filenames, URLs and transcripts).
  • Convert CSV files to Parquet format and upload them to the Hugging Face repository.

Completion Criteria:

  • Images uploaded to AWS S3 with URLs.
  • Metadata CSVs created and uploaded to Hugging Face.

Card Reviewer:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

  • Status
    Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions