Skip to content

Phase 6: implement offline benchmark harness (run_benchmark) #12

Description

@monthop-gmail

run_benchmark() ยัง raise NotImplementedError — metrics กับ ground truth พร้อมแล้ว เหลือ harness

งาน

  • mode offline: ประกอบ service ด้วย DeterministicEmbedder + IdentityReranker แล้ว ingest FIXTURE_CORPUS
  • รันทุก case ผ่านทั้ง 3 arm ด้วย service.with_strategy()
  • aggregate เป็น ArmResult ด้วย evaluation.metrics
  • fail ทันทีถ้ามี leakage ก่อนจะรายงานตัวเลขคุณภาพใด ๆ (§26)
  • print รายงานที่อ่านรู้เรื่องลง stdout และคืน exit code ตาม report.passed

เงื่อนไขว่าเสร็จ

  • make benchmark รันได้โดยไม่ต้องมี API key และไม่ต้องมี DB (หรือใช้ DB local ก็ได้ — ระบุให้ชัด)
  • รันสองครั้งได้ผลเท่ากันเป๊ะ (deterministic)

เตือนไว้ในรายงานด้วย

ตัวเลข offline เทียบได้กับ offline ด้วยกันเท่านั้น — DeterministicEmbedder เป็น hash ไม่มี semantic
Hit@K ของมันบอกคุณภาพ retrieval จริงไม่ได้ ใช้จับ regression ของอัลกอริทึมเท่านั้น

อ้างอิง: §15, §18, §25, §26 · evaluation/benchmark.py

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    evaluationGround truth, metrics, benchmark (§15-§18)

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions