Skip to content

feat: move assignment reference-file extraction to the mark-jobs worker - #604

Open
NoahFreelove wants to merge 9 commits into
masterfrom
feat/assignment-file-extract-worker-offload
Open

NoahFreelove wants to merge 9 commits into
masterfrom
feat/assignment-file-extract-worker-offload

Conversation

@NoahFreelove

Copy link
Copy Markdown
Contributor

Summary

Author reference-file content extraction (pdf-parse / pdfjs+canvas — the CPU/RAM-heavy part of completeAssignmentFileUpload) no longer runs on the mark-api request thread. The complete endpoint now finalizes S3, marks the row status=READY / extractionStatus=PENDING, and enqueues a job on a new mark.file-extract BullMQ queue; the extraction half runs on the mark-jobs fleet via AssignmentFileService.runAssignmentFileExtraction and persists extractedText + terminal READY/FAILED.

  • New queue + job constants (mirrored api/jobs), FileExtractJobPayload, QUEUE_METADATA entry
  • Producer split in assignment-file.service.ts; enqueue failure is tolerated (row stays PENDING, upload still succeeds)
  • JobExecutorService dispatches the queue; jobs app registers it through the declarative workerSpecs table: FILE_EXTRACT_CONCURRENCY (default 2), lockDuration 300s, maxStalledCount=0 (poison-file OOM fails permanently instead of cascading to more pods)
  • Chart defaults: FILE_PROCESSING_BUDGET_BYTES on mark-jobs, commented FILE_EXTRACT_CONCURRENCY, heavy-tier example updated

This is backend-only: the v2 AssignmentFile pipeline has no frontend caller yet, and its sole server-side consumer (question generation) already reads only extractionStatus=READY rows, so the sync→async contract change breaks no consumer.

Behavior details

  • S3 download errors rethrow so the job's retry budget (attempts=3) covers transient infra blips; extraction errors are deterministic and persist FAILED terminally. The method is idempotent (fresh row load by id; single overwrite update).
  • The worker trusts nothing from the payload beyond ids — bucket/key/mime are re-derived from the row.
  • Structured logging on every path, including the raw-update fallback.

Deploy follow-up (apps-faculty-deploy, separate PR)

  • config/mark-jobs-heavy/values.yaml.gotmpl: JOB_WORKER_QUEUES: "mark.attempt.heavy,mark.file-extract", plus FILE_EXTRACT_CONCURRENCY: "2" and FILE_PROCESSING_BUDGET_BYTES: "1073741824"
  • Fast tier (config/mark-jobs): unchanged — its explicit queue list must NOT gain mark.file-extract
  • Ordering: deploy this image before (or with) the heavy-tier queue-list change — a pod listing an unknown queue crash-loops at boot by design
  • Until the overlay ships, jobs enqueue and wait in Redis (no TTL) and drain when the heavy tier picks the queue up; nothing user-visible in the window
  • S3 egress for mark-jobs already exists (grading downloads submissions today); no network-policy change

Notes for the deferred PENDING-row sweeper

  • A permanently failed job sits in the failed set (removeOnFail: count 100) under jobId extract:<fileId>; a later add with the same jobId is silently deduplicated. The sweeper must remove the stale failed job first (or use a nonce jobId) or re-enqueues will silently no-op.
  • A stall-killed (OOM) job permanent-fails without transitioning the row, leaving it PENDING — indistinguishable from never-enqueued. Fine while headless; add a failed-listener DB mark when the multipart UI lands.

Testing

  • apps/api: src/job-queue + assignment v2 file suites — 144 passed (2 pre-existing environment skips)
  • apps/jobs: worker + constants-sync + process-safety — 93 passed
  • Chart verified with helm template (budget key renders; JOB_WORKER_QUEUES stays unset by default)

Registers mark.file-extract in the workerSpecs table with a dedicated
handleFileExtractJob handler (concurrency 2, 5-min lock,
maxStalledCount=0). Updates all spec assertions that enumerate exact
worker counts (6→7) and adds local-routing + forward-routing coverage
for the new queue.
…dget

Extraction errors are deterministic and persist FAILED, but an S3
download error is infrastructure: rethrow it so the remaining BullMQ
attempts recover instead of terminally failing the file. Also surface
and rethrow a failure of the raw fallback update.
@skills-network-bot

Copy link
Copy Markdown

This pull request has had no activity for 30 days, so it has been marked stale. It will be closed in 14 days unless it becomes active. To keep it open, push a commit, leave a comment, or remove the stale label. If it is blocked, a short note on what it is waiting for is enough.

@skills-network-bot skills-network-bot Bot added the stale No activity for 30 days. Closes after 14 more unless something changes. label Oct 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stale No activity for 30 days. Closes after 14 more unless something changes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant