Skip to content

feat(llm): expose Lemonade's global_timeout so long-document prefill stops hitting a fixed ceiling #4171

Description

@kovtcharov

Lemonade v11.9.0 replaced its fixed 120-second prefill timeout with a 600-second default that is configurable via global_timeout. GAIA never had a lever for this, and long-document Q&A is exactly the workload that hits it.

Why this matters here. The canonical GAIA failure of this shape is #1030 — document Q&A on long PDFs. A large PDF turned into a single large prefill is precisely what a fixed 120s ceiling truncates, and until now the only response was to make the prompt smaller. A configurable ceiling turns "long PDF silently fails" into a policy decision GAIA can make per profile.

Worth scoping:

  • Decide where the value belongs. It likely pairs with the existing per-device context pins (GPU_CTX_SIZE / NPU_CTX_SIZE in src/gaia/llm/lemonade_client.py), since a machine running a 64K context needs a longer prefill budget than one capped at 32K — the NPU and GPU profiles should probably not share one number.
  • Surface a real error when the budget is exceeded rather than a truncated answer, per the fail-loudly rule.

Should be validated against the rag_quality eval category on a genuinely long document, not a synthetic prompt.

Proposed from the Lemonade v11.9.0/v2026.39.1 release scan; maintainer triage decides scope. cc @kovtcharov-amd

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestllmLLM backend changes

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions