Skip to content

llm_hub: tests with a fake LiteLLM proxy started inside the job - #4

Merged
arash77 merged 3 commits into
llm-hub-context-window-checkfrom
llm-hub-ci-test2
Oct 2, 2026
Merged

arash77 merged 3 commits into
llm-hub-context-window-checkfrom
llm-hub-ci-test2

Conversation

@arash77

@arash77 arash77 commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Proposal for bgruening#1943: tests that run the tool for real, with no CI change.

  • test-data/fake_litellm.py: a small fake LiteLLM proxy in front of vLLM (Python standard library only). It uses the real formats: /model/info, /utils/token_counter, streaming /chat/completions, API key check, vLLM's HTTP 400 for prompts over the context length and for images sent to text-only models.
  • The test models (provider galaxy-test) exist only in the test data table. Only for them, the tool command starts the fake proxy inside the job and points the tool to it. Real jobs are not affected.
  • 5 new tests: input fits; input over the published limit; input over the backend limit with no published limit (fallback message); image sent to a text-only model; retry after a server error.

…ide the job

Test models (provider galaxy-test) exist only in the test data table. For
them the command starts test-data/fake_litellm.py (standard library only)
inside the job and points the tool to it, so the tests run the tool for
real without any CI change: a normal run, an input over the limit, and a
model without a limit.
@arash77 arash77 changed the title TEST: llm_hub tests with an in-job fake LiteLLM, no CI change (do not merge) llm_hub: tests with a fake LiteLLM proxy started inside the job Oct 2, 2026
The fake now checks the API key and model, counts the content with a
tokenizer (the chat endpoint adds the chat template like vLLM), returns
vLLM's HTTP 400 errors for prompts over the context length and for images
sent to text-only models, and fails once with 503 for one model. Five
tests cover: input fits, input over the published limit, input over the
backend limit with no published limit, image to a text-only model, and a
retry after a server error.
Comment thread tools/llm_hub/llm_hub.xml Outdated
#set prompt = ''
#end if

#if str($model.fields.provider) == 'galaxy-test'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a big fan of having test only behavior in the command section.
Maybe we could at least move it into a macro (in an extra file)?
So it is clearly indicated that this only for testing purposes.
@bgruening what is your opinion?

This could be good compromise because otherwise we would need to change the CI behavior, letting it start a mock server, right ? @arash77

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, moved it to test_macros.xml (cfb2784). The command now only has @START_FAKE_LITELLM@.

And yes, without this the CI would have to start the mock server itself. That means changing the shared pr.yaml for one tool, so it would run for every tool.

@IvoLeist IvoLeist Oct 2, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe if at one point we realise that multiple tools would benefit of a mock server test we could think about a way to integrate this in pr.yaml but for now I would not touch it.

Keeps the test-only part out of the command section.

Co-authored-by: Ivo <Ivo.leist@googlemail.com>
@arash77
arash77 merged commit cfb2784 into llm-hub-context-window-check Oct 2, 2026
10 checks passed
@arash77
arash77 deleted the llm-hub-ci-test2 branch October 5, 2026 09:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants