feat(workflow): support multimodal agent trajectories - #1606
Conversation
Prepare image prompts with the Hugging Face processor and preserve vision tensors and token type IDs so exported agent trajectories remain trainable. Key changes: - Support processor-backed Chat Completions and Responses prompts - Export multimodal training fields across interaction chains - Add the Geometry3K agent workflow and focused regression tests
| assert resp is not None, "Model response is not set." | ||
| if ( | ||
| self.prompt_token_ids is not None | ||
| and self.prompt_token_ids != resp.input_tokens |
There was a problem hiding this comment.
This check does not validate the prompt tokens vLLM actually used.
For multimodal requests, VLLMBackend sends messages to /v1/chat/completions, where vLLM independently renders the prompt. However, RemoteInfEngine constructs ModelResponse.input_tokens from the original local req.input_ids, not from prompt IDs returned by vLLM.
Therefore, this comparison still passes if AReaL and vLLM expand the same image differently. This PR should not present the current client-side comparison as providing the same guarantee.
The preferred long-term solution is to migrate the vLLM rollout transport to its token-in/token-out API, i.e. /inference/v1/generate so we can compare training prompt with vLLM’s final rendered prompt before sampling which applies to text-input as well.
Above design is too large for this PR and deserves a separate issue, I'd suggest now remove or qualify the claim that this check guarantees rollout/training prompt equality
There was a problem hiding this comment.
Thanks, agreed. The previous wording overstated what this comparison can guarantee for vLLM because ModelResponse.input_tokens does not contain prompt IDs independently returned by the vLLM server.
I have narrowed the scope of this PR to the SGLang rollout backend and updated the description accordingly.
Multimodal vLLM agent workflow support is deferred to a follow-up. Before enabling it, we should migrate the vLLM transport to its token-in/token-out generate API so that the request IDs are authoritative and can be validated without relying on the current chat-completions reconstruction.
| async def agenerate(self, request): | ||
| self.requests.append(request) | ||
| return ModelResponse( | ||
| input_tokens=list(request.input_ids), |
There was a problem hiding this comment.
Following up comments in types.py, this fake engine always echoes request.input_ids as ModelResponse.input_tokens, so it reproduces the client-side assumption rather than verifying tokens used by vLLM.
The current test proves only that to_tensor_dict() can compare two lists. It does not prove that resp.input_tokens came from the inference server.
I suggest adjust the test and PR description so they do not claim server-side prompt-token validation unless the response contains the actual prompt IDs used by vLLM
| rid=str(uuid.uuid4()), | ||
| metadata=metadata if not is_omitted(metadata) else {}, | ||
| tokenizer=self.tokenizer, | ||
| processor=self.processor, |
There was a problem hiding this comment.
Interrupted multimodal generation does not appear to resume from the partial output.
RemoteInfEngine handles an aborted request by appending the partial generated tokens to req.input_ids and resubmitting it. However, when vision_msg_vllm is present, VLLMBackend ignores the grown req.input_ids and resends the unchanged original messages to /v1/chat/completions.
The resulting trajectory can contain:
- a partial generation from the original prompt; followed by
- another generation restarted from that same original prompt.
We need to ensure that the actual prompt sent to vLLM on retry includes the partial generated tokens. This also needs an adapter-level abort/resubmit test that verifies the second request continues from the first request’s output.
There was a problem hiding this comment.
Agreed. The current multimodal vLLM chat-completions path resends the original messages and therefore cannot guarantee that an aborted request resumes from the grown req.input_ids.
This PR is now scoped to SGLang only, so I am not enabling or attempting to repair the multimodal vLLM path here. The vLLM follow-up should move to the token-in/token-out generate API and include an adapter-level abort/resubmit test verifying that the second request contains the partial output from the first request.
| tokenizer=self.tokenizer, | ||
| processor=self.processor, | ||
| image_data=image_data if has_images else None, | ||
| vision_msg_vllm=([vision_messages_for_vllm] if has_images else None), |
There was a problem hiding this comment.
Multimodal requests are routed through /v1/chat/completions, but AReaL's vLLM wrapper applies _wait_if_paused() only to /v1/completions.
This means multimodal agent requests can enter while generation is paused for a weight update, bypassing the lifecycle protection used by the abort/update/resume path.
We need to ensure that the endpoint used by multimodal rollout requests is guarded by the same pause mechanism before enabling this path for training.
There was a problem hiding this comment.
Agreed. The existing vLLM lifecycle protection is attached to the completions path, while the multimodal path uses /v1/chat/completions, so enabling it here would bypass the expected pause/update/resume guard.
Multimodal agent workflow support in this PR is now limited to SGLang. Before vLLM support is enabled in a follow-up, the selected vLLM endpoint must participate in the same pause/update/resume lifecycle, with
coverage for requests arriving during a weight update.
| "The tokenizer chat template must return text before VLM processing." | ||
| ) | ||
|
|
||
| images = [load_image(image) for image in image_data] |
There was a problem hiding this comment.
Security related: image_data can contain arbitrary HTTP(S) URLs, and load_image() fetches them from the proxy worker.
A client that can access the rollout proxy can therefore make workers request link-local metadata endpoints or internal cluster services. An unreachable URL can also occupy the preprocessing thread because this call supplies no explicit timeout.
The preserved URL is not transported correctly to vLLM either: the current backend later treats the URL string as base64 and wraps it in a data: URI.
I suggest reject non-inline images and accept only bounded data URIs or raw base64. If remote fetching is required, it needs explicit scheme, host and IP allowlisting, redirect revalidation, finite timeouts, and byte/pixel limits.
There was a problem hiding this comment.
I applied a narrowly scoped security fix: scheme-bearing image URLs are now rejected before image preprocessing, and the supported request contract is limited to inline data URIs or raw base64 image data.
This prevents the rollout worker from fetching arbitrary HTTP(S) resources and also removes the path where a remote URL could later be incorrectly wrapped as base64 for the inference backend.
| global _message_preprocessors, _prefix_matcher | ||
| config = _engine.config | ||
| tokenizer = load_hf_tokenizer(config.tokenizer_path) | ||
| processor, tokenizer = load_hf_processor_and_tokenizer(config.tokenizer_path) |
There was a problem hiding this comment.
If AutoProcessor.from_pretrained() fails, load_hf_processor_and_tokenizer() catches the exception and returns processor=None. This setup accepts that state.
A later image request then follows the text-only prompt-preparation path and may still be sent to the vLLM multimodal route, but its exported trajectory has no vision tensors or multimodal token-type IDs.
I suggest fail explicitly whenever image_data is non-empty but no image processor is available. Text-only deployments can continue accepting processor=None; an image request must not silently become a text-only training sample.
There was a problem hiding this comment.
Fixed. _prepare_prompt now fails immediately when image data is present but no multimodal processor is available. Text-only requests may still run with processor=None, but an image request can no longer be exported as a text-only training trajectory.
| mm_token_type_ids + [0] * resp.output_len, | ||
| dtype=torch.long, | ||
| ).unsqueeze(0) | ||
| result["multi_modal_input"] = [self.multi_modal_input or {}] |
There was a problem hiding this comment.
Performance/scalability related, this preserves the correct engine batch shape i.e. one multimodal-input dict per sequence, but the current ownership model duplicates the full vision tensors across turns.
Each interaction reprocesses and retains the conversation's complete image set. With individual export, a K-turn episode serializes K copies of the tensors. With concat, the wire payload may contain only the leaf, but the session cache still holds K copies until export.
Add a standalone Geometry3K Chat Completions agent and a dedicated SGLang example while preserving the existing VisionRLVR entry point. Reject unsafe inline image inputs and missing multimodal processors, with focused unit coverage for request and image validation.
Keep the review fix limited to rejecting remote image URLs and avoid introducing custom decoding or resource limits.
Keep backend-managed image URLs working for tokenizer-only clients while requiring a local processor for trainable v1 multimodal trajectories. Scope remote URL rejection to the local processor path so the shared client does not regress v2 inference.
| """Build model-ready prompt tokens and vision tensors with an HF processor.""" | ||
| from transformers.image_utils import load_image | ||
|
|
||
| if any("://" in image for image in image_data): |
There was a problem hiding this comment.
[NIT] The new validation blocks scheme-bearing URLs, but it does not enforce the stated inline-only contract.
transformers.load_image() also accepts local filesystem paths. A value such as /shared/private-image.png does not contain ://, so it passes this check and is opened by the proxy worker.
Suggested fix: positively decode the value as base64 instead of passing arbitrary strings to load_image()
There was a problem hiding this comment.
Change the local processor path to positively decode the input as base64 and construct the image from in-memory bytes, rejecting any value that is not valid inline image data. The tokenizer-only backend forwarding path will remain unchanged.
Prevent local multimodal preprocessing from interpreting request strings as filesystem paths. Validate base64 input and construct images from in-memory bytes before invoking the Hugging Face image helper.
|
LGTM |
Keep the SGLang tokenizer configuration advisory while preventing strict multimodal agent trajectories from entering the unsupported vLLM path. Key changes: - Warn when SGLang multimodal rollout may reprocess expanded tokens - Reject strict multimodal agent requests on the vLLM backend - Cover both OpenAI request paths and SGLang launch behavior Refs: #1606
Adapt the processor-backed proxy path from areal-project#1606 so multimodal agent interactions export the same prompt tokens and vision tensors consumed by training. Key changes: - Process inline images outside the async event loop - Validate image payloads and processor output shapes - Export and serialize multimodal training fields - Normalize mixed text and vision batches
Backport the focused single-turn agent from areal-project#1606 so the multimodal proxy path can be tested without multi-turn concat or tool feedback. Key changes: - Add the Geometry3K Chat Completions agent - Add a PPOTrainer entry point for proxy rollouts - Cover image injection, reward scoring, and request construction
Adapt the processor-backed proxy path from areal-project#1606 so multimodal agent interactions export the same prompt tokens and vision tensors consumed by training. Key changes: - Process inline images outside the async event loop - Validate image payloads and processor output shapes - Export and serialize multimodal training fields - Normalize mixed text and vision batches
Backport the focused single-turn agent from areal-project#1606 so the multimodal proxy path can be tested without multi-turn concat or tool feedback. Key changes: - Add the Geometry3K Chat Completions agent - Add a PPOTrainer entry point for proxy rollouts - Cover image injection, reward scoring, and request construction
Adapt the processor-backed proxy path from #1606 so multimodal agent interactions export the same prompt tokens and vision tensors consumed by training. Key changes: - Process inline images outside the async event loop - Validate image payloads and processor output shapes - Export and serialize multimodal training fields - Normalize mixed text and vision batches
Backport the focused single-turn agent from #1606 so the multimodal proxy path can be tested without multi-turn concat or tool feedback. Key changes: - Add the Geometry3K Chat Completions agent - Add a PPOTrainer entry point for proxy rollouts - Cover image injection, reward scoring, and request construction
Description
Enable trainable multimodal trajectories for OpenAI-compatible agent rollouts using the SGLang rollout backend.
The OpenAI client now prepares image prompts with the Hugging Face processor, carries processor-expanded prompt tokens, multimodal token type IDs, and vision tensors through interaction chains, and exports them for training. Image preprocessing runs through
asyncio.to_threadso asynchronous rollout execution remains non-blocking.This PR also adds a standalone
Geometry3KAgentcomponent with anasync run(data, **extra_kwargs)entry point. Likeareal.workflow.openai.math_agent.MathAgent, it sends one request throughAsyncOpenAI.chat.completions.create()using the proxy-provided client settings. The existingexamples/vlm/geometry3k_grpo.pyremains the VisionRLVR example and keeps its existing YAML and behavior.SGLang Geometry3K agent example
The multimodal agent workflow introduced by this PR is supported with the SGLang rollout backend. Use the dedicated example and configuration:
python examples/vlm/geometry3k_agent_grpo.py \ --config examples/vlm/geometry3k_agent_grpo.yamlMultimodal agent workflow support for vLLM is intentionally deferred to a follow-up. That work should migrate the vLLM rollout transport to its token-in/token-out API and cover prompt-token validation, interrupted-generation resubmission, and pause/update/resume lifecycle protection before vLLM is enabled for this workflow.
Related Issue
N/A
Type of Change
Checklist
pre-commit run --all-files)/review-prcommand/create-prBreaking Change Details (if applicable):
None.
Additional Context
Key changes:
RolloutWorkflow.Risk areas:
Validation:
ruff checkpassed for the modified Python files.ruff format --checkpassed for the modified Python files.git diff --checkpassed.Not run: