Skip to content

feat(workflow): support multimodal agent trajectories - #1606

Merged
sitabulaixizawaluduo merged 8 commits into
mainfrom
feature/vlm-agent-workflow
Aug 26, 2026
Merged

sitabulaixizawaluduo merged 8 commits into
mainfrom
feature/vlm-agent-workflow

Conversation

@sitabulaixizawaluduo

@sitabulaixizawaluduo sitabulaixizawaluduo commented Aug 14, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Enable trainable multimodal trajectories for OpenAI-compatible agent rollouts using the SGLang rollout backend.

The OpenAI client now prepares image prompts with the Hugging Face processor, carries processor-expanded prompt tokens, multimodal token type IDs, and vision tensors through interaction chains, and exports them for training. Image preprocessing runs through asyncio.to_thread so asynchronous rollout execution remains non-blocking.

This PR also adds a standalone Geometry3KAgent component with an async run(data, **extra_kwargs) entry point. Like areal.workflow.openai.math_agent.MathAgent, it sends one request through AsyncOpenAI.chat.completions.create() using the proxy-provided client settings. The existing examples/vlm/geometry3k_grpo.py remains the VisionRLVR example and keeps its existing YAML and behavior.

SGLang Geometry3K agent example

The multimodal agent workflow introduced by this PR is supported with the SGLang rollout backend. Use the dedicated example and configuration:

python examples/vlm/geometry3k_agent_grpo.py \
    --config examples/vlm/geometry3k_agent_grpo.yaml

Multimodal agent workflow support for vLLM is intentionally deferred to a follow-up. That work should migrate the vLLM rollout transport to its token-in/token-out API and cover prompt-token validation, interrupted-generation resubmission, and pause/update/resume lifecycle protection before vLLM is enabled for this workflow.

Related Issue

N/A

Type of Change

  • 🐛 Bug fix
  • ✨ New feature
  • 💥 Breaking change
  • 📝 Documentation update
  • ♻️ Refactoring
  • ⚡ Performance improvement
  • ✅ Test coverage improvement

Checklist

  • I have read the Contributing Guide
  • Pre-commit hooks pass (pre-commit run --all-files)
  • Relevant tests pass; new tests added for new functionality
  • Documentation updated through the runnable example and PR usage instructions
  • Branch is up to date with main
  • Self-reviewed via /review-pr command
  • This PR was created by a coding agent via /create-pr
  • This PR is a breaking change

Breaking Change Details (if applicable):

None.

Additional Context

Key changes:

  • Support processor-backed multimodal prompts for Chat Completions and Responses on SGLang.
  • Preserve multimodal training fields across interaction chains.
  • Validate processor prompt tokens against the tokens sent to SGLang before export.
  • Add a standalone Geometry3K Chat Completions agent without depending on RolloutWorkflow.
  • Add a separate SGLang Geometry3K agent training entry point and YAML while preserving the existing VisionRLVR example.
  • Preserve Geometry3K's existing 90% accuracy and 10% format reward.
  • Serialize an empty Responses API tools list when no tools are configured.

Risk areas:

  • Multimodal token alignment across parent and child interactions.
  • Processor-specific vision tensor shapes and token type IDs.
  • Compatibility of text-only agent rollouts when a processor is available.

Validation:

  • ruff check passed for the modified Python files.
  • ruff format --check passed for the modified Python files.
  • git diff --check passed.

Not run:

  • Pytest was intentionally not run; the relevant UTs are prepared for manual execution.
  • GPU and distributed integration suites were not run because they require Linux CUDA or multi-node hardware.

Prepare image prompts with the Hugging Face processor and preserve vision tensors and token type IDs so exported agent trajectories remain trainable.

Key changes:
- Support processor-backed Chat Completions and Responses prompts
- Export multimodal training fields across interaction chains
- Add the Geometry3K agent workflow and focused regression tests
assert resp is not None, "Model response is not set."
if (
self.prompt_token_ids is not None
and self.prompt_token_ids != resp.input_tokens

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This check does not validate the prompt tokens vLLM actually used.

For multimodal requests, VLLMBackend sends messages to /v1/chat/completions, where vLLM independently renders the prompt. However, RemoteInfEngine constructs ModelResponse.input_tokens from the original local req.input_ids, not from prompt IDs returned by vLLM.

Therefore, this comparison still passes if AReaL and vLLM expand the same image differently. This PR should not present the current client-side comparison as providing the same guarantee.

The preferred long-term solution is to migrate the vLLM rollout transport to its token-in/token-out API, i.e. /inference/v1/generate so we can compare training prompt with vLLM’s final rendered prompt before sampling which applies to text-input as well.

Above design is too large for this PR and deserves a separate issue, I'd suggest now remove or qualify the claim that this check guarantees rollout/training prompt equality

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, agreed. The previous wording overstated what this comparison can guarantee for vLLM because ModelResponse.input_tokens does not contain prompt IDs independently returned by the vLLM server.
I have narrowed the scope of this PR to the SGLang rollout backend and updated the description accordingly.
Multimodal vLLM agent workflow support is deferred to a follow-up. Before enabling it, we should migrate the vLLM transport to its token-in/token-out generate API so that the request IDs are authoritative and can be validated without relying on the current chat-completions reconstruction.

async def agenerate(self, request):
self.requests.append(request)
return ModelResponse(
input_tokens=list(request.input_ids),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Following up comments in types.py, this fake engine always echoes request.input_ids as ModelResponse.input_tokens, so it reproduces the client-side assumption rather than verifying tokens used by vLLM.

The current test proves only that to_tensor_dict() can compare two lists. It does not prove that resp.input_tokens came from the inference server.

I suggest adjust the test and PR description so they do not claim server-side prompt-token validation unless the response contains the actual prompt IDs used by vLLM

rid=str(uuid.uuid4()),
metadata=metadata if not is_omitted(metadata) else {},
tokenizer=self.tokenizer,
processor=self.processor,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interrupted multimodal generation does not appear to resume from the partial output.

RemoteInfEngine handles an aborted request by appending the partial generated tokens to req.input_ids and resubmitting it. However, when vision_msg_vllm is present, VLLMBackend ignores the grown req.input_ids and resends the unchanged original messages to /v1/chat/completions.

The resulting trajectory can contain:

  • a partial generation from the original prompt; followed by
  • another generation restarted from that same original prompt.

We need to ensure that the actual prompt sent to vLLM on retry includes the partial generated tokens. This also needs an adapter-level abort/resubmit test that verifies the second request continues from the first request’s output.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The current multimodal vLLM chat-completions path resends the original messages and therefore cannot guarantee that an aborted request resumes from the grown req.input_ids.
This PR is now scoped to SGLang only, so I am not enabling or attempting to repair the multimodal vLLM path here. The vLLM follow-up should move to the token-in/token-out generate API and include an adapter-level abort/resubmit test verifying that the second request contains the partial output from the first request.

tokenizer=self.tokenizer,
processor=self.processor,
image_data=image_data if has_images else None,
vision_msg_vllm=([vision_messages_for_vllm] if has_images else None),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Multimodal requests are routed through /v1/chat/completions, but AReaL's vLLM wrapper applies _wait_if_paused() only to /v1/completions.

This means multimodal agent requests can enter while generation is paused for a weight update, bypassing the lifecycle protection used by the abort/update/resume path.

We need to ensure that the endpoint used by multimodal rollout requests is guarded by the same pause mechanism before enabling this path for training.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The existing vLLM lifecycle protection is attached to the completions path, while the multimodal path uses /v1/chat/completions, so enabling it here would bypass the expected pause/update/resume guard.
Multimodal agent workflow support in this PR is now limited to SGLang. Before vLLM support is enabled in a follow-up, the selected vLLM endpoint must participate in the same pause/update/resume lifecycle, with
coverage for requests arriving during a weight update.

Comment thread areal/experimental/openai/client.py Outdated
"The tokenizer chat template must return text before VLM processing."
)

images = [load_image(image) for image in image_data]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security related: image_data can contain arbitrary HTTP(S) URLs, and load_image() fetches them from the proxy worker.

A client that can access the rollout proxy can therefore make workers request link-local metadata endpoints or internal cluster services. An unreachable URL can also occupy the preprocessing thread because this call supplies no explicit timeout.

The preserved URL is not transported correctly to vLLM either: the current backend later treats the URL string as base64 and wraps it in a data: URI.

I suggest reject non-inline images and accept only bounded data URIs or raw base64. If remote fetching is required, it needs explicit scheme, host and IP allowlisting, redirect revalidation, finite timeouts, and byte/pixel limits.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I applied a narrowly scoped security fix: scheme-bearing image URLs are now rejected before image preprocessing, and the supported request contract is limited to inline data URIs or raw base64 image data.
This prevents the rollout worker from fetching arbitrary HTTP(S) resources and also removes the path where a remote URL could later be incorrectly wrapped as base64 for the inference backend.

global _message_preprocessors, _prefix_matcher
config = _engine.config
tokenizer = load_hf_tokenizer(config.tokenizer_path)
processor, tokenizer = load_hf_processor_and_tokenizer(config.tokenizer_path)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If AutoProcessor.from_pretrained() fails, load_hf_processor_and_tokenizer() catches the exception and returns processor=None. This setup accepts that state.

A later image request then follows the text-only prompt-preparation path and may still be sent to the vLLM multimodal route, but its exported trajectory has no vision tensors or multimodal token-type IDs.

I suggest fail explicitly whenever image_data is non-empty but no image processor is available. Text-only deployments can continue accepting processor=None; an image request must not silently become a text-only training sample.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. _prepare_prompt now fails immediately when image data is present but no multimodal processor is available. Text-only requests may still run with processor=None, but an image request can no longer be exported as a text-only training trajectory.

mm_token_type_ids + [0] * resp.output_len,
dtype=torch.long,
).unsqueeze(0)
result["multi_modal_input"] = [self.multi_modal_input or {}]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Performance/scalability related, this preserves the correct engine batch shape i.e. one multimodal-input dict per sequence, but the current ownership model duplicates the full vision tensors across turns.

Each interaction reprocesses and retains the conversation's complete image set. With individual export, a K-turn episode serializes K copies of the tensors. With concat, the wire payload may contain only the leaf, but the session cache still holds K copies until export.

Add a standalone Geometry3K Chat Completions agent and a dedicated SGLang example while preserving the existing VisionRLVR entry point.

Reject unsafe inline image inputs and missing multimodal processors, with focused unit coverage for request and image validation.
Keep the review fix limited to rejecting remote image URLs and avoid introducing custom decoding or resource limits.
@sitabulaixizawaluduo
sitabulaixizawaluduo marked this pull request as draft August 17, 2026 07:40
@sitabulaixizawaluduo
sitabulaixizawaluduo marked this pull request as ready for review August 20, 2026 07:57
@sitabulaixizawaluduo sitabulaixizawaluduo added the safe-to-test Ready to run unit-tests in a PR. label Aug 20, 2026
Keep backend-managed image URLs working for tokenizer-only clients while requiring a local processor for trainable v1 multimodal trajectories. Scope remote URL rejection to the local processor path so the shared client does not regress v2 inference.
@sitabulaixizawaluduo sitabulaixizawaluduo added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Aug 20, 2026
Comment thread areal/experimental/openai/client.py Outdated
"""Build model-ready prompt tokens and vision tensors with an HF processor."""
from transformers.image_utils import load_image

if any("://" in image for image in image_data):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[NIT] The new validation blocks scheme-bearing URLs, but it does not enforce the stated inline-only contract.

transformers.load_image() also accepts local filesystem paths. A value such as /shared/private-image.png does not contain ://, so it passes this check and is opened by the proxy worker.

Suggested fix: positively decode the value as base64 instead of passing arbitrary strings to load_image()

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Change the local processor path to positively decode the input as base64 and construct the image from in-memory bytes, rejecting any value that is not valid inline image data. The tokenizer-only backend forwarding path will remain unchanged.

Prevent local multimodal preprocessing from interpreting request strings as filesystem paths. Validate base64 input and construct images from in-memory bytes before invoking the Hugging Face image helper.
@Adiactive

Copy link
Copy Markdown
Contributor

LGTM

@sitabulaixizawaluduo sitabulaixizawaluduo added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Aug 21, 2026
@sitabulaixizawaluduo sitabulaixizawaluduo added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Aug 21, 2026
Keep the SGLang tokenizer configuration advisory while preventing strict multimodal agent trajectories from entering the unsupported vLLM path.

Key changes:
- Warn when SGLang multimodal rollout may reprocess expanded tokens
- Reject strict multimodal agent requests on the vLLM backend
- Cover both OpenAI request paths and SGLang launch behavior

Refs: #1606
@sitabulaixizawaluduo sitabulaixizawaluduo added safe-to-test Ready to run unit-tests in a PR. and removed safe-to-test Ready to run unit-tests in a PR. labels Aug 24, 2026
@sitabulaixizawaluduo
sitabulaixizawaluduo merged commit 94ce165 into main Aug 26, 2026
13 checks passed
@sitabulaixizawaluduo
sitabulaixizawaluduo deleted the feature/vlm-agent-workflow branch August 26, 2026 03:11
Adiactive added a commit to Adiactive/AReaL that referenced this pull request Aug 26, 2026
Adapt the processor-backed proxy path from areal-project#1606 so multimodal agent
interactions export the same prompt tokens and vision tensors consumed by
training.

Key changes:
- Process inline images outside the async event loop
- Validate image payloads and processor output shapes
- Export and serialize multimodal training fields
- Normalize mixed text and vision batches
Adiactive added a commit to Adiactive/AReaL that referenced this pull request Aug 26, 2026
Backport the focused single-turn agent from areal-project#1606 so the multimodal
proxy path can be tested without multi-turn concat or tool feedback.

Key changes:
- Add the Geometry3K Chat Completions agent
- Add a PPOTrainer entry point for proxy rollouts
- Cover image injection, reward scoring, and request construction
Adiactive added a commit to Adiactive/AReaL that referenced this pull request Aug 26, 2026
Adapt the processor-backed proxy path from areal-project#1606 so multimodal agent
interactions export the same prompt tokens and vision tensors consumed by
training.

Key changes:
- Process inline images outside the async event loop
- Validate image payloads and processor output shapes
- Export and serialize multimodal training fields
- Normalize mixed text and vision batches
Adiactive added a commit to Adiactive/AReaL that referenced this pull request Aug 26, 2026
Backport the focused single-turn agent from areal-project#1606 so the multimodal
proxy path can be tested without multi-turn concat or tool feedback.

Key changes:
- Add the Geometry3K Chat Completions agent
- Add a PPOTrainer entry point for proxy rollouts
- Cover image injection, reward scoring, and request construction
HwVanICI pushed a commit that referenced this pull request Aug 26, 2026
Adapt the processor-backed proxy path from #1606 so multimodal agent
interactions export the same prompt tokens and vision tensors consumed by
training.

Key changes:
- Process inline images outside the async event loop
- Validate image payloads and processor output shapes
- Export and serialize multimodal training fields
- Normalize mixed text and vision batches
HwVanICI pushed a commit that referenced this pull request Aug 26, 2026
Backport the focused single-turn agent from #1606 so the multimodal
proxy path can be tested without multi-turn concat or tool feedback.

Key changes:
- Add the Geometry3K Chat Completions agent
- Add a PPOTrainer entry point for proxy rollouts
- Cover image injection, reward scoring, and request construction

This branch was previously deployed

1 inactive deployment
AReaL-unittests — b3886c5c Deployed Aug 24, 2026 by sitabulaixizawaluduo via Run AReaL unit tests (sglang) #806
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

safe-to-test Ready to run unit-tests in a PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants