Website - Paper - Benchmark - Doc
- 2026-07-27: We added results from Claude Code Opus-5 and Codex GPT-5.6-sol
- 2026-07-03: We released the benchmark.
- 2026-07-01: We released our paper and website.
HealthAgentBench is a terminal-based benchmark suite for evaluating agents on realistic health tasks. Each task drops an agent into a terminal environment where it must inspect data, use tools, reason, and act to solve a concrete clinical or biomedical problem, then a task-specific verifier scores the result. The figure below shows the main results from frontier agents on this benchmark.
Task success rate across all 54 Γ 3 trials (Wilson 95% CI), with cost and time per task.
HealthAgentBench currently ships seven task categories:
| Category (name in this codebase) | # Tasks | What is the task |
|---|---|---|
X-ray Report Correction (xray_report_correction) |
10 | Correct a chest X-ray radiology report for the latest MIMIC-CXR study, scored with the CheXprompt LLM judge verifier. |
Pathology Tumor Area Selection (tumor_area_selection_pathology) |
10 | Predict the set of tumor-containing tiles over public whole-slide H&E pathology images. |
EHR Format Conversion (ehr_to_meds_etl) |
1 | ETL raw MIMIC-IV EHR data into the MEDS common data format. |
CT Abnormality Classification (ct_abnormality) |
10 | Patient-level chest-CT abnormality detection built on the CT-RATE dataset. |
Clinical Trial Matching (clinical_trial_matching) |
9 | Identify every clinical trial a patient is eligible for from a candidate pool (TREC Clinical Trials 2021, set-recall). |
EHR Data Quality Auditing (ehr_data_quality) |
8 | Flag rows containing injected data-quality errors in a corrupted MIMIC-IV EHR subset. |
EHR Event Modelling (ehr_event_modelling) |
6 | Predict future clinical events over longitudinal EHR timelines (Stanford SHAH lab's EHRSHOT benchmark). |
| Total | 54 |
Each task has its own README.md under tasks/ with the task's category, success criteria, data/credentials, and commands to run that task or its whole category.
HealthAgentBench/ # repo root
βββ README.md
βββ pyproject.toml # package + dependency config (uv; requires Python >=3.12)
βββ .env.example # template for gated-dataset credentials
βββ assets/ # figures & media (banner, hero chart, and shared task category data assets)
βββ website/ # Astro leaderboard / docs site
βββ LICENSE
βββ SECURITY.md
βββ tasks/ # 54 Harbor tasks, one flat directory per task
βββ xray_report_correction_case_*/ # 10 tasks - Longitudinal X-ray report correction
βββ tumor_area_selection_pathology_slide_*/ # 10 tasks β WSI tumor-tile selection
βββ ct_abnormality_valid_*/ # 10 tasks β chest-CT abnormality detection
βββ clinical_trial_matching_task_*/ # 9 tasks β patient β trial matching
βββ ehr_data_quality_task_*/ # 8 tasks β flag injected EHR errors
βββ ehr_event_modelling_*/ # 6 tasks β future clinical-event prediction
βββ ehr_to_meds_etl/ # 1 task β MIMIC-IV β MEDS ETL
All 54 tasks live as flat, sibling directories directly under tasks/ β the task
directory name is prefixed with its category (e.g. xray_report_correction_case_01,
ct_abnormality_valid_16_a_1).
Every task follows the same Harbor layout (task.toml + instruction.md +
environment/ + tests/) plus a README.md describing that task, how to run it, and
how to run its whole category.
# Clone the repo
git clone https://github.com/microsoft/HealthAgentBench.git
cd HealthAgentBench
# Install dependencies
uv sync --all-extrasPython version requirement: >=3.12.
Some of the task categories above require gated datasets and
per-user credentials before the container can run. Obtain credentials following the
instructions below and fill in the credentials in .env (there is a file
.env.example showing the template).
- ehr_event_modelling β EHRSHOT
(Redivis, Stanford SHAH lab). Apply for access from EHRSHOT; once approved, create a
Redivis API token and set it
as
REDIVIS_API_TOKENin.env. - ct_abnormality β CT-RATE
(Hugging Face, OpenRAIL gated). Accept the dataset agreement from CT-RATE, then set your
Hugging Face token as
HF_TOKENin.env. - xray_report_correction β MIMIC-CXR v2.1.0
(radiology reports) and MIMIC-CXR-JPG v2.1.0
(JPG frames + metadata), both on PhysioNet. Once approved, set your PhysioNet
username as
PN_USERand password asPN_PASSin.env.
X-ray Report Correction is scored by an LLM-based judge based on the CheXprompt verifier. We use GPT-5.4 as the default judge. Configure one of two paths in .env (the
verifier auto-detects which is present): (a) vanilla OpenAI β set CHEXPROMPT_OPENAI_API_KEY and optionally CHEXPROMPT_OPENAI_BASE_URL; or (b) Azure OpenAI β
set CHEXPROMPT_AZURE_OPENAI_API_KEY, CHEXPROMPT_AZURE_OPENAI_ENDPOINT, and CHEXPROMPT_AZURE_OPENAI_API_VERSION. To change the default judge, set CHEXPROMPT_DEPLOYMENT to a different model or deployment name.
The coding agent also needs credentials. For subscription authentication, run claude setup-token and set CLAUDE_CODE_OAUTH_TOKEN for Claude Code, or log in to Codex on the host and set CODEX_FORCE_AUTH_JSON=1 to use $HOME/.codex/auth.json. Alternatively, configure ANTHROPIC_API_KEY for Claude Code or OPENAI_API_KEY for Codex. ANTHROPIC_BASE_URL and OPENAI_BASE_URL are optional and are only needed when using third-party endpoints. See .env.example for the template.
Make sure you have .env at this repo directory so that the containers can read the environment variables.
set -a; source .env; set +aWe use Harbor to run evaluation. Refer to the Harbor Background section for more details especially if you want to evalute additional agents.
--n-attempts specifies how many independent times each task is run and --n-concurrent specifies how many
trials run in parallel.
uv run harbor run \
--path tasks \
--agent claude-code \
--model claude-opus-4-8 \
--agent-kwarg reasoning_effort=xhigh \
--agent-kwarg disallowed_tools="WebSearch WebFetch" \
--n-attempts 1 --n-concurrent 5 \
--jobs-dir <the output directory> \Harbor will print out a mean column that records the success rate and all the tasks' rewards (1 or 0) across the runs. You will find full results in the output directory. We also record additional metrics in addition to binary pass in verifier/metrics.json in each task subdirectory in output directory.
Note:
- The harbor runs above will download data and mount the data to
assets/<task_category>/to speed up multiple task setups using the same data assets. You will need at least 30GB on disk available to download the data assets. - This repo does not host labels, but the harbor runs above will fetch gold labels which will appear under
tasks/<task_name>/testsafter the run. - We suggest disallowing web browsing / web fetching when running this benchmark, so the agent can't cheat by searching for gold labels online. Harbor's built-in Codex agent has no web-search toggle, so you need to create a thin wrapper that subclasses it and adds a CliFlag mapping a kwarg (eg. disable_web_search) to Codex's
-c web_search="disabled"config override (see the Codex config reference and the Harbor agents docs). - Check the exception errors (if any) reported by Harbor. We treat agenttimeout trials as failures (reward=0) when reporting the overall success rate.
uv run harbor run \
--path tasks/xray_report_correction_case_01 \
--agent claude-code \
--model claude-opus-4-8 \
--agent-kwarg reasoning_effort=xhigh \
--agent-kwarg disallowed_tools="WebSearch WebFetch" \
--n-attempts 1 --n-concurrent 1Point --path at tasks/ and glob the category name prefix with
--include-task-name (globs are supported; quote it so your shell doesn't expand
the *):
uv run harbor run \
--path tasks \
--include-task-name "xray_report_correction_*" \
--agent claude-code \
--model claude-opus-4-8 \
--agent-kwarg reasoning_effort=xhigh \
--agent-kwarg disallowed_tools="WebSearch WebFetch" \
--n-attempts 1 --n-concurrent 5Category prefixes for --include-task-name: clinical_trial_matching_*, ct_abnormality_*, ehr_data_quality_*, ehr_event_modelling_*, ehr_to_meds_etl, tumor_area_selection_pathology_*, xray_report_correction_*
This project uses Harbor as the terminal-task execution and evaluation substrate. Harbor provides a consistent trial lifecycle (agent run, verifier run, and artifacts), while HealthAgentBench adds domain-specific health tasks, Harbor task environments, and benchmark integrations. Refer to the pointers below if you would like to run on additional agents supported by Harbor.
Important pointers:
- Harbor repo: https://github.com/harbor-framework/harbor
- Harbor docs/wiki: https://deepwiki.com/harbor-framework/harbor
- Stable Harbor version used here:
0.8.0(see theharbor==pin inpyproject.toml)
If you use HealthAgentBench in your research, please cite:
@misc{liu2026healthagentbench,
title = {HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents},
author = {Liu, Qianchu and Zhang, Sheng and Qin, Guanghui and Valanarasu, Jeya Maria Jose and Rokuss, Maximilian and Lu, Mingyu and Ossowski, Timothy and Chaves, Juan Manuel Zambrano and Wong, Cliff and Argaw, Peniel and Hasija, Yashna and Wei, Mu and Yim, Wen-wai and Liu, Qin and Jing, Zilin and Entenmann, Jason and Usuyama, Naoto and Naumann, Tristan and Poon, Hoifung},
year = {2026}
}
