English | 中文
This is the mandatory low-context router for InferenceX work. Pick the one page that owns the task, then follow only its source links. Repository source files and workflows remain authoritative.
Paths and shell commands in these guides are relative to inferencex-e2e/ unless stated otherwise. The Python manifest, lockfile, and .python-version live in that project directory. Repository-wide policy and GitHub workflows remain at the repository root.
| Task | Open first | Then inspect |
|---|---|---|
| Add a model or GPU benchmark | Config reference | closest benchmark script, launcher, master YAML, changelog |
| Modify an existing config | Config reference | validation schema, generator, runtime consumer |
| Add a runner | Runner setup | configs/runners.yaml, launcher |
| Change srt-slurm or llm-d | Recipe reference | Recipe YAML, master config, srtctl mapping, launcher |
| Change MTP | Draft-model precision | MTP sibling, draft model, chat-template path |
| Validate a matrix | CI procedures | generator CLI and Pydantic validation |
| Dispatch or monitor a run | CI procedures | e2e-tests.yml, run logs, artifacts |
| Prepare a PR sweep | CI procedures | run-sweep.yml, labels, changelog delta |
| Reuse a green sweep | CI procedures | reuse gate, source artifacts, merge helper |
| Add or debug evals | Eval and AgentX procedures | EVALS.md, eval templates, score validator |
| Run AgentX | Eval and AgentX procedures | agentic config, trace source, live-run skill |
| Inspect a result or ingest | Recovery and results procedures | artifact schema, collector, app ingest workflow |
| Recover failed ingest | Recovery and results procedures | recovery tool, source-run artifacts, ancestry rules |
| Debug a runner or workspace | Recovery and results procedures | launcher cleanup, .claude/commands/ cluster playbooks, cluster logs |
| Page | Open it for |
|---|---|
index.md / index_zh.md |
This task router and its Chinese counterpart |
architecture.md / architecture_zh.md |
Config-to-result flow, ownership boundaries, artifacts, and InferenceX-app handoff |
power_model README |
Installation, CLI usage, supported systems, and power-model assumptions |
ci-procedures.md / ci-procedures_zh.md |
Matrix generation, validation, dispatch, PR sweeps, reuse, staging, and artifact downloads |
eval-agentx-procedures.md / eval-agentx-procedures_zh.md |
Eval and AgentX selection, execution, scoring, evidence, and live-run diagnosis |
agentx-standalone.md / agentx-standalone_zh.md |
Install the pinned AgentX client and replay traces against an existing server without CI or Slurm |
results-and-ingestion.md / results-and-ingestion_zh.md |
Published-result lookup, artifact identities and schemas, app ingestion, dedupe, and provenance |
recovery-results-procedures.md / recovery-results-procedures_zh.md |
Result processing, ingest verification and recovery, runner cleanup, and failure classification |
testing.md / testing_zh.md |
Local checks, smoke runs, evidence standards, and review gates |
troubleshooting.md / troubleshooting_zh.md |
Failure-layer diagnosis, known cases, safe remediation, and stop conditions |
PR_REVIEW_CHECKLIST.md / PR_REVIEW_CHECKLIST_zh.md |
CODEOWNER review and exact sign-off requirements |
| Reference | Owns |
|---|---|
AGENTS.md |
Mandatory low-context agent policy and benchmark invariants |
CONTRIBUTING.md |
PR review, CODEOWNER sign-off, sweep reuse, merge, and post-merge duties |
.github/AGENT_OPERATIONS.md |
Translation terms, sweep labels, dispatch, eval selection, power, metrics, and artifacts |
configs/CONFIGS.md |
Master-config schema, search spaces, runners, and topology fields |
.github/workflows/README.md |
Generator examples, workflow operation, and reuse policy |
infx/evals/EVALS.md |
Eval task, execution, collection, and validation contracts |
benchmarks/multi_node/srt-slurm-recipes/RECIPES.md |
Disaggregated recipe registration and master-config coupling |
utils/runner_setup/RUNNER_SETUP.md |
Runner provisioning and setup |
MODELS.md |
Supported models, hardware coverage, and naming |
klaud.md / klaud_zh.md |
Klaud Cold selection, ownership, validation and recovery |
klaud-reporting.md / klaud-reporting_zh.md |
Klaud PR body, progress comments, numeric comparisons and final preflight |
benchmarks/srt_agentic.sh |
AgentX trace replay client shared by single- and multi-node srt-slurm recipes |
- Open only the focused page and source sections needed for the task.
- Do not load large YAML, JSON, logs, generated matrices, or whole reference files when a filtered view answers the question.
- Source code, workflow YAML, schemas, launchers, and collectors win over explanatory docs.
- When behavior changes, update the nearest English guide first and its
_zh.mdcounterpart in the same change.