Search the atlas

Esc to close · Cmd/Ctrl + K to open

Why these models were selected

Roles and weight updates come from the paper case record. Selection reasons are written only when the paper or code states them; unknown is never inferred as fact.

Paper model roles and selection-evidence status
ModelRoleWeights updatedSelection basis
GPT-4baselineNoNot recorded: verify in the paper/code
GPT-3.5 TurbobaselineNoNot recorded: verify in the paper/code
Llama 2 70B ChatbaselineNoNot recorded: verify in the paper/code
Vicuna 13BbaselineNoNot recorded: verify in the paper/code

AgentBench: Evaluating LLMs as Agents

2023-08-07 · Verified

Paper info

Evolution targets
Agent evaluation, Actor policy
Benchmark
AgentBench, ALFWorld, WebShop, OSWorld
Category
Self-evolving agent, Reasoning, Tool use
Model relations
4

Start research from this paper

Choose a research mode to carry this paper into the workspace context.

Reproducibility status

Code Not verified
Checkpoint Not verified
Config Not verified
Environment Not verified

No structured reproducibility record has been added yet.

Model-role topology

This view only shows recorded model-role relations; it does not invent workflow edges.

Baseline
GPT-4GPT-3.5 TurboLlama 2 70B ChatVicuna 13B

This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.

Model roles

GPT-4

Baseline · weights not updated · GPT-4 benchmarked as an AgentBench agent.

GPT-3.5 Turbo

Baseline · weights not updated · GPT-3.5 Turbo benchmarked as an AgentBench agent.

Llama 2 70B Chat

Baseline · weights not updated · Llama-2-70B-Chat benchmarked as an open AgentBench agent.

Vicuna 13B

Baseline · weights not updated · Vicuna-13B benchmarked as an open AgentBench agent.