AgentBoard: An Analytical Evaluation Board of LLM Agents
2024-01-23 · Verified
Paper info
Start research from this paper
Choose a research mode to carry this paper into the workspace context.
Reproducibility status
No structured reproducibility record has been added yet.
Model-role topology
This view only shows recorded model-role relations; it does not invent workflow edges.
This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.
Model roles
GPT-4
Baseline · weights not updated · GPT-4 benchmarked on AgentBoard.
GPT-3.5 Turbo
Baseline · weights not updated · GPT-3.5 Turbo benchmarked on AgentBoard.
Llama 2 70B Chat
Baseline · weights not updated · Llama-2-70B-Chat benchmarked on AgentBoard.
Qwen 7B Chat
Baseline · weights not updated · Qwen-7B-Chat benchmarked on AgentBoard.