An AI agent is an LLM wired into a loop where it can take actions, calling tools, reading the results, and deciding what to do next, rather than just producing one text response to one prompt.
A plain chat model answers a question and stops. An agent is set up to keep going: it can decide to search the web, run a calculation, read a file, or call an API, look at what came back, and decide on its next step, repeating until it judges the task done.
Agents are what let an LLM go beyond answering questions from memory into actually doing things, booking, searching, editing, executing code, which is the direction most serious AI product work is heading.
An agent runs a loop: the LLM is given the task and a list of available tools, it decides whether to answer directly or call a tool, the tool executes and returns a result, that result is added back into the model's context, and the loop repeats until the model produces a final answer or hits a stopping condition.
Every extra step in an agent's loop is another LLM call, so agent tasks are typically far slower and more expensive per completed task than a single chat response, and each step is a place accuracy can compound in the wrong direction.
The tasks this site measures are all single-turn, one prompt in, one recorded response out, deliberately simpler than a full agent loop, so the seconds shown here are the floor an agentic version of the same task would build on top of, not the whole picture.