Tool calling is a model's ability to output a structured request to run a specific function, with specific arguments, instead of (or alongside) a plain text answer, letting application code execute it and feed the result back in.
The model is told, in advance, about a set of functions it's allowed to use: their names, what they do, and what arguments they take. When it decides one is useful, instead of guessing at an answer, it outputs a structured call like get_weather(city="Mumbai") for the surrounding application to actually run.
It's the mechanism that turns a text-only model into something that can look up real data, take real actions, or run real calculations, reliably enough that application code can parse and execute the request without guessing at the model's intent.
The model is provided a schema (usually JSON Schema) describing each available tool's name, description, and parameters. During generation, it can choose to emit a tool call matching that schema instead of free text; the calling application validates it, executes the real function, and returns the result as a new message the model reads before continuing.
A model that hallucinates an argument or calls the wrong tool can trigger a real action with real consequences, which is why schema validation and permission checks sit between the model's output and actual execution in any serious system.