Skip to content

Agents and Tool Use

The loop

messages = [{"role": "user", "content": task}]

for step in range(1, max_steps + 1):
    response = fake_llm.chat(messages, tools=tools, temperature=0)
    message = response["choices"][0]["message"]

    if not message.get("tool_calls"):
        return (message["content"], step)        # prose means finished

    call = message["tool_calls"][0]
    args = json.loads(call["function"]["arguments"])
    result = registry[call["function"]["name"]](**args)
    messages += [message, {"role": "tool", "content": str(result)}]

raise RuntimeError("the agent did not finish")   # the termination proof

The limit is written before the loop, not after the first incident. Falling out of the loop is a failure, so it raises — returning the last tool result dresses failure as an answer.

Memory

Each step appends two messages and the whole history is re-sent, so cost grows quadratically with steps.

trimmed = [messages[0]] + messages[-keep_recent:]   # keep the task, keep the recent

Never drop messages[0]. An agent that forgets the task finishes a different one, confidently.

Plans

Validate the whole plan before executing any of it:

  • not empty
  • not longer than the limit
  • every step names a tool that exists

Finding an invalid step at step 6 means five steps of side effects have already happened.

MCP

Tools become runtime data from an external source, so:

merged[server_name + "." + tool_name] = tool    # namespace, never merge flat

Two servers offering search is normal; one silently shadowing the other is a bug the operator cannot see.

Three guards, all of them

GuardCatches
step limitan unbounded loop
cost budgetexpensive steps inside the step limit
no-progress detectorthe same call repeating forever
signature = (tool_name, tuple(sorted(args.items())))
if self.seen[signature] > max_repeats:
    return "no progress"

Permissioning

Everything the agent does, your process does. Permission at the dispatcher, not in the prompt: a prompt is a request, an absent tool is a boundary.

if tool_name not in allowed:                       # allow-list, never deny-list
    return (False, "not allowed")
if tool_name.startswith(("write_", "delete_")) and tool_name not in writable:
    return (False, "read-only")

Reads and writes are different permissions. Most agent tasks need only reads.

Human in the loop

Gate by consequence, not by tool:

ActionAsk?
readno
reversible writeno
irreversible or externally visiblealways
if (tool_name, tuple(sorted(args.items()))) not in approvals:
    return ("ask", describe(tool_name, args))

Approval is per call, not per tool — yes to "email 2 people" is not yes to "send email". And describe the action in terms someone will actually read: "Send an email to 4,182 recipients?" rather than "Run tool send_email?".

This cheatsheet is the summary. If you want to build it yourself, the Agents and Tool Use course walks you through it in the browser — the first lesson is free.