Agents and Tool Use
The loop
messages = [{"role": "user", "content": task}]
for step in range(1, max_steps + 1):
response = fake_llm.chat(messages, tools=tools, temperature=0)
message = response["choices"][0]["message"]
if not message.get("tool_calls"):
return (message["content"], step) # prose means finished
call = message["tool_calls"][0]
args = json.loads(call["function"]["arguments"])
result = registry[call["function"]["name"]](**args)
messages += [message, {"role": "tool", "content": str(result)}]
raise RuntimeError("the agent did not finish") # the termination proof
The limit is written before the loop, not after the first incident. Falling out of the loop is a failure, so it raises — returning the last tool result dresses failure as an answer.
Memory
Each step appends two messages and the whole history is re-sent, so cost grows quadratically with steps.
trimmed = [messages[0]] + messages[-keep_recent:] # keep the task, keep the recent
Never drop messages[0]. An agent that forgets the task finishes a different one,
confidently.
Plans
Validate the whole plan before executing any of it:
- not empty
- not longer than the limit
- every step names a tool that exists
Finding an invalid step at step 6 means five steps of side effects have already happened.
MCP
Tools become runtime data from an external source, so:
merged[server_name + "." + tool_name] = tool # namespace, never merge flat
Two servers offering search is normal; one silently shadowing the other is a bug the
operator cannot see.
Three guards, all of them
| Guard | Catches |
|---|---|
| step limit | an unbounded loop |
| cost budget | expensive steps inside the step limit |
| no-progress detector | the same call repeating forever |
signature = (tool_name, tuple(sorted(args.items())))
if self.seen[signature] > max_repeats:
return "no progress"
Permissioning
Everything the agent does, your process does. Permission at the dispatcher, not in the prompt: a prompt is a request, an absent tool is a boundary.
if tool_name not in allowed: # allow-list, never deny-list
return (False, "not allowed")
if tool_name.startswith(("write_", "delete_")) and tool_name not in writable:
return (False, "read-only")
Reads and writes are different permissions. Most agent tasks need only reads.
Human in the loop
Gate by consequence, not by tool:
| Action | Ask? |
|---|---|
| read | no |
| reversible write | no |
| irreversible or externally visible | always |
if (tool_name, tuple(sorted(args.items()))) not in approvals:
return ("ask", describe(tool_name, args))
Approval is per call, not per tool — yes to "email 2 people" is not yes to "send email". And describe the action in terms someone will actually read: "Send an email to 4,182 recipients?" rather than "Run tool send_email?".