The idea in one minute#
A model can only produce text. It becomes able to do things when your program offers it tools — functions described by a name, a purpose and a JSON schema for their arguments — and runs a loop: send the conversation and the tool descriptions; if the model answers with a tool call, execute it and append the result to the conversation; ask again; stop when the model answers in plain text or a limit is reached.
That loop is the whole mechanism behind “agents”. It is about fifty lines. Everything hard about it is engineering you already know: validating untrusted input, bounding time and steps, handling errors, and deciding what the program is permitted to do.
An analogy#
A researcher working by correspondence with an assistant in the archive. The researcher (the model) cannot enter the archive; they can only write requests — “fetch file 12”, “add these two figures”. The assistant (your program) decides whether each request is allowed, carries it out, and writes back what was found. The exchange continues until the researcher has enough to write the answer. The assistant, not the researcher, holds the keys.
A picture#
flowchart TB
START["User task"] --> CALL["Call the model<br/>messages + tool definitions"]
CALL --> R{"Response"}
R -->|"final text"| DONE["Return the answer"]
R -->|"tool call: name + JSON arguments"| V{"Known tool?<br/>Arguments valid?<br/>Allowed?"}
V -->|"no"| ERR["Append an error result<br/>so the model can correct itself"]
V -->|"yes"| RUN["Run the tool<br/>with a timeout"]
RUN --> RES["Append the result<br/>as a 'tool' message"]
ERR --> LIM{"Steps or budget<br/>exhausted?"}
RES --> LIM
LIM -->|"no"| CALL
LIM -->|"yes"| STOP["Stop: return what we have,<br/>marked incomplete"]
class START,DONE neutral
class CALL io
class R,V,LIM queue
class RUN,RES compute
class ERR,STOP warnHow it really works#
Describing a tool#
The model sees, for each tool, a name, a description written for the model, and a schema:
{
"type": "function",
"function": {
"name": "get_gpu_status",
"description": "Return memory use and temperature for one GPU on one node.",
"parameters": {
"type": "object",
"properties": {
"node": {"type": "string"},
"gpu": {"type": "integer", "minimum": 0}
},
"required": ["node", "gpu"]
}
}
}The response then contains, instead of text, something like
{"name": "get_gpu_status", "arguments": "{\"node\":\"n7\",\"gpu\":2}"} with a call ID. You
run it and send back a message with role tool, that ID, and the result as text.
The descriptions are part of your prompt: clear names, one job per tool, and precise argument descriptions matter more than clever orchestration.
The loop’s rules#
| Rule | Why |
|---|---|
| Arguments are untrusted. Parse strictly, validate types and ranges | The model can produce anything, and its input may contain text written by an attacker |
| Unknown tool or bad arguments → return the error to the model | It usually corrects itself on the next turn |
| Limit steps | A confused model loops |
| Limit time: a context for the whole task and a timeout per tool | A stuck tool must not hang the task |
| Limit cost: count tokens across calls | Context grows every step; so does the bill |
| Bound result size: truncate or summarize large tool output | A 2 MB log pasted into the conversation destroys the context window |
| Run independent tool calls concurrently | Models often request several at once; an error group fits (IV.06) |
| Make side effects explicit: read-only tools run freely; writes need a policy or a human approval | The program holds the keys |
| Record every step | A trace of model calls and tool calls is the only way to debug it |
Safety, concretely#
- Least privilege. A tool that runs shell commands or SQL gives the model everything that
process can do. Prefer narrow tools (
get_order(id)) over general ones (run_sql(query)). - Prompt injection. Text returned by a tool — a web page, a document, an email — may contain instructions. The model cannot reliably tell them from yours. Treat any task that mixes untrusted content with consequential tools as requiring a check outside the model.
- Idempotency. A retried step may run a tool twice. Give writes an idempotency key.
- Sandbox code execution and file access.
The Model Context Protocol#
Writing a custom integration for every model and every tool does not scale. MCP is an open protocol in which a server exposes tools (plus resources and prompts) and any compatible client — an IDE, an agent framework, your loop — can discover and call them over stdio or HTTP. There is an official Go SDK for both sides. A Go service that already has a good internal API becomes usable by agents by exposing it as an MCP server, and Go’s strengths — a single binary, concurrency, strict typing of arguments — fit that job well.
Frameworks, or not#
Genkit, Google’s ADK, Eino and LangChainGo provide the loop, tool plumbing, tracing and multi-agent patterns. They are worth it when you need their integrations. For a single loop with a handful of tools, writing it yourself is short, has no dependencies, and leaves nothing hidden — which matters when the thing you are debugging is non-deterministic.
Observing it#
One trace per task; one span per model call and per tool call, with token counts, tool names, durations and outcomes. Watch steps per task, the tool error rate, and tokens per task. The conventions are in Observability V.04.
Where this goes next#
Planning, memory across tasks, multiple cooperating agents, evaluation of agent behaviour — the subject of the Agentic Engineering path, which is in preparation. This lesson is the foundation it will build on: a loop, a registry, and limits.
Code#
The “model” here is a scripted stand-in so the program is deterministic and runs offline; swap it for the client from lesson 06 and the loop is unchanged. Watch it make a bad call, receive the error, and correct itself.
// agent.go — the tool-calling loop: registry, validation, limits, and error feedback.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"strings"
"time"
)
type Message struct {
Role string // system, user, assistant, tool
Content string
Call *ToolCall // set when the assistant asks for a tool
}
type ToolCall struct {
Name string
Args json.RawMessage
}
type Tool struct {
Name, Description string
ReadOnly bool
Run func(ctx context.Context, args json.RawMessage) (string, error)
}
// decode parses arguments strictly: unknown fields are an error.
func decode(args json.RawMessage, into any) error {
dec := json.NewDecoder(strings.NewReader(string(args)))
dec.DisallowUnknownFields()
return dec.Decode(into)
}
var fleet = map[string][]int{"n7": {71, 68, 93, 70}} // node → GPU temperatures
func tools() map[string]Tool {
return map[string]Tool{
"gpu_temperature": {
Name: "gpu_temperature", Description: "Temperature in C of one GPU on one node.", ReadOnly: true,
Run: func(_ context.Context, args json.RawMessage) (string, error) {
var a struct {
Node string `json:"node"`
GPU *int `json:"gpu"`
}
if err := decode(args, &a); err != nil {
return "", fmt.Errorf("bad arguments: %w", err)
}
temps, ok := fleet[a.Node]
if !ok {
return "", fmt.Errorf("unknown node %q", a.Node)
}
if a.GPU == nil || *a.GPU < 0 || *a.GPU >= len(temps) {
return "", fmt.Errorf("gpu must be between 0 and %d", len(temps)-1)
}
return fmt.Sprintf(`{"celsius": %d}`, temps[*a.GPU]), nil
},
},
"drain_node": {
Name: "drain_node", Description: "Remove a node from service.", ReadOnly: false,
Run: func(context.Context, json.RawMessage) (string, error) { return `{"drained": true}`, nil },
},
}
}
// scriptedModel stands in for an LLM: it looks at the conversation and decides the next move.
func scriptedModel(msgs []Message) Message {
var results []string
for _, m := range msgs {
if m.Role == "tool" {
results = append(results, m.Content)
}
}
switch len(results) {
case 0: // first attempt: a wrong argument name, as real models sometimes produce
return Message{Role: "assistant", Call: &ToolCall{"gpu_temperature", json.RawMessage(`{"node":"n7","index":2}`)}}
case 1: // it reads the error and corrects itself
return Message{Role: "assistant", Call: &ToolCall{"gpu_temperature", json.RawMessage(`{"node":"n7","gpu":2}`)}}
case 2: // it decides to act
return Message{Role: "assistant", Call: &ToolCall{"drain_node", json.RawMessage(`{"node":"n7"}`)}}
default:
return Message{Role: "assistant", Content: "GPU 2 on n7 is at 93 C, above the 85 C limit. Drain result: " + results[len(results)-1]}
}
}
var ErrStepLimit = errors.New("step limit reached")
// Run is the loop. approve decides whether a tool with side effects may run.
func Run(ctx context.Context, task string, registry map[string]Tool, maxSteps int, approve func(ToolCall) bool) (string, []Message, error) {
msgs := []Message{
{Role: "system", Content: "You operate a GPU fleet. Use tools; do not guess."},
{Role: "user", Content: task},
}
for step := 1; step <= maxSteps; step++ {
if err := ctx.Err(); err != nil {
return "", msgs, err
}
reply := scriptedModel(msgs)
msgs = append(msgs, reply)
if reply.Call == nil {
return reply.Content, msgs, nil // plain text: the task is finished
}
call := *reply.Call
var result string
tool, known := registry[call.Name]
switch {
case !known:
result = fmt.Sprintf(`{"error": "no tool named %q"}`, call.Name)
case !tool.ReadOnly && !approve(call):
result = `{"error": "not approved: this action changes the system"}`
default:
tctx, cancel := context.WithTimeout(ctx, 2*time.Second) // a stuck tool cannot hang the task
out, err := tool.Run(tctx, call.Args)
cancel()
if err != nil {
result = fmt.Sprintf(`{"error": %q}`, err.Error()) // the model sees the error and can retry
} else {
result = out
}
}
if len(result) > 2000 {
result = result[:2000] + "...[truncated]" // never let one result flood the context
}
fmt.Printf(" step %d: %s(%s) → %s\n", step, call.Name, call.Args, result)
msgs = append(msgs, Message{Role: "tool", Content: result})
}
return "", msgs, ErrStepLimit
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
task := "Check GPU 2 on node n7 and take it out of service if it is overheating."
fmt.Println("policy: writes are NOT approved")
answer, msgs, err := Run(ctx, task, tools(), 6, func(ToolCall) bool { return false })
fmt.Printf(" answer: %s\n messages: %d err: %v\n\n", answer, len(msgs), err)
fmt.Println("policy: writes approved")
answer, _, err = Run(ctx, task, tools(), 6, func(c ToolCall) bool { return c.Name == "drain_node" })
fmt.Printf(" answer: %s\n err: %v\n\n", answer, err)
fmt.Println("step limit of 2")
_, _, err = Run(ctx, task, tools(), 2, func(ToolCall) bool { return true })
fmt.Println(" err:", err)
}Remember this#
- An agent is a loop: call the model, run the tool it asks for, append the result, repeat.
- Tool arguments are untrusted input: parse strictly, validate, and feed errors back.
- Limit steps, time, tokens and result size. Separate read-only tools from ones with effects.
- The program holds the keys; the model only asks.
- Trace every step.
Try it#
- Run
agent.go. Follow the three runs: where does the model recover from its own mistake, and where does policy stop it? - Make the scripted model request two read-only tool calls in one turn and run them concurrently with a limit, preserving the order of results.
- Replace
scriptedModelwith the client from lesson 06 against a real tool-calling model. What changes in the loop? (Very little — that is the point.)
Check yourself#
- What are the two possible kinds of model response in the loop?
- Why are tool errors returned to the model instead of aborting?
- Name four limits a production loop must enforce.