# Day 00a: Six ideas that explain every AI agent, before any code

*Somewhere in the first hour of reading about AI agents, you meet a diagram. A box labelled Thought, an arrow to Action, an arrow to Observation, and an arrow curling back to the start. The captions say the agent reasons, plans and decides. It makes building one sound like work for a research lab.*

*Then someone asks a plain question: when the agent "thinks", what does the model actually receive? The honest answer is a list of messages, written as JSON, sent in an ordinary web request. The reply is more JSON. Nothing in between remembers anything.*

*That gap, between the vocabulary and the mechanism, is what this primer closes: six ideas, explained from zero, each with something you can see for yourself.*

* * *

## The Short Version

*   **A model call is text in, text out**, over HTTP. It predicts likely text, which is why it can be fluently wrong.
    
*   Models read and are priced in **tokens**, chunks often smaller than a word.
    
*   **The context window** is the most one call can carry. Anything outside it does not exist for the model.
    
*   **The model remembers nothing between calls.** Your code resends the conversation.
    
*   **Tool calling does not let the model act.** It lets the model ask your code to act.
    
*   **An agent is a loop** around a model call: ask, run any requested tool, send the result back, repeat until it answers. Your code owns the limits.
    

* * *

## Before you start

**What you will learn.** The six ideas, tried for real where possible, and the vocabulary for the next thirty days.

**Prerequisites.** None to read it. You have used a chat assistant, and you can read a little Python. The optional "Try it yourself" boxes need a macOS or Linux terminal (or WSL on Windows) with `curl` and `python3`, and a free [OpenRouter](https://openrouter.ai) account: one web address that forwards requests to hundreds of models, some of them free.

**Time.** About twenty-five minutes to read, fifteen more for the experiments.

**How to read it.** Each idea has the same parts: plain words, an everyday picture, what goes over the wire, the common mistake, why it matters. At "Predict, then check", decide before opening the reveal.

### Get a free OpenRouter key

The experiments need an **API key**, a secret string that tells OpenRouter which account is asking. Getting one takes two minutes and no payment.

1.  Go to [openrouter.ai](https://openrouter.ai) and sign up, with an email address or one of the sign-in buttons it offers.
    
2.  Open [openrouter.ai/keys](https://openrouter.ai/keys) and click **Create Key**.
    
3.  Name it something you will recognise, such as `cartographer`. Leave the credit limit empty; `:free` models cost nothing.
    
4.  Copy the key at once. It is shown only one time.
    
5.  Free models need no payment, but they are rate-limited: a busy one may ask you to wait.
    
6.  Never paste the key into a file, a chat or a screenshot. Load it straight into your terminal instead:
    

```bash
read -rs "?OpenRouter key: " k && export OPENROUTER_API_KEY="$k" && unset k
```

Paste the key when asked; nothing shows as you paste. That is zsh, the macOS default; in bash, start with `read -rsp "OpenRouter key: " k`. The key now sits in an **environment variable**, a named value your terminal hands to the commands it runs, never printed and never saved. It lasts until you close that terminal, and Day 01 loads it the same way.

If the key ever leaks, delete it on the Keys page and create a new one. The old one stops working the moment you delete it.

## The picture the word suggests

The word "agent" suggests an assistant that remembers what you told it, looks at your files, and goes off and does things. Each part of that picture is wrong in a way that decides how you build one. It does not remember: your code resends everything. It cannot see your files, only the text you send. It cannot act: it can only ask, and your code acts. The six ideas below replace the picture one piece at a time.

## Idea 1: A model call is text in, text out

**In plain words.** Your code sends an HTTP request to a **provider**, the service that runs the model, holding a model name and a list of messages. The useful part of the response is one field: the generated text. The model does not browse, run code or open files. It returns the text it predicts should come next, one small piece at a time.

**An everyday picture.** The word suggestions above your phone keyboard: the same job, guessing what comes next, at an enormously larger scale.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/11be3f25-f95a-4fd0-a21d-da25feae9ae2.png align="center")

**On the wire.** A complete request. Each message has a **role** saying who wrote it: `user` for you, `assistant` for the model, `system` for standing instructions at the top.

```json
{
  "model": "nvidia/nemotron-3-super-120b-a12b:free",
  "messages": [
    {"role": "user", "content": "Name one Python web framework, in one word."}
  ]
}
```

The response, trimmed to the fields that matter. The shape is real; your wording and numbers will differ.

```json
{
  "choices": [
    {
      "message": {"role": "assistant", "content": "Flask"},
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": <number>,
    "completion_tokens": <number>,
    "total_tokens": <number>
  }
}
```

Three fields carry everything. `choices[0].message.content` is the answer: the first entry of `choices`, its `message`, its `content`. `finish_reason` says why it stopped: `stop` means finished, `length` means cut off at a size limit. `usage` counts tokens, Idea 2.

> **Try it yourself.** One real model call. Paste it into the terminal where you loaded your key.

```bash
curl -s https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "nvidia/nemotron-3-super-120b-a12b:free", "messages": [{"role": "user", "content": "Name one Python web framework, in one word."}]}' \
  | python3 -m json.tool
```

| Part | What it does |
| --- | --- |
| `curl -s <url>` | Sends an HTTP request to the endpoint, quietly |
| `-H "Authorization: Bearer ..."` | A header carrying your key, read from the environment variable |
| `-H "Content-Type: application/json"` | Says the body is JSON |
| `-d '{...}'` | The body: the model name and the messages |
| `python3 -m json.tool` | Prints the reply neatly; it changes nothing |

**Predict, then check.** How much of what comes back will be the answer itself?

Check your prediction

One word, inside a page of bookkeeping: an id, the model name, the finish reason, the token counts.

**What you should see.** After a few seconds, JSON with a framework name in `content`, `finish_reason` set to `stop`, and three counts under `usage`. A `reasoning` field, if present, is not the answer. If you see an error instead:

| You see | It means |
| --- | --- |
| `401` | The key is missing or wrong in this terminal; run the `read` line again |
| `429` | The free model is busy; wait a little and retry |
| The model id is unknown or unavailable | Free models come and go; swap in any id ending in `:free` from [openrouter.ai/models](https://openrouter.ai/models) |

> **Common mistake.** Treating a fluent answer as a checked one. The model produces what is *likely*, not what is *verified*, so when it does not know, it still produces something likely. A confident, plausible, wrong answer is **hallucination**. No prompt removes it entirely. **Grounding**, making the model answer from text you gave it, reduces it.

**Why it matters for agents.** Everything an agent does is built from this call, and grounding is why cartographer, the agent this series builds, only cites code it has read off disk.

## Idea 2: Tokens are the unit everything is measured in

**In plain words.** Models do not read characters or words. Text is split into **tokens**: often a whole short word, sometimes a fragment of a long one, sometimes one punctuation mark. The model reads and writes tokens, and paid providers charge per token both ways. The splitter is the model's **tokenizer**, and different models split the same text differently.

**An everyday picture.** The data on a phone plan: what you pay for, and what runs out. Tokens are both the price and the limit.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/38c3d5f2-ba36-43ce-9e8a-df76042118a7.png align="center")

**On the wire.** You send text; the provider reports the cost in `usage`: `prompt_tokens` for what you sent, `completion_tokens` for what the model wrote, `total_tokens` for the sum.

**A worked example by hand.** A tokenizer might keep "the" whole and cut "unbelievably" into pieces like "un", "believ" and "ably"; where it cuts depends on the model. So eight words need not be eight tokens, and code, with its brackets and indentation, usually costs more than its word count suggests.

**Predict, then check.** Your question in the first experiment had eight words. Will `prompt_tokens` say 8?

Check your prediction

Very likely not. Words and tokens do not line up one to one, and the provider also counts the formatting that wraps your message. Run the experiment again and read the real number.

To watch a tokenizer cut text live, try [Tiktokenizer](https://tiktokenizer.vercel.app/); its counts are for particular models, so they illustrate, not predict.

> **Common mistake.** Estimating cost or size in words. Providers bill and limit in tokens, and only `usage` tells you the true count.

**Why it matters for agents.** An agent makes many calls, each billed and limited in tokens. Day 03 measures a real repository instead of guessing.

## Idea 3: The context window is one page

**In plain words.** Every model has a **context window**: the most tokens one request can hold, your messages and the reply together. It is not memory and not storage. It is the size of the one page the model sees, and anything off that page does not exist for that call. Sizes differ by model, so check yours.

**An everyday picture.** An exam where you bring one sheet of notes and write your answer on the same sheet.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/a0272d9c-12ae-4895-b30f-6cd3f774d5e2.png align="center")

**A worked example by hand.** Toy numbers, chosen for easy arithmetic, not taken from any real model. Pretend the window is 100 tokens.

| What is on the page | Tokens | Running total |
| --- | --- | --- |
| Instructions (the system message) | 15 | 15 |
| One source file | 60 | 75 |
| Your question | 10 | 85 |
| Room left for the reply | 15 | 100 |

Add a second 60-token file and the input alone is 145: it does not fit, and the request fails with an error that no rewording fixes. Keep one file but ask for a long explanation, and the reply runs out of room after 15 tokens, stopping mid-sentence with `finish_reason` set to `length`.

> **Common mistake.** Thinking a better prompt fixes a full window. The problem is arithmetic, so the fix is sending less: a slice instead of the whole file, a summary, only the files that matter. Chat apps may quietly drop your oldest messages to make room, which is why a very long chat seems to forget its start.

**Why it matters for agents.** A real codebase is far bigger than any page, so an agent that reads code must choose what goes on it. That choice is most of what this series builds.

## Idea 4: The model remembers nothing

**In plain words.** Send two requests in a row and the second knows nothing about the first; the endpoint keeps no conversation for you. A chat feels continuous because your code keeps the conversation as a list and resends all of it, plus the new message, every time. The model is not remembering. It is rereading. This is what **stateless** means.

**An everyday picture.** A help line where a different person answers every call and nobody keeps notes. To continue, you read the whole conversation out again, every time.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/87ddc3fc-479f-42ef-a347-d75a8e8229ec.png align="center")

**Predict, then check.** Three separate requests: the first tells the model a name, the second asks for it back, the third sends the whole conversation in one list. What will the second answer?

> **Try it yourself: prove it forgets.** Three requests, each printing only the answer.

```bash
curl -s https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nvidia/nemotron-3-super-120b-a12b:free", "messages": [{"role": "user", "content": "My name is Asha. Reply only with OK."}]}' \
  | python3 -c "import sys, json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
```

```bash
curl -s https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nvidia/nemotron-3-super-120b-a12b:free", "messages": [{"role": "user", "content": "What is my name?"}]}' \
  | python3 -c "import sys, json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
```

```bash
curl -s https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nvidia/nemotron-3-super-120b-a12b:free", "messages": [{"role": "user", "content": "My name is Asha. Reply only with OK."}, {"role": "assistant", "content": "OK"}, {"role": "user", "content": "What is my name?"}]}' \
  | python3 -c "import sys, json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
```

What you should see

The first replies OK. The second says some version of "I don't know your name". The third says Asha. Wording varies; the pattern does not. Nothing changed between the second and third except what you sent. If the second names Asha, it guessed: rerun it. If a command prints a KeyError, the request failed: swap the last line for python3 -m json.tool to see the error.

**A worked example by hand.** Toy numbers again: every question is 10 tokens, every answer 20.

| Turn | What your code sends | Input tokens |
| --- | --- | --- |
| 1 | Q1 | 10 |
| 2 | Q1, A1, Q2 | 40 |
| 3 | Q1, A1, Q2, A2, Q3 | 70 |

You typed 30 tokens of questions but sent 120 tokens of input, and the gap widens every turn.

> **Common mistake.** "But my chat app remembers me." The app does, not the model: it saves notes and puts them back into the request. Some providers offer APIs that store a conversation for you, but the model still receives all of it on every call.

**Why it matters for agents.** An agent's working memory is a list your code owns: what goes back in, what gets cut, what it costs. Day 02 builds it.

## Idea 5: Tool calling lets the model ask, not act

**In plain words.** If the model only produces text, how does an agent read a file? In the request, you describe functions it may ask for, each a **tool** with a name, a description and its arguments (its **tool schema**). The model can then reply with a **tool call**, a structured request to run one, instead of an answer. Your code runs it and sends the result back as a new message. The model asked; your program acted.

**An everyday picture.** A doctor on the phone says "take your temperature and tell me the number". You hold the thermometer and report back. The doctor only asked.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/1329842e-9598-48e7-b69e-d04ed83b893b.png align="center")

**On the wire.** The request gains a `tools` list. One entry, in the format Day 06 uses:

```json
{
  "type": "function",
  "function": {
    "name": "read_file",
    "description": "Read one file from the repository and return its text.",
    "parameters": {
      "type": "object",
      "properties": {
        "path": {
          "type": "string",
          "description": "Path relative to the repository root."
        }
      },
      "required": ["path"]
    }
  }
}
```

When the model wants it, `content` is `null`, `finish_reason` is `tool_calls`, and the message carries the request:

```json
{
  "role": "assistant",
  "content": null,
  "tool_calls": [
    {
      "id": "call_1",
      "type": "function",
      "function": {
        "name": "read_file",
        "arguments": "{\"path\": \"src/app.py\"}"
      }
    }
  ]
}
```

Your code appends that message, runs `read_file`, and appends the result with the role `tool`, tied to the request by its id:

```json
{"role": "tool", "tool_call_id": "call_1", "content": "<the text of src/app.py>"}
```

Then it sends everything again. The id here is made up.

**Predict, then check.** Look at `arguments`. Is it a JSON object?

Check your prediction

No. It is a string containing JSON, written by the model like any other text. Your code must parse it and expect it to be malformed, name a missing file, or ask for a tool you never described.

> **Common mistake.** Saying the model "ran" a tool. It can only request what you described; nothing runs unless your code runs it. One reply can request several tools, and each gets its own result message.

**Why it matters for agents.** This is where the control sits: the tools you describe are all the agent can do, and checking each request is your job. Day 06 makes this round trip by hand.

## Idea 6: An agent is a loop

**In plain words.** Put the last two ideas together and you have an agent:

1.  Send the conversation and the list of available tools.
    
2.  If the reply is an answer, stop.
    
3.  If the reply is a tool request, run it, append the result to the conversation, and go back to step 1.
    

That is the whole of it. The model chooses the next step. Your code decides the rest: which tools exist, how many times the loop may run, what happens at that limit. Each pass is a **turn**.

![](https://cdn.hashnode.com/uploads/covers/69eb0bc51e45c4e0da9b9f00/54d1e8db-1153-48ca-bef7-23dadf2dbf77.png align="center")

**An everyday picture.** The phone doctor asks for your temperature, then about your throat, then diagnoses. You decide how many tests you will do.

**On the wire: one run, traced by hand.** An agent with `list_files` and `read_file` tools is asked for a project's entry point.

| Turn | Your code sends | The model replies | Your code then |
| --- | --- | --- | --- |
| 1 | system, your task, the tools | tool call: `list_files` | runs it, appends the file list |
| 2 | all of the above, plus that result | tool call: `read_file` on one file | runs it, appends the file's text |
| 3 | all of the above, plus that result | an answer, `finish_reason` `stop` | stops and prints it |

Every turn sends everything before it: Idea 4, inside the loop.

**Predict, then check.** What happens if the model never replies with an answer, and keeps asking for one more file?

Check your prediction

Nothing in the model stops it. The loop runs, spending tokens, until your code stops it, which is why every real loop has a ceiling, such as a maximum number of turns.

> **Common mistake.** Reading "the agent decided" as something beyond this loop. The model returned a tool request and the loop ran it. Thought, Action, Observation is this loop drawn with bigger words.

**Why it matters for agents.** A loop the model can keep alive needs a limit you control. Day 06 builds the loop with two such limits.

## What an agent is not

*   **Not a chatbot.** A chat window is an interface. Some chat apps run a tool loop behind it, and the loop is what makes the agent.
    
*   **Not a workflow.** A workflow is a fixed sequence your code decides in advance, and it is often the better choice.
    
*   **Not a smarter model.** It is the same model, with a loop and tools around it.
    

The question that sorts them is who decides the next step:

|  | Who decides the next step | Good for | Breaks when |
| --- | --- | --- | --- |
| Chat model | You, every turn | A single question | The task needs many steps |
| Workflow | Your code, fixed in advance | Predictable, repeatable jobs | The path depends on what it finds |
| Agent | The model, inside a loop you bound | The next step depends on the last result | You cannot afford it to explore |

Reading an unfamiliar codebase is the third kind: you cannot know the second file to open until you have read the first.

## Words you will meet in the next thirty days

Every later post defines its terms again, in this wording.

| Term | What it means | Built on |
| --- | --- | --- |
| Model | The program that predicts text; you reach it over the network | Day 01 |
| Provider | The service that runs models and answers your requests | Day 01 |
| Endpoint | The web address a request is sent to | Day 01 |
| API key | A secret string that identifies your account; never write it in a file | Day 01 |
| Message | One entry in the conversation: a role plus its text | Day 01 |
| Role | Who wrote a message: `system`, `user`, `assistant` or `tool` | Day 01 |
| Prompt | The text you send for the model to respond to | Day 01 |
| System prompt | Standing instructions in the `system` message at the top | Day 01 |
| Hallucination | A fluent answer that is not backed by anything the model was given | Day 01 |
| Grounding | Making the model answer from text you supplied | Day 01 |
| Transcript | The conversation list your code keeps and resends | Day 02 |
| Stateless | Keeps nothing between calls | Day 02 |
| Token | The chunk of text a model reads, writes and is billed in | Day 03 |
| Context window | The most tokens one request and its reply can hold | Day 03 |
| Structured output | Making the model reply in a fixed JSON shape you can check | Day 04 |
| JSON Schema | A standard way to describe the shape a piece of JSON must have | Day 04 |
| Temperature | A setting that controls how much randomness goes into picking each token | Day 05 |
| Tool | A function you describe to the model by name, description and arguments | Day 06 |
| Tool call | The model's structured request to run one of your tools | Day 06 |
| Turn | One pass of the agent loop: send, reply, act | Day 06 |
| MCP | Model Context Protocol, a standard way to offer tools to many agents | Day 09 |
| Graph | An agent written as named steps and the edges between them | Day 15 |
| State | Everything the agent has accumulated so far, in one place | Day 16 |
| Checkpoint | A saved copy of state, so a run can resume after a crash | Day 19 |
| Human in the loop | Pausing the agent for a person to approve a step | Day 20 |
| Memory | What is kept beyond one conversation, and how it is put back in | Day 22 |
| Retrieval | Fetching only the relevant text, so it fits on the page | Day 23 |
| Eval | A repeatable test of the agent's answers | Day 25 |
| Trace | A record of every call and tool run in one agent run | Day 26 |
| Guardrail | A limit on what the agent is allowed to do | Day 29 |

And a translation for what you will read elsewhere:

| When someone says | It means |
| --- | --- |
| "The agent thought about it" | The model generated text in reply to the conversation |
| "The agent decided to read a file" | The model returned a tool call, and the loop ran it |
| "The agent remembers" | Your code resent earlier messages, or put saved notes back into the request |
| "The agent forgot" | That text was not in the request it was sent |
| "The agent used a tool" | Your code ran a function the model asked for |
| "The agent observed the result" | Your code appended the tool's output as a message |
| "The context is full" | The request plus the reply would exceed the window |
| "It hallucinated" | It produced likely text that nothing it was given supports |
| "Thought, Action, Observation" | Reply, tool call, tool result: the same loop |

## Where each idea comes back

| Idea | Where you build it |
| --- | --- |
| 1\. Text in, text out | Day 01, the raw call; Day 04, making the text checkable; Day 05, how each next piece is picked |
| 2\. Tokens | Day 03, measuring a real repository and a cost meter; Day 27, cost control |
| 3\. The context window | Day 03, the hard wall; Day 23, retrieval as a tool |
| 4\. The model remembers nothing | Day 02, the transcript; Day 22, memory |
| 5\. Tool calling | Day 06, by hand; Day 08, tool descriptions; Days 09 to 13, MCP |
| 6\. The agent loop | Day 06, the loop and its limits; Days 15 to 21, the loop as a graph |

## Check yourself

Answer each before opening it. If one will not come, reread that idea.

**1.** A model confidently describes a function that does not exist. Why, and what reduces it?

Answer

It produces likely text, not checked text (Idea 1). Give it the real source to answer from.

**2.** Your question is twelve words. Will prompt\_tokens be 12?

Answer

Almost certainly not. Tokens are not words, and formatting counts too (Idea 2).

**3.** A toy window holds 100 tokens. Instructions 15, a file 70, a question 10. How much room is left for the reply, and what if it needs more?

Answer

5 tokens. The reply is cut off with finish\_reason length (Idea 3). Send less.

**4.** Turn two of a chat clearly "remembers" turn one. Where does that memory live?

Answer

In your program's list, resent with every request (Idea 4).

**5.** Questions are 10 tokens, answers 20. How many input tokens does turn three send?

Answer

70: two questions, two answers, the new question (Idea 4).

**6.** An agent "read a file". Who opened it?

Answer

Your code, after the model returned a tool call (Idea 5).

**7.** What is in a tool call's arguments field, and why check it?

Answer

A string of JSON the model wrote; it can be malformed or name things that do not exist (Idea 5).

**8.** What stops an agent loop running forever, and who owns it?

Answer

A limit in your code, such as a maximum number of turns (Idea 6).

**9.** What one question separates a chatbot, a workflow and an agent?

Answer

Who decides the next step.

## Recap

1.  "The agent decided" means the model returned a tool request and the loop ran it.
    
2.  When something "forgets", look at the list it was sent.
    
3.  When something is too big, it is a context-window problem; no prompt fixes it.
    
4.  When an answer is confidently wrong, ask what text it was given.
    
5.  Trust the `usage` block, not an estimate.
    
6.  Your code owns the ceiling on every loop.
    

## What's next

Day 01 writes the first agent: the request from Idea 1, sent from about forty lines of Python, reading the same `content` field. [Today's dev notes](https://github.com/Harshavardhan17/cartographer/blob/main/docs/devnotes/day-00a.md) list what Day 01 needs on your machine.

## Further reading

*   [OpenRouter API reference](https://openrouter.ai/docs/api-reference/overview): the fields used above.
    
*   [OpenRouter tool calling](https://openrouter.ai/docs/guides/features/tool-calling): the full round trip.
    
*   [OpenRouter error codes](https://openrouter.ai/docs/api-reference/errors): every status code.
    
*   [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents): workflows versus agents.
    

* * *

*Which of these six surprised you, or which one did the explainers you have read get wrong? It tells me which idea needs more room in the posts that follow.*

* * *

*All views in this post are my own. The scenarios are illustrative composites, and code examples are not drawn from any specific codebase.*
