# Day 02: An LLM is a stateless function. Everything else is you

*I pointed the forty-line version of cartographer at a second file and asked the obvious follow-up: how does this one relate to the last one?*

*It asked me which file I meant.*

*My first instinct was that I had broken something. A flag I had dropped, a session id I had forgotten to thread through. I went looking for it for longer than I am comfortable admitting. There is no session id. There was never going to be one. The endpoint had no idea I had called it ninety seconds earlier, and no amount of configuration was going to give it one.*

*Every piece of memory you have ever seen in an AI product is somebody, somewhere, deciding what goes into an array before a POST. That is the entire mechanism. Once you have seen it you cannot unsee it, and a surprising amount of vocabulary stops being mysterious.*

* * *

## The Short Version

**Code for this day:** [cartographer at day-02](https://github.com/Harshavardhan17/cartographer/tree/day-02)

*   The model is **a pure function of the array you send**. It retains nothing between requests, and nothing you can configure changes that.
    
*   The forty-line version sent one message. Now it sends a list, and **the caller owns that list** — no library, no service, no hidden field.
    
*   Append the **assistant's reply as well as the user's**. Omitting it produces a subtler bug than amnesia.
    
*   Your agent now has **exactly one piece of state**, and it is a JSON file. Debug it with `cat` before you reach for anything else.
    
*   Re-sending a transcript is **quadratic**: ten equal turns bill you fifty-five turns' worth of input, not ten.
    
*   "Memory" is a product word. **Every implementation of it is a policy for what to put in the array**, and choosing that policy well is most of the rest of this series.
    

## The follow-up that goes nowhere

Here is the version I wrote first, and it is worth staring at because it looks correct.

```python
def turn(messages: list[dict[str, str]], content: str) -> str:  # Broken. Do not ship this.
    return ask(messages + [{"role": "user", "content": content}])
```

A list of messages goes in. A reply comes out. The signature reads like a conversation.

But `messages + [...]` builds a throwaway list, sends it, and drops it. The caller's list is never touched, so every call starts from the same place. This survives a read-through precisely because the type annotation implies history — you have to notice that nothing ever grows.

The endpoint this series uses is the OpenAI-shaped chat completions endpoint, and on that shape a request is self-contained. Whatever you did not put in the array did not happen.

## There is no session. There is a list.

The fix is three lines and it is the whole idea of the day.

```python
def turn(messages: list[dict[str, str]], content: str) -> str:
    messages.append({"role": "user", "content": content})
    reply = ask(messages)
    messages.append({"role": "assistant", "content": reply})
    return reply
```

It mutates the list it was handed. That is deliberate: there is one transcript, the caller holds it, and every append happens in one function you can put a breakpoint in.

The second append is the line people drop, and its failure mode is worse than no memory at all. Leave it out and the model receives a list of questions with no answers attached — a transcript in which it apparently never spoke. So it re-answers things it already answered, contradicts a conclusion it reached two turns ago, and generally behaves like a model that is bad at reasoning. You will spend an evening rewriting the system prompt, the standing instructions in the first message, before you find it.

Making both appends the responsibility of one function is not style. It removes the state where only half the exchange got recorded.

Around `turn`, the rest of `src/cartographer/agent.py` changes less than you would expect. Yesterday's template goes, replaced by `SYSTEM`, the system message that opens every list and tells the model it is mapping a repository one file at a time and should refer back to earlier files by name. `ask` keeps its body but now takes the whole list and sends it as the `messages` field, so it no longer decides what the model sees. `main` turns every path on the command line into one turn, then asks a closing question that only a model holding every earlier file can answer: which would you read first? `tests/test_agent.py` is replaced by a test that swaps `ask` for a fake and asserts both the order of roles and that the second call saw four messages. [Today's dev notes](https://github.com/Harshavardhan17/cartographer/blob/main/docs/devnotes/day-02.md) walk through every line.

## Your agent has exactly one piece of state

Because the model is a pure function of that array, the array is the complete state of your program. So write it down:

```python
TRANSCRIPT.write_text(json.dumps(messages, indent=2))
```

That line ends `main`, and `transcript.json` goes into `.gitignore`, because it holds whole source files and is rebuilt on every run. This is the highest-value five minutes of the day, and almost nobody's tutorial does it. When the output is wrong, the question you actually need answered is not *why did the model say that*. It is *what did the model see*. There is exactly one artifact that answers it, and now it is on disk.

Nearly every bad answer I have chased at this stage has been visible in that file within thirty seconds: a file that was never appended, a source read as an empty string because the path was relative to the wrong directory, a system message clobbered by a later assignment. None of those are model problems, and all of them look like model problems until you open the transcript.

It also makes runs replayable. The input is a JSON array; you can re-send it byte for byte, change one message, and send it again.

## The bill is a triangle, not a line

Now the part that gets expensive. Every call carries the whole array, so you pay for turn one again on turn two, and again on turn three.

If each turn adds roughly *T* tokens, the chunks of text a model reads and bills by, the *n*th request carries about *nT* input tokens, and the run as a whole costs *T·n(n+1)/2*. Ten turns is fifty-five turns' worth of input. Twenty turns is two hundred and ten. The curve is a triangle, and you are billed for its area.

For a chat assistant that is survivable, because a turn is a sentence. For an agent reading source files, each turn is an entire file, so *T* is enormous and the triangle gets frightening at single-digit *n*. That is a measurement problem before it is a design problem, which is what tomorrow is about.

There are three honest strategies, and they trade against each other:

|  | Send everything | Summarise older turns | Keep pointers, re-read on demand |
| --- | --- | --- | --- |
| **Input cost over n turns** | grows with the square | flat once the cap bites | flat, plus one call per lookup |
| **What you give up** | the context window, then your budget | detail you did not know mattered | round trips and latency |
| **Right when** | short runs, and while learning | chat, where old turns are chatty | the content is on disk and cheap to re-read |

For a codebase the third column is clearly right — the repository is still sitting there, and re-reading a file costs a disk read. That is the direction this project goes, and it is why the agent eventually has to be able to fetch things for itself rather than being handed them.

## What I'd do

1.  **Pass the transcript as an argument.** The moment it becomes a hidden attribute on an object, you stop being able to see what you sent.
    
2.  **Append both sides in one function.** Make the half-recorded exchange unrepresentable.
    
3.  **Dump the array to disk on every run**, from the first day. It is four lines and it is your only observability until much later.
    
4.  **Assert on the sequence of roles in a test.** Mine checks that two turns produce system, user, assistant, user, assistant — it would have caught the broken version above immediately.
    
5.  **Do not install a memory library** until you can name which column of that table you need. All three are a page of code; the hard part is the choice, and no library makes it for you.
    

* * *

*If this was useful, a clap helps other people find it. I am curious about the other side of this one: what is the longest-running agent conversation you have kept alive, and what did you eventually have to throw away to keep it affordable?*

* * *

*All views in this post are my own. The scenarios are illustrative composites, and code examples are not drawn from any specific codebase.*
