# Day 00: Learn AI agents by building one, in 30 days

*Someone hands you an empty file and says: build an agent. Perhaps you have read the explainers and can recite that an agent is "an LLM in a loop with tools". Perhaps you have read none of them, and the word only means software that does things on its own. Either way the cursor blinks, and you do not know what the first line should be.*

*Most people in that spot open a framework's quickstart. Three lines later something impressive happens, and the interesting part — the loop, the decision about what to send, the thing that goes wrong at scale — happened inside a function you never opened. That is not your fault. Almost everything written about agents demonstrates a library, so what it teaches is vocabulary rather than mechanism.*

*This is a thirty-day path to closing that gap. One concept a day, each one built by hand before any framework touches it, all of it assembling into a single real tool by the end. You will write the agent loop before you use a framework that hides it, and by the time you do adopt one you will be able to say what it is doing for you.*

* * *

## The Short Version

*   **Thirty days, one concept each**, from a bare model call to a production-shaped agent with evaluation and cost control.
    
*   Everything is built **by hand first**. Week one uses no SDK and no framework — forty lines of Python over raw HTTP.
    
*   It all assembles into one real tool: **cartographer**, an agent that reads a codebase you have never seen and tells you where to start.
    
*   Every day ships **build instructions you can follow**, and the code is public: [github.com/Harshavardhan17/cartographer](https://github.com/Harshavardhan17/cartographer).
    
*   Every model call runs on a **free tier**. Following along costs nothing.
    
*   Day 25 scores the finished agent against repositories with **known-good answers**. If it turns out to be bad at this, that is the post.
    

* * *

## Why build one thing instead of thirty demos

Most tutorials teach agent concepts with a fresh toy per concept: a weather tool for tool calling, a PDF chatbot for retrieval. Each one works, and none of them teach you much, because a toy cannot push back. You never hit the wall that makes the concept necessary — so you learn the API and not the reason.

So this series builds a single tool across all thirty days, chosen because it fights you.

**cartographer** points at a repository you have never seen and returns an onboarding map: where execution actually begins, which small fraction of the files carry the system, one data path traced end to end, and where your first change would go. Everyone who has joined a team knows the week this replaces.

It is a good teacher because of one property: **a repository does not fit**. A mid-sized open-source project runs to several million tokens, the chunks of text a model reads and is billed by. There is no context window, the most one request can hold, large enough, no prompt clever enough, and no bigger model next year that makes it go away. The agent must decide what to read.

That single constraint turns every concept in the curriculum from a demonstration into a necessity. Tool calling becomes the only way in. Retrieval stops being a stage you were told to add. Cost control becomes arithmetic. You meet each idea at the moment you need it, which is the only way any of it sticks.

It has a second useful property: **there is a right answer**. Point it at a codebase you know well and you can check the map, which is what makes a real evaluation possible on Day 25 instead of a vibe-based "seems to work".

## What you will build, week by week

**Week one: the model, with nothing in front of it.** The agent loop in about forty lines of Python over raw HTTP, with no SDK and no agent library. You will handle the field that separates a finished answer from a truncated one, build conversation memory by hand and see why "memory" is a misleading word, do the token arithmetic that governs every design decision later, make the model return typed records instead of prose, and hand-roll tool calling so you have seen exactly what crosses the wire. By Friday you have a working agent and a clear list of everything wrong with it.

**Week two: MCP.** The protocol that lets any editor or client talk to your agent instead of it living inside one script. This is a good moment to learn it, because MCP removed its initialization handshake and protocol-level sessions in the [2026-07-28 revision](https://modelcontextprotocol.io/specification/2026-07-28/changelog) — which means a large share of what is currently written about how MCP works describes a protocol that no longer exists. You will learn the version that is true today, build a server worth keeping, and handle the security question that arrives the moment an agent reads code it did not write.

**Week three: LangGraph, and why a loop is not enough.** Reading a codebase is not a pipeline. You read a file, fail to understand it, and go back for more context — that is a cycle, and a cycle is the thing a chain cannot express. You will also add the two features that turn a script into software: surviving a crash four thousand files into a scan, and stopping to ask a human before doing something consequential.

**Week four: the half that decides whether it is real.** Memory as three separate problems rather than one word. Retrieval as a tool the agent chooses to call. An evaluation set with known-good answers, so "did it get worse?" has an answer. Tracing, cost routing, and the seven ways this breaks in production: the part most tutorials skip.

## How to follow along

Each day has two pieces: a post explaining the concept and why it matters, and [dev notes](https://github.com/Harshavardhan17/cartographer/blob/main/docs/devnotes/day-00.md) that say which file to create in which folder, what to type into it, and what every line does.

Today's build has no agent code. It creates the repository, and every file in it has a job. `pyproject.toml` names the project, its Python version and its one runtime dependency. `src/cartographer/__init__.py` makes the folder an importable package and holds the version. `.gitignore` keeps the environment, caches and any secrets out of history. `LICENSE` makes the code reusable. `README.md` says what the tool will do and admits nothing works yet. `.github/workflows/ci.yml` lints, type-checks and tests every push, and `tests/test_package.py` is one real test. It also clones click 8.5.0, the repository the agent studies all series, pinned so every day measures the same code.

You need Python 3.12, a terminal, and about an hour a day. Every model call runs against free models through one OpenAI-compatible endpoint, so there is no bill.

You do not need prior agent experience. If agents are new to you, read the primer first, "Six ideas that explain every AI agent, before any code": about fifteen minutes, no code, covering what a model call is, tokens, the context window, why the model remembers nothing, tool calling, and the loop that makes it an agent. If they are familiar, start at Day 01. You do need to be comfortable reading Python.

## Two rules for the writing

**Nothing gets described that has not been built.** The code ships before the post does. When something does not work, that is the post, and there will be several, because thirty days of building something hard is not thirty days of wins.

**No invented numbers.** Any claim about accuracy comes from the Day 25 evaluation, against repositories with known-good answers, with the methodology published alongside it. Where a figure is illustrative, the sentence will say so.

Tomorrow: the forty-line agent, and why your first one should not be a framework.

* * *

*If you have lost a week to an unfamiliar codebase — and most of us have — which part actually cost you the time: finding the entry point, or working out which files mattered? The answer changes what gets built first, so I would genuinely like to know.*

* * *

*All views in this post are my own. The scenarios are illustrative composites, and code examples are not drawn from any specific codebase.*
