Skip to main content

Command Palette

Search for a command to run...

Day 01: Why your first agent should be 40 lines, not a framework

Updated
•9 min read•View as Markdown
Day 01: Why your first agent should be 40 lines, not a framework

The first agent I ever got working, I did not understand. I had installed a framework, pasted eleven lines out of its quickstart, and watched a correct answer appear. Then I changed the prompt slightly and it printed nothing. Not an error. An empty string, exit code zero.

I spent the next forty minutes in a debugger. Nine frames down, past a callback manager and two abstract base classes, I found what I had been hunting: a dictionary getting built, and an HTTP POST. That was it. That was the whole thing the framework had been doing for me.

The bug was not what bothered me. What bothered me was that I could not have guessed where to look, because I had never once seen the request. I had learned a vocabulary — chains, runnables, callbacks — sitting on top of a fact I had never been shown.

So the first working version of cartographer, the codebase-mapping agent I am building over thirty days, has nothing in the way. One file, one dependency, forty lines.


The Short Version

Code for this day: cartographer at day-01

  • An agent is an HTTP POST in a loop. Everything else is retries, ergonomics and scheduling.

  • Frameworks abstract the part that was already easy and hide the part you need visible when it breaks.

  • Day one is forty lines: read a file, post it, print the reply. httpx is the only import that is not standard library.

  • Call raise_for_status() before you touch the body, or every auth and rate-limit failure reaches you as KeyError: 'choices'.

  • The reason to start raw is not purity. It is that you will end up reading a framework's source anyway, and this is the map you read it with.

  • The honest ending: it works, and the output is prose nobody can check. That, not the absent framework, is the real problem.


An agent is a POST. That is the entire secret.

Strip every layer off and what reaches the network is one request with two fields that matter: a model id, and a list of messages, each one a role (system, user or assistant) plus its text. What comes back is JSON, and the part you want is buried at a fixed address inside it. Every model provider worth using speaks this shape, which is why one endpoint can front hundreds of models.

That is not a simplification for teaching purposes. It is the actual protocol. The rest of this series — tools, protocols, graphs, memory, retrieval — is thirty days of arguing about which strings go into that array and in what order. If you do not have the array clearly in your head, none of the later arguments have anywhere to land.

The forty lines

I am running everything in this series through OpenRouter, one endpoint at https://openrouter.ai/api/v1 speaking the OpenAI wire protocol, on free models — so anyone can reproduce the whole thing for nothing. Nothing here is specific to that choice; change the URL and the model string and it works elsewhere.

import os
import pathlib
import sys

import httpx

OPENROUTER_URL = "https://openrouter.ai/api/v1/chat/completions"
MODEL = os.environ.get("CARTOGRAPHER_MODEL", "nvidia/nemotron-3-super-120b-a12b:free")
TARGET = pathlib.Path("src/cartographer/agent.py")

PROMPT = """You are reading one file from a repository you have never seen.

Say what this file is for, what it depends on, and what would break if it
were deleted. Where the file does not say, answer that you cannot tell.

--- {path} ---
{source}
"""

def ask(prompt: str) -> str:
    response = httpx.post(
        OPENROUTER_URL,
        headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
        json={"model": MODEL, "messages": [{"role": "user", "content": prompt}]},
        timeout=120.0,
    )
    response.raise_for_status()
    body = response.json()
    return str(body["choices"][0]["message"]["content"])

def main() -> None:
    if not TARGET.exists():
        sys.exit(f"cannot find {TARGET} from {pathlib.Path.cwd()}")
    print(ask(PROMPT.format(path=TARGET, source=TARGET.read_text())))

if __name__ == "__main__":
    main()

That is all of src/cartographer/agent.py, the one file today adds, and it reads top to bottom in four parts.

The constants say where and what. OPENROUTER_URL is the chat completions endpoint, the address that takes a list of messages and returns a reply. MODEL reads the model id from an environment variable, with a default. TARGET is the one file to read.

PROMPT is a template. {path} and {source} are placeholders that .format() fills with the file's name and contents. Its instructions ask three concrete questions and, deliberately, permit "I cannot tell", because a model that is not allowed to say that fills the gap with something plausible.

ask is the model call. It posts a JSON body with exactly two fields, the model id and a one-message list, with the key in an Authorization header and a two-minute timeout, because free models are slow. It checks the status, parses the reply, and digs the text out of choices[0].message.content. The str() around that lookup is for the type checker, which will not accept an untyped value from a function promising a string.

main checks the file exists, fills the template, asks, and prints. The if __name__ guard runs main only when the file is executed as a program, so importing it from a test sends nothing.

The second new file, tests/test_agent.py, checks without a network that the filled prompt contains the file's text and the permission to be uncertain. Today's dev notes explain every line, including the commands to run it.

The target is hardcoded to the file you are looking at, so the agent's first job is to describe itself. That is a deliberately easy first test: you already know the right answer, so you can judge the output the moment it appears.

The model id comes from the environment with a default rather than being pinned in source, because free-tier availability churns constantly. When one model is busy you swap it in your shell, not in a commit.

Three failures a framework would have swallowed

No key. os.environ['OPENROUTER_API_KEY'] raises KeyError at the exact line that needed it. Subscript, not .get(). A missing credential should not be allowed to become an empty string that travels three functions before failing.

A bad key, or a busy model. Without raise_for_status(), a 401 or a 429 gives you a perfectly valid JSON body containing an error object and no choices — so the failure surfaces as KeyError: 'choices' from the last line of ask, pointing at the one place where nothing is wrong. One method call converts that into an exception naming the status code.

An empty answer. choices can come back as an empty list. Under raw HTTP you see that immediately, because you are the one doing the indexing. Under a framework you get None from a helper and start wondering about your prompt.

None of these are exotic. They are the first three things that happen to everybody, and all three are cheap to diagnose only because there is no layer between the traceback and the wire.

So when is a framework right?

Not never. It is right at the point where you are paying for something you have already built badly.

Raw HTTP Provider SDK Agent framework
You own the request body the retry and typing layer almost nothing
You see when it breaks status code and body a typed exception a stack you have not read
Costs you boilerplate per feature portability across providers the shape of your program
Right when learning, or the loop is the product one provider, production traffic you need durability, branching and resumption, and know why

The third column becomes correct later in this series, when the agent has to survive a crash halfway through a scan and resume where it stopped. Writing that yourself is a bad use of a month. But arriving there having never seen the request means you cannot debug it, and you will need to.

It works, and I cannot tell if it is right

Running it prints a fluent paragraph about agent.py. It reads well. It is also completely unverifiable: there is no field to assert on, no count to compare, nothing a test can hold. Prose is unfalsifiable by construction, and a mapping tool whose output cannot be checked is a tool nobody should trust with a codebase they do not know.

That gap is the honest state of the project tonight, and it is a more interesting problem than any missing framework. A framework would have produced the same unverifiable paragraph, faster, from behind a wall.

What I'd do

  1. Write the POST by hand once, before installing anything. Half an hour, and it never stops paying.

  2. Put raise_for_status() in from the start. It is the difference between a diagnosis and a hunt.

  3. Read the key with a subscript, not .get(). Fail at the boundary.

  4. Point the first version at a file you already understand, so bad output is obvious rather than plausible.

  5. Adopt the framework the day you can name the feature you need from it — resumption, branching, durability — and not one day earlier.


If this was useful, a clap helps other people find it. I want to hear the other side too: what was the first thing that made you reach for an agent framework, and was it something you had already tried to build yourself?


All views in this post are my own. The scenarios are illustrative composites, and code examples are not drawn from any specific codebase.