Course 0 — What an AI agent actually is (and how the local-model side works)
Read this before Course 1. You don't need to install anything yet. This is the mental model that makes the install courses make sense.
By the end of this course, you will be able to:
There is no install in this course. Read, think, ask questions.
A chatbot is a thing you ask questions of.
You write a question. It writes a reply. You write another. It writes another. Every reply is generated fresh from the conversation history. When you close the tab, the chatbot forgets you.
Examples: ChatGPT on the web, Gemini in a browser, Claude.ai in a browser.
A chatbot is a typewriter. You type, it types back. Nothing else happens.
An agent is a thing that does things.
You give it a goal. It figures out the steps. It calls tools, runs code, sends messages, opens files, makes changes. It remembers what it did last week. It can run on a schedule without you asking.
Examples: Hermes running on your laptop. The Afterschool Agent we're building in this course.
An agent is a worker. You give it a job, and it works on that job — across days, across sessions, across the file system on your computer.
A chatbot answers. An agent does.
The programmes you know — ChatGPT, Gemini, Claude — are chatbots. They've made AI easier to talk to, but they don't know anything about your life, they can't run on your laptop, and they can't do anything in the world.
The Afterschool Agent is built on Hermes, which is an agent. It knows your name, your interests, your challenges. It suggests readings. It sends a Morning Minute at 6 AM. It drafts a study plan when you ask. It runs on your laptop, so it works even when your Wi-Fi is down.
That's the difference. Everything else in this course is about installing that, connecting it, and using it.
Every AI agent has four parts. Hermes is no exception.
The model is the part that actually "thinks." It's a large mathematical file trained on a big chunk of the internet. You ask it a question, it generates an answer.
Hermes can talk to many models:
Course 2 covers LM Studio (the local side). Course 3 covers OpenCode Go (the cloud side).
Memory is what the agent remembers across sessions.
When you chat with a chatbot, the conversation ends when you close the tab. When you chat with an agent on your laptop, the agent saves:
This is what makes the agent feel like it knows you. Cloud chatbots don't have this — they treat every conversation as a stranger.
Tools are the agent's hands.
A model can only generate text. It can't send a message, search the web, open a file, or run a calculation. Tools let the agent do those things.
Hermes ships with tools for:
When the agent decides it needs a tool, it stops, calls the tool, and then resumes the conversation. This is the "loop" the next lesson is about.
The schedule is the agent's calendar.
An agent doesn't just react when you talk to it. It can also act on its own at times you set:
This is the "cron jobs" piece. The agent runs like a quiet employee, even when you're not looking at it.
This is what makes a chatbot into an agent.
You write. The model writes. That's it.
you → model → you
You write a goal. The agent:
you → agent → model → decision
↙ ↘
tool call reply
↓
tool result
↘
model (with new info)
↘
reply
You don't see the loop. You see a chat window. But inside, the agent is iterating: read, think, act, look, repeat.
You can ask Hermes things like:
For each of those, the agent will likely loop through 2-5 tool calls before it answers you. The reply you see is the end of a short story.
Most AI runs in the cloud. ChatGPT, Claude, Gemini — they all live in data centres. Your question gets sent to a server somewhere, the server runs the model, and the reply comes back.
That's fine for casual use. It's a problem for anything personal.
When you send a question to a cloud model:
For a chat about which movie to watch, who cares.
For a chat about your sleep, your anxiety, your faith, your kids — you might care.
A local model runs on your laptop. The model file is on your hard drive. Your prompt never leaves your machine. The reply is generated on your machine. Nothing travels.
This is the same privacy as writing in a notebook instead of sending a text.
Local models are smaller than cloud models. They can't know as much. They can't reason as deeply. But for everyday conversation, planning, journaling, and the eight essentials — they're enough.
| Cloud model | Local model | |
|---|---|---|
| Smartness | High | Medium |
| Cost per query | $0.0001–0.01 | Free |
| Privacy | Your data travels | Your data stays |
| Internet needed | Yes | No |
| Speed | Instant (usually) | Slower (small models) |
| Battery/CPU | None on your end | Uses your laptop |
Hermes uses both. By default, simple tasks go to the local model. Hard questions (anything that needs deep reasoning) get routed to a cloud model. The agent decides for you.
Local models need a way to run on your laptop. Two popular tools do this.
LM Studio is a desktop app. You download it once, like any other app. Inside, you browse a catalogue of models, click one to download, and chat with it from inside the app.
It also has a "Server" mode: when you turn it on, your laptop becomes a small HTTP server that other apps (like Hermes) can talk to. That's how Hermes uses LM Studio.
When to use LM Studio:
Ollama is a command-line tool (no graphical interface). You install it once, then you type ollama run lfm2-700m and it runs.
It's faster to launch, uses less memory, and is preferred by people who like the terminal. It also exposes a server on the same port as LM Studio, so Hermes can talk to either.
When to use Ollama:
Course 2 walks you through LM Studio because:
You can switch to Ollama later if you want. Hermes treats both as the same thing — they'll show up as different "local model providers" in the agent's settings.
The model is the actual "brain" of the agent. Different models have different sizes, speeds, and capabilities.
Model files are measured in parameters — billions of values the model learned during training. Bigger numbers mean smarter models, but also slower and more memory-hungry.
| Model | Size | Disk | RAM | Speed | Smartness |
|---|---|---|---|---|---|
| LFM2-700M | 0.7B | 1.4 GB | 2 GB | Fast | OK for chat, planning, rituals |
| Qwen 1.5B | 1.5B | 2.5 GB | 4 GB | Fast | Better reasoning |
| Llama 3.1 8B | 8B | 4.7 GB | 8 GB | Medium | Quite smart |
| Mistral 22B | 22B | 13 GB | 24 GB | Slow | Very smart |
For this course, use LFM2-700M. It's:
When you outgrow it, swap it for a bigger model. But don't rush to a bigger model — the small one is more private, faster, and uses less battery.
Bigger models are smarter but:
You can switch models anytime. Start small. Graduate when you need to.
Hermes sits between you and the model. When you ask a question, Hermes:
You don't see any of this. You see a chat. The agent does the rest.
You: "Plan tomorrow for me."
What happens:
All of that takes 2-4 seconds.
You: "Read my notes from yesterday and tell me what I should review for tomorrow's exam."
What happens:
This time, the loop has 2-3 tool calls. Longer. But still invisible to you.
We touched on tools. Now the schedule.
A cron job is a task that runs on a schedule. Not in response to a request — just on time.
Examples from your everyday life:
The agent's cron jobs:
Each one is a separate task. Each one runs without you. The agent does the work, then sends you the result.
You might think "that's just a reminder." It's not. A reminder is something you set. A cron job is a task that the agent does for you, then notifies you when it's done.
The difference:
The agent does the work. The notification is just the delivery.
The schedule is the fourth part. Without it, the agent only responds when you talk to it. With it, the agent has a daily rhythm. It becomes a quiet presence in the teen's life.
Here's the full picture.
┌─────────────────────────────────────┐
│ HERMES AGENT │
│ │
│ ┌─────────────┐ ┌───────────────┐ │
you talk to it ───────► │ the model │ │ memory │ │
│ │ (local or │ │ (your name, │ │
it sends you ────────── │ cloud) │ │ interests, │ │
▲ │ │ │ │ history) │ │
│ │ └─────────────┘ └───────────────┘ │
│ reply │ │
│ │ ┌─────────────┐ ┌───────────────┐ │
│ │ │ tools │ │ schedule │ │
│ │ │ (files, web, │ │ (cron jobs │ │
│ │ │ messages) │ │ at fixed │ │
│ │ │ │ │ times) │ │
│ │ └─────────────┘ └───────────────┘ │
└─────────│─────────────────│────────────┘
↓ ↓
local model Telegram /
(your laptop) WhatsApp
After Course 4, the agent is yours. It runs on your laptop. It knows your essential. It sends you a Morning Minute every morning. It plans your afternoon. It talks to you about the eight essentials until you tell it to stop.
That's the goal. This course just made sure you understand the goal.
You now have the mental model. Courses 1-4 are about installing the parts.
When you finish this course, take a break. Come back when you're ready to install.
Next course: Course 1 — Download and install Hermes Desktop.
| Term | Meaning |
|---|---|
| Agent | An AI that does things, not just talks. Has tools, memory, and a schedule. |
| Chatbot | An AI that only talks. No persistent memory, no tools, no schedule. |
| Model | The mathematical file that generates text. Local or cloud. |
| Local model | A model that runs on your laptop. |
| Cloud model | A model that runs on someone else's server. |
| LM Studio | A desktop app for downloading and running local models. |
| Ollama | A command-line tool for running local models. |
| Tool | Something the agent can do besides talk — read a file, send a message, run code. |
| Cron job | A task that runs on a schedule, not in response to a request. |
| Memory | What the agent remembers across sessions. |
| OpenCode Go | A paid gateway that lets Hermes use big cloud models for a few cents per query. |