Professional IT Solutions

Hermes, Your Personal AI Assistant (Running on Free AI Models)

AI

Everybody talks about AI assistants, but most of what you see is a chat window that forgets you the moment you close the tab. I wanted something different: an assistant that runs on my side, remembers things, and writes to me first when there is something I should know. That is how I ended up with Hermes. And the part people don't believe when I tell them: I run it on free AI models.

What Hermes Is

Hermes Agent is an open-source AI agent made by Nous Research. It is not a chatbot inside a browser. You install it on a machine you control, connect it to an AI model, and then you talk to it from the terminal or from a messaging app like Telegram, Discord, Slack, WhatsApp, Signal or even plain email.

The important word is agent. Hermes doesn't only answer questions. It can run commands, search the web, keep notes about what it learned, and do things on a schedule without you asking every time. If you want to read more, the official documentation (opens in a new tab) is very good.

How I Set It Up

The installation is honestly the easy part. On Linux it is one command that downloads the installer and sets everything up. After that you tell Hermes which AI model to use, and this is where it gets interesting.

Hermes is not tied to one company. It works with a long list of providers, and two of them are perfect for a setup that costs nothing:

  • NVIDIA, which hosts a catalog of models you can call through an API key from their developer platform
  • OpenRouter, which is one account and one API key for models from many different labs, including a number of free ones

You create an account on each, generate an API key, give the keys to Hermes, and that's it. No credit card was needed for what I use. The free models are not toys either. Some of them are very capable, more than enough for an assistant that reads, writes, summarizes and runs a few commands.

Never Stuck on One Model

Free models have one obvious weakness: limits. You can hit a rate limit, a model can be overloaded, or a provider can simply have a bad hour. If your assistant depends on a single model, at that moment you don't have an assistant.

Hermes solves this with fallback providers. You give it a list of models in the order you prefer, and when the first one doesn't answer, it automatically moves to the next one, and then the next. It even does this in the middle of a conversation, without losing what you were talking about.

So my setup is not "one free model". It is a chain of several very capable models from NVIDIA and OpenRouter, and at any moment at least a few of them are ready to do the job. In practice I almost never notice which one actually answered. The task just gets done.

I Built My Own Reminder App First

Hermes was not my first try at this. Before it, I made a reminder application myself, using n8n and an LLM, and I used it through Slack. I wrote about n8n before, in the article on how I automated my freelance job hunt, and it is still one of my favorite tools. The reminder app worked, and I used it every day.

But for this purpose Hermes is far more flexible, and simply better. In n8n every new ability is another piece of workflow that somebody has to design, build and then maintain. Hermes comes with things that are very hard to build into n8n: it remembers what we talked about, it lets me change or cancel a reminder just by saying so, it can search the web or run a check on my network in the same conversation, and it moves to another model on its own when one stops answering.

That doesn't make n8n a worse tool. For a fixed process that runs the same way every time, it is still what I reach for. A personal assistant is just a different kind of problem, and Hermes was made for exactly that.

What I Actually Use It For

Reminders. I will be honest, this is the biggest benefit for me, and it sounds too simple to be impressive. I write "remind me on Thursday at ten to call the client about the renewal" the same way I would write it to a colleague, and on Thursday at ten I get the message. No app to open, no form to fill in. Because it is that easy, I actually use it, and things stopped slipping.

The morning weather report. Every morning Hermes sends me the current weather and the forecast for the day. A small thing, but it is the first message I read with my coffee.

Checking the network like a professional would. This one I like a lot. Hermes can check the state of my local network and the internet connection and send me a proper report. Not just "internet works", but the kind of checks an admin would run by hand, written up in a few clear sentences. For somebody who does networking for a living, it is like having a junior colleague who never gets bored of the routine checks.

Looking things up. When I need some detail from the internet and I don't have time to dig, I ask Hermes. It searches, reads the pages, and comes back with a short answer instead of ten open tabs.

Regular reports. Anything that needs to be checked and summarized every day or every week can be turned into a scheduled task. You describe what you want once, and the report arrives on time from then on.

A Few Things Worth Knowing

These are the features that, in my opinion, make Hermes different from a normal chatbot:

  • Scheduling in plain language. One-time or recurring tasks, created just by asking. You can pause, change or remove them the same way.
  • Memory. It keeps notes about you, your preferences and your environment, and uses them in the next conversation.
  • Skills. When it figures out how to do something, it can save that as a skill and reuse it later, so it gets better the longer it runs.
  • Many ways to reach it. More than twenty messaging platforms through one gateway, plus the terminal.
  • Web search and browsing built in as tools.
  • Safety controls. Dangerous commands need your approval, and you can limit who is allowed to talk to it.

Can It Run on a Local Model?

Yes, and for some people this will be the most important part. Hermes can also use a local LLM, a model that runs entirely on your own hardware. In that case nothing leaves your network: no API key, no provider, no limits.

The condition is hardware. You need a machine that is available locally with a powerful enough GPU, which in practice means enough VRAM to load the model you want. The bigger and smarter the model, the more VRAM it asks for. If you have such a machine, a local model is a great option, and you can even combine the two: local model first, free cloud models as the fallback.

What to Keep in Mind

A few honest notes, so you don't expect magic.

Free models come with limits, and those limits change. This is exactly why the fallback chain is not a nice extra but the thing that makes the whole setup usable.

An agent that can run commands is a serious tool. Give it only the access it needs, keep the approvals on, and think twice before you let it near anything important. I treat it the same way I would treat a new colleague with a fresh account.

And when you use a cloud model, your messages go to that provider. For reminders and weather that is fine. For anything sensitive, that is the moment to think about a local model.

The Bottom Line

A personal AI assistant doesn't have to be an expensive subscription. With Hermes, a couple of free API keys and an hour of setup, I got an assistant that reminds me, reports to me and looks things up for me, and that keeps working even when one of the models behind it doesn't.

The technology is impressive, but what I really noticed after a few weeks is much simpler: I forget fewer things.

What is the one small task you would hand over to an assistant tomorrow, if it cost you nothing to try?

  • Hermes Agent
  • AI Assistant
  • Self-Hosting
  • OpenRouter
  • NVIDIA
  • n8n
  • Automation