Everyone asked their own AI. I put one in the room.
Three AI agents I designed, built and run for OpenSverige. One researches with members, one thinks with the board, one writes with the editors. They answer in Discord threads, where everyone can see.
- role
- Gustaf Garnow, founder. Product design, build and operations, solo.
- what
- 3 agents for 3 audiences: members, the board, the editors.
- proof
- 17 of 26 members came back · 167 threads · evals 7 → 13 of 14
- when
- July to October 2026 · live
- built with
- Hermes Agent
- Grok
- Discord
- Obsidian
- Tailscale
- Mac mini
People pasted their chatbot’s answers into the chat.
whyclose
Someone would ask a question in the server, and someone else would go off to ChatGPT or Claude, come back and paste in a wall of text. The research happened outside the room, there were no sources anyone could follow up on, and the conversation usually ended right there. I wanted the agent inside the conversation instead, as one more participant that everyone can see, question and build on.
We killed the agent that watched members.
whyclose
My first attempt ran on OpenClaw, behind the scenes, and kept track of member activity: who posted, how often, who had gone quiet. It was rigid, and as the server owner it never told me anything I could act on. Worse, it pointed in the wrong direction. I didn’t want agents that report on the community. I wanted agents that talk with it, where everyone can see. It was shut down on 18 August.
Ask in a thread. Everyone can join.
howclose
Mentioning Supergrok-Handen opens a Discord thread around the question. Other members jump in, push back and keep brainstorming, and the agent stays in the thread with web and X search while they talk. It runs on Grok for one reason: Grok can search X, and in AI the news breaks on X first. It only ever answers in public, because people start using it once they see others use it. Every reply follows the same format: a short notice instead of an essay, no emojis, links masked so the thread doesn’t fill up with preview cards, sources last and ten lines at most.
A newsletter, delivered in Discord.
whyclose
Every morning at 07:00 Supergrok-Handen posts the day’s LLM signals in #llm-news: the three stories that matter for builders, one line on what each means in practice, and quick hits with sources. It’s written short on purpose, for the morning coffee. I took it away for a while, and several people asked where it had gone, lurkers included. That’s how I know it gets read.
Three agents. Three sets of rules.
members
Supergrok-Handen researches with members.
It searches the web and X, answers in threads with sources last, and is allowed opinions and dry humour. It runs on Grok because Grok can read X, where AI news breaks first. Every morning it also posts the day’s LLM signals. 167 threads since July.
the board
Paragraf is the association’s sidekick. The board decides.
It is collective intelligence between members and the board, aimed at a board that can run itself. Before it answers, it rates what is at stake: one line for small things, and a fixed four-point template with a section reference for tax, contracts and exclusions. When the answer isn’t in its sources, it says so and points to where it is.
the editors
Redaktionen writes. A human publishes.
Editors tag ideas, and every idea has to pass three yes-questions and 48 hours of sparring before it becomes a draft. Memory and profiling are switched off, and it never invents quotes, numbers or people.
Code checks every message, before and after the model.
whyclose
None of these checks is AI. They are plain Python in Hermes plugins and hooks, so they give the same result every time, they can be tested, and they don’t depend on the model behaving. The model writes the answers. The code decides what reaches it and what goes out.
beforePersonal ID numbers never leave the Mac mini.
Before a board question is sent to Grok, a hook looks for Swedish personal ID numbers: ten or twelve digits with a valid date, coordination numbers included, the association’s own organisation number excluded. Each one is replaced with [PNR] and the question is still answered. The audit log records that a number was removed, never the number.
afterEvery section Paragraf cites has to exist.
When Grok has answered, the same plugin checks every § against the statutes (§ 1–13) and every law reference against the list of sources Paragraf is allowed to use. A reference that doesn’t exist gets a visible warning under the answer in Discord, and every check is logged with the date, the references and the result.
toolsEach agent gets only the tools its job needs.
Supergrok-Handen has web search, X search and image generation. Paragraf has its own files, the web, a decision log and scheduled reminders. Redaktionen has no file tools at all, and nine tools are blocked in code, including writing files, running code, sending messages and rewriting its own skills.
quota15 questions per member a day.
A hook counts each member’s unique questions per day by a hash of the question, not its text, and stops at 15. It keeps a shared Grok account fair for everyone.
aggregateThe weekly report counts people. It never names them.
Every Sunday a script reads the week’s questions, counts them and the number of people who asked, strips mentions and cuts each question to 110 characters. Names are counted but never printed. Grok turns that into a public report in #supergrok.
One Mac mini. Three isolated agents.
howclose
Each agent is a separate Hermes Agent gateway with its own folder, Discord app, tools and rules, so a problem in one never touches the others. They run on a Mac mini under launchd, which restarts a gateway that crashes, and a nightly health check plus a test after every reboot cover the rest. Short questions go to a faster Grok model to keep replies quick. Paragraf reads the statutes and the law texts it relies on from an Obsidian vault, and I build and run all of it over Tailscale from my laptop.
I tested the guard. The guard had bugs.
howclose
To know whether Paragraf actually follows its rules, I wrote an eval: 14 real board questions whose answers are in the statutes, three of them traps. Every check is deterministic code, so no model grades another model. Does each cited section exist, is the stakes rating right, is the answer within its length, does it ever repeat an ID number. The first run found three bugs, all in my own guard and none in the model: a real law flagged as unknown, a citation format nobody checked, and a warning stuck under a correct answer. After the fix the guard passes 13 of 13 tests and the answers went from 7 to 13 of 14. The last one is a judgment call, and I left it as a fail rather than move the goalposts.
Conversations, not lookups.
26 members tried Supergrok-Handen between July and September. The 17 filled in came back on another day and asked again.
in short
Agents in the room, not behind it.
- Three agents, three audiences
- Threads, not DMs
- Guardrails before and after the model
- A daily signal at 07:00
- Evals: 7 → 13 of 14
- 17 of 26 came back