portfolio OpenSverige sidekicks
OpenSverige · sidekicks live

Everyone asked their own AI. I put one in the room.

Three AI agents I designed, built and run for OpenSverige. One researches with members, one thinks with the board, one writes with the editors. They answer in Discord threads, where everyone can see.

role
Gustaf Garnow, founder. Product design, build and operations, solo.
what
3 agents for 3 audiences: members, the board, the editors.
proof
17 of 26 members came back · 167 threads · evals 7 → 13 of 14
when
July to October 2026 · live
built with
Hermes Agent
Grok
Discord
Obsidian
Tailscale
Mac mini

One question, from a Discord thread to a Mac mini and back. Example answer.0:42 · no sound

the gap

People pasted their chatbot’s answers into the chat.

whyclose

Someone would ask a question in the server, and someone else would go off to ChatGPT or Claude, come back and paste in a wall of text. The research happened outside the room, there were no sources anyone could follow up on, and the conversation usually ended right there. I wanted the agent inside the conversation instead, as one more participant that everyone can see, question and build on.

Illustration of a Discord channel: a member asks about gaussian splatting, another pastes a long answer copied from ChatGPT, and then nothing happens for three hours.
Before the agents. The research happened in another tab.
what we cut

We killed the agent that watched members.

whyclose

My first attempt ran on OpenClaw, behind the scenes, and kept track of member activity: who posted, how often, who had gone quiet. It was rigid, and as the server owner it never told me anything I could act on. Worse, it pointed in the wrong direction. I didn’t want agents that report on the community. I wanted agents that talk with it, where everyone can see. It was shut down on 18 August.

Illustration of a member activity dashboard listing anonymous member numbers, last seen and posts per week, struck through in red and stamped decommissioned 18 August.
The first version tracked activity. It went in August.
the design

Ask in a thread. Everyone can join.

howclose

Mentioning Supergrok-Handen opens a Discord thread around the question. Other members jump in, push back and keep brainstorming, and the agent stays in the thread with web and X search while they talk. It runs on Grok for one reason: Grok can search X, and in AI the news breaks on X first. It only ever answers in public, because people start using it once they see others use it. Every reply follows the same format: a short notice instead of an essay, no emojis, links masked so the thread doesn’t fill up with preview cards, sources last and ten lines at most.

Example of a Discord thread in #supergrok: three members and Supergrok-Handen in the same conversation, with short sourced answers and links.
One thread, the people and the agent together. Example.
the morning signal

A newsletter, delivered in Discord.

whyclose

Every morning at 07:00 Supergrok-Handen posts the day’s LLM signals in #llm-news: the three stories that matter for builders, one line on what each means in practice, and quick hits with sources. It’s written short on purpose, for the morning coffee. I took it away for a while, and several people asked where it had gone, lurkers included. That’s how I know it gets read.

A real post in #llm-news from Supergrok-Handen at 07:01: LLM signals for 1 October 2026 with a summary, three numbered stories with source links and a list of quick hits, in Swedish.
The real post from 1 October, in Swedish. Three stories, quick hits, sources.
three audiences

Three agents. Three sets of rules.

membersSupergrok-Handen researches with members.

It searches the web and X, answers in threads with sources last, and is allowed opinions and dry humour. It runs on Grok because Grok can read X, where AI news breaks first. Every morning it also posts the day’s LLM signals. 167 threads since July.

the boardParagraf is the association’s sidekick. The board decides.

It is collective intelligence between members and the board, aimed at a board that can run itself. Before it answers, it rates what is at stake: one line for small things, and a fixed four-point template with a section reference for tax, contracts and exclusions. When the answer isn’t in its sources, it says so and points to where it is.

the editorsRedaktionen writes. A human publishes.

Editors tag ideas, and every idea has to pass three yes-questions and 48 hours of sparring before it becomes a draft. Memory and profiling are switched off, and it never invents quotes, numbers or people.

the rules

Code checks every message, before and after the model.

whyclose

None of these checks is AI. They are plain Python in Hermes plugins and hooks, so they give the same result every time, they can be tested, and they don’t depend on the model behaving. The model writes the answers. The code decides what reaches it and what goes out.

Pipeline: Discord's role gate and mention requirement, then checks before the model (personal ID numbers replaced with PNR, a daily quota, silent notes), then Grok with limited tools, then checks after the model (every statute section and law reference verified, everything logged).
Every check is plain Python in Hermes hooks. Same input, same result.
beforePersonal ID numbers never leave the Mac mini.

Before a board question is sent to Grok, a hook looks for Swedish personal ID numbers: ten or twelve digits with a valid date, coordination numbers included, the association’s own organisation number excluded. Each one is replaced with [PNR] and the question is still answered. The audit log records that a number was removed, never the number.

afterEvery section Paragraf cites has to exist.

When Grok has answered, the same plugin checks every § against the statutes (§ 1–13) and every law reference against the list of sources Paragraf is allowed to use. A reference that doesn’t exist gets a visible warning under the answer in Discord, and every check is logged with the date, the references and the result.

toolsEach agent gets only the tools its job needs.

Supergrok-Handen has web search, X search and image generation. Paragraf has its own files, the web, a decision log and scheduled reminders. Redaktionen has no file tools at all, and nine tools are blocked in code, including writing files, running code, sending messages and rewriting its own skills.

quota15 questions per member a day.

A hook counts each member’s unique questions per day by a hash of the question, not its text, and stops at 15. It keeps a shared Grok account fair for everyone.

aggregateThe weekly report counts people. It never names them.

Every Sunday a script reads the week’s questions, counts them and the number of people who asked, strips mentions and cuts each question to 110 characters. Names are counted but never printed. Grok turns that into a public report in #supergrok.

the machine

One Mac mini. Three isolated agents.

howclose

Each agent is a separate Hermes Agent gateway with its own folder, Discord app, tools and rules, so a problem in one never touches the others. They run on a Mac mini under launchd, which restarts a gateway that crashes, and a nightly health check plus a test after every reboot cover the rest. Short questions go to a faster Grok model to keep replies quick. Paragraf reads the statutes and the law texts it relies on from an Obsidian vault, and I build and run all of it over Tailscale from my laptop.

Diagram: the OpenSverige Discord server with three channels connects to three agents on a Mac mini running Hermes Agent. The agents call Grok, Paragraf reads an Obsidian vault, and Gustaf's laptop reaches the Mac mini over Tailscale.
Discord, a Mac mini, Grok. And an Obsidian vault for the board.
the eval

I tested the guard. The guard had bugs.

howclose

To know whether Paragraf actually follows its rules, I wrote an eval: 14 real board questions whose answers are in the statutes, three of them traps. Every check is deterministic code, so no model grades another model. Does each cited section exist, is the stakes rating right, is the answer within its length, does it ever repeat an ID number. The first run found three bugs, all in my own guard and none in the model: a real law flagged as unknown, a citation format nobody checked, and a warning stuck under a correct answer. After the fix the guard passes 13 of 13 tests and the answers went from 7 to 13 of 14. The last one is a judgment call, and I left it as a fail rather than move the goalposts.

Eval results: 14 board questions with before and after marks. The guard went from 9 of 11 to 13 of 13 tests, the answers from 7 of 14 to 13 of 14.
Before and after the fix. Two guard tests were added with it.
the numbers

Conversations, not lookups.

167threads started with the agent
4member messages per thread, on average
17/26members came back and asked again
806questions from members since July

26 members tried Supergrok-Handen between July and September. The 17 filled in came back on another day and asked again.

in short

Agents in the room, not behind it.

  • Three agents, three audiences
  • Threads, not DMs
  • Guardrails before and after the model
  • A daily signal at 07:00
  • Evals: 7 → 13 of 14
  • 17 of 26 came back