A Sovereign AI Agency Experiment

A Sovereign AI Agency Experiment

TL;DR: This is a early blog on the self-hosted AI agent stack under construction: the Strands Agents framework, IBM’s Granite 4.1 running locally via Ollama, and IRC as the human-in-the-loop channel. The example agent is in the repo, earnings_analyst, which summarizes SEC earnings filings, a deliberately low-stakes testbed. The point of the testbed is agents that reason over data too sensitive to hand to a hyperscaler, and this is the harness they’ll run on once it’s proven out. This post covers what’s built so far, and the code lives in the statixs-agency repository.


Why sovereign, and why now

The agent actually running for me in the statixs-agency repo today is earnings_analyst. It watches SEC EDGAR for a company’s most recent earnings filing, asks a locally-hosted model to summarize it, and posts the summary to IRC. There’s nothing sensitive about a public earnings release, so on its own it doesn’t need to be sovereign at all. It’s built that way anyway, because it’s the testbed for a harness meant for data that does.

Every AI demo so far follows the same shape. An agent calls a frontier model hosted by a hyperscaler, the model reasons over your data, and a tool call somewhere down the chain takes action through the harness. It works fine, right up until the data is the kind of thing that is sensitive or classified. None of that should be shipped to a third-party inference endpoint.

Sovereign in this post means three concrete things:

  • Inference runs locally. No data leaves the host running the model.
  • The agent framework is open and self-hostable. No SaaS control plane.
  • Communication channels are protocols you already operate. No vendor messaging fabric.

The interesting question is whether that combination is actually usable today, or whether you give up too much capability to make it worthwhile. After a few weeks of building, the answer looks like yes, with some sharp edges, and earnings_analyst is the evidence so far. Security-shaped agents (a honeypot triage agent, a MalZoo triage agent) are next, once the harness has more mileage on it; see “What’s next” below.


The stack

Three components, deliberately small. Here’s how they fit together, with inference, framework, and alerting all inside the same perimeter:

   SEC EDGAR    ─────▶┌────────────────────────────┐
   (public data)      │  earnings_analyst  (agent) │─┐
                      └────────────────────────────┘ │ Strands
                                                      │ tool loop
                      ┌────────────────────────────┐ │
                      │  Granite 4.1  (Ollama,     │◀┘
                      │  local GPU, no SaaS)       │
                      └──────────────┬─────────────┘
                                     │ send_irc_message()
                                     ▼
                      ┌────────────────────────────┐
                      │  ircd  (homelab IRC server)│
                      └──────────────┬─────────────┘
                                     │ #alerts channel
                                     ▼
                      ┌────────────────────────────┐
                      │  operator @ irssi          │
                      └────────────────────────────┘

   └─── inference, framework, and alerting stay on-prem; EDGAR is public ───┘

Strands Agents is the agent framework. It’s open source, written in Python, and the surface area is small enough that you can read the whole library in an afternoon. The mental model is the same as most modern agent frameworks, an Agent with a system prompt and a list of @tool-decorated Python functions, but it doesn’t assume a specific model provider. Bedrock is the default, but earnings_analyst points strands.models.ollama.OllamaModel at a local endpoint instead (a Docker container with Ollama), so the swap is a model-config line rather than a rewrite.

Granite 4.1:8b, served through Ollama, is the model. IBM released the 4.x series under Apache 2.0, the smaller variants run comfortably on a single workstation GPU, and the function-calling format is documented and stable. Granite isn’t Claude, and it won’t one-shot a complex multi-tool plan the way a frontier model does. For a narrow, well-scoped task, like turning one earnings press release into a plain-language summary, it holds up well. Where it doesn’t hold up is the subject of a section below.

IRC, talked to via irssi on the human side, is the communication channel. This is the part that surprises people, so the next section is dedicated to it.


Why IRC

The first time this setup came up with a colleague, the response was “you’re using a 1988 protocol because…?” The reasons stack up faster than you’d think.

IRC is text-only, line-oriented, and trivial to script against. The entire client implementation in statixs-agency/tools/irc.py is around a hundred lines of Python with no dependencies beyond the standard library. Compare that to a Slack bot, which needs an app registration, OAuth scopes, signing secrets, and a webhook endpoint reachable from the public internet. For a sovereign deployment that shouldn’t need an outbound HTTP path to a SaaS, that’s the wrong shape.

It also runs on infrastructure you already control. A single ircd instance on a homelab box handles every agent and every operator. There’s no third-party retention of message content, no analytics pipeline ingesting your alert stream, and no cloud egress.

The audit trail is just text. irssi keeps logs. Every alert an agent sends, every operator response, and every back-and-forth between agents on the same channel is plain text, so it’s grep-able, diffable, and replayable. For a security tool that’s exactly what you want.

It’s also bidirectional. The same channel the agent uses to alert is the channel an operator uses to ask the agent a follow-up question, so there’s no separate “agent UI” to build.

The IRC tool is intentionally stateless. It opens a fresh TCP connection per call, registers, joins, posts, and quits, which means there’s no long-lived client to crash, no reconnect logic to get wrong, and no shared state between agents:

@tool
def send_irc_message(message: str, channel: Optional[str] = None) -> str:
    """Send a message to an IRC channel on the homelab IRC server.

    Opens a fresh connection, registers, joins the channel, posts the
    message, and disconnects. Use this for alerts and status updates from
    agents.
    """
    return _send_irc_message_impl(message, channel)

One detail worth calling out: any agent that relays externally-sourced text into a channel, will eventually try to send something with an embedded newline or control character. If you don’t strip CR/LF, the agent has just been turned into an IRC command injection primitive by whatever produced that text. The tool sanitises every line:

def _sanitize(text: str) -> str:
    # Strip CR/LF and other control chars. Without this, an attacker-controlled
    # value (a malicious DNS query, a process arg) could inject extra IRC commands.
    return "".join(c for c in text if c >= " " and c not in "\r\n")

This is the same class of bug as log injection or HTTP header injection. It’s boring and well-understood, and it’s absolutely the kind of thing that gets missed when “the agent” is treated as a magic box rather than a piece of software that calls real network functions.


What the agents look like

The repository layout is deliberately flat. Each agent is a directory with its own agent.py and tools.py; a runner.py is added for agents that need to run unattended. Tools that only one agent uses live next to the agent. Tools that more than one agent uses get promoted to the top-level tools/ directory:

statixs-agency/
├── agents/
│   └── earnings_analyst/
│       ├── __init__.py
│       ├── agent.py          # Agent definition and entry point
│       ├── runner.py         # Event-driven runner (consumes earnings-radar events)
│       └── tools.py          # Agent-specific tools (@tool functions)
├── tools/
│   ├── __init__.py
│   └── irc.py                # IRC notification (shared)
├── config/
│   └── settings.py           # Shared settings from environment
└── tests/                    # Unit tests per agent and per shared tool

One agent is wired up so far: earnings_analyst finds a company’s most recent SEC EDGAR earnings filing, fetches the press release text, and asks Granite for a plain-language summary. It runs two ways: on demand (python -m agents.earnings_analyst.agent CRWD), or unattended via runner.py, which watches a directory for events dropped by a companion project, earnings-radar, and summarizes each new filing as it lands. It uses the shared IRC tool for delivery, and it’s the reason the shared-vs-agent-specific split in tools/ exists at all: send_irc_message is generic, while the EDGAR lookup and fetch tools are specific to this agent and live in its own tools.py.

The agent definition itself is close to Strands’ own boilerplate, with one addition, an explicit local model:

from strands import Agent
from strands.models.ollama import OllamaModel

from config import settings
from tools.irc import send_irc_message
from .tools import find_earnings_filings, fetch_report_text

SYSTEM_PROMPT = """You are an equity research analyst..."""

agent = Agent(
    model=OllamaModel(host=settings.OLLAMA_HOST, model_id=settings.OLLAMA_MODEL_ID),
    system_prompt=SYSTEM_PROMPT,
)

Strands handles the tool-call loop when there is one. earnings_analyst actually skips it for data fetching, which is the subject of the next section.


Where Granite holds up, and where it doesn’t

A few honest observations from running this for a few weeks, all specific to earnings_analyst since that’s the only agent with real mileage on it so far.

It holds up well on narrow summarization over a single, already-fetched document. Feed it one earnings press release and ask for a roughly 200-word plain-language summary covering revenue, EPS, guidance, and margins, and Granite 4.1 does this reliably and stays inside the word budget.

It does not hold up on multi-step tool orchestration, at least not at the model size in use here. An early version let the model call the EDGAR-lookup tools directly as part of its own reasoning loop, chaining “find the filing, fetch the exhibit, decide if it’s the right one” itself. An 8B model lost the thread constantly: repeated tool calls, fetched the wrong exhibit, or stalled mid-chain. The fix was structural, not a bigger model. The EDGAR lookup and text fetch now happen deterministically in plain Python, and the model is only invoked once, on text that’s already been retrieved, purely to summarize it. The model never drives the tool loop for this agent; it’s a single-shot completion over a large prompt.

That same lesson set the scope for the agent’s first version: summarize only the single most recent filing. A quarter-over-quarter comparison would need the model to reason across two documents at once, which is exactly the kind of multi-step reasoning that fell over above. That’s parked as a separate task, likely backed by a stronger model, once there’s enough summary history to compare against, rather than something to force onto the local model today.

The takeaway matches what a lot of people are converging on. A local model is a fine agent runtime for one well-scoped decision made over data you hand it directly. Ask it to plan and chain tool calls on its own, and the wheels come off faster than you’d expect from an 8B-class model.


What’s next

earnings_analyst was the stab at a sovereign AI agent platform. Some security-shaped agents worth exploring:

  • A honeypot-event agent that consumes Cowrie session logs (the same Cowrie stack from the SSH honeypots post) and triages sessions worth a human look.
  • A MalZoo-output agent that reads from the analysis_record schema defined in the data contracts post and decides whether a sample warrants escalation. The contract is what makes this tractable, because without a formal schema the agent is guessing at field semantics.

Both will need the same discipline earnings_analyst forced on the design: keep the model out of the tool-orchestration loop, fetch deterministically, and reserve the model for the one judgment call it’s actually good at. Or it’s time to spend some money on a stronger homelab that can run bigger models locally.

The statixs-agency repository is public now, with the shared IRC tool, earnings_analyst, and a worked example of running Granite through Ollama instead of the Bedrock default. Fair warning: it’s a one-agent harness so far, an early cut rather than something to point at your live telemetry yet.


Final thoughts

The point of the experiment was never to prove that homelab ran local models can replace frontier models for heavy general-purpose agent work. They can’t for most budgets and pretending otherwise is how you end up with a tool that confidently reasons with wrong data.

The point was to test whether a sovereign stack (local inference, open framework, boring protocols) can carry a useful slice of agent workload, before trusting it with anything sensitive. So far the answer is yes, provided you keep each agent narrow, keep the model out of multi-step tool orchestration, and treat the agent’s outputs as software that needs the same input-validation discipline as anything else on a network.

More to come as statixs-agency shapes up, starting with the security agents this whole thing was built for.

Happy hunting :)