# Sigil

**Agents that are watched, not just prompted.** — In production

[Ask about licensing](https://matthicks.com/schedule)

What it is

Sigil runs the whole agent: the conversation, the model calls, tools, memory, context, safety, multi-agent work, and the live stream to every client. An application brings its own models, tools, and storage.

Most frameworks trust the model and hope the prompt holds. Sigil watches what the agent actually does and holds it to the rules around it. That is what lets small, cheap, local models do work that usually needs a frontier model, and it makes frontier models steadier too.

What it does

Any model. The major hosted providers, or a model on your own hardware. Sigil knows what each one costs and what it is good at, and sends the work accordingly.

Context that holds. A model is never handed more than it can take, and what gets summarized along the way can be pulled back in full when the agent needs it.

Memory that lasts. Conversations that run for months keep what matters and recall it when it becomes relevant, with a history behind every fact rather than a silent overwrite.

Tools that declare themselves. Every tool states what it touches and what it would cost to be wrong, so the framework can hold it to that.

Work that is checked. An agent sets goals with criteria, and before it reports success, those criteria are tested against what it actually did.

Many agents. Work is handed to other agents in their own conversations, and long-running jobs survive a restart.

Live everywhere. Replies and tool progress stream to every device in a conversation, and a dropped connection picks up where it left off.

Safety and control

A person approves the calls that matter, one call at a time, and an agent cannot approve itself. Tenants are kept hard apart. Agents cannot reach inward to the machine they run on. Every turn runs under limits on effort and spend, and stopping one costs nothing further.

How it is proven

Sigil is hardened against public agent benchmarks: LongMemEval for memory, SWE-bench Verified for coding, and AgentBench. The point is not a leaderboard position. Weak models expose faults that a frontier model would paper over, so much of this work runs on a local 9-billion-parameter model, and every failure is traced back to a general fix rather than a patch for that benchmark.

LongMemEval-S asks whether an assistant can answer questions about details buried in months of past conversation. Sigil answers 78.0% of all 500 correctly, and the facts an answer needs are in front of the model 89.6% of the time, for about 470 tokens of recalled context per question.

SWE-bench Verified hands an agent real bug reports from open-source Python projects and checks its fix against the project's own tests. Sigil resolves 355 of the 500, working offline, with no way to look up the fix that was actually shipped.

Licensing

Sigil is private and closed source, and runs real products in production today. It is available for commercial licensing.
