# Inside ClarkCant: how one message becomes an answer, a widget or a task

A source-level tour of ClarkCant: one runtime and many clients, a conductor that decides, Pi behind a single adapter, widgets as a vocabulary, a policy that only narrows, memory you can delete, and voice that is only a voice.

Duy Nguyen /zuey/ · 2026-10-05T10:09:10.065Z

Source: https://clarkcant.cc/blog/clarkcant-architecture-deep-dive

People ask me what ClarkCant actually is under the Orb. The honest answer: a Node.js runtime that owns a SQLite database, an agent loop it borrows from Pi, and a lot of rules about who is allowed to decide what. The chat window is just the part you see.

This post walks through the architecture the way the code does it, not the way a pitch deck would. I'll follow one message from the composer to the reply, then look at the pieces that make that trip safe: widgets, policy, memory, voice, and the sessions and nodes you are never supposed to notice. Every file path here is real. Where something is designed but not built, I'll say so.

One rule from DESIGN.md frames everything below: don't make the user understand the architecture to get work done. Which is a funny thing to write a few thousand words about. But you're an engineer, so you asked.

## One runtime, many clients

ClarkCant is not a desktop app with a helper process. From the first commit it has been a portable runtime with many conversation clients. apps/runtime is the headless node: the composition root, an authenticated command gateway over HTTP and WebSocket, node identity, lifecycle and health. apps/web is the browser client. apps/desktop is a thin Electron shell with typed IPC, contextIsolation on, nodeIntegration off, and no scheduler of its own. Run the same runtime on a VPS without Electron and that's the server distribution.

Each node has its own identity, its own SQLite database in WAL mode, its own credentials and its own folders. A conversation has exactly one home node. Nodes can be paired and hand each other scoped work over NodeLink, but they never share a filesystem or a live database. The architecture doc says it better than I can: seamlessness lives in communication, delegation and presentation, not in hiding every difference or copying secrets around.

The runtime executes TypeScript directly on Node 22.19 or newer, and SQLite comes from the built-in node:sqlite, so the core needs no native modules. That was on purpose. Fewer things to compile means fewer things to audit.

ClarkCant repository numbers on 5 October 2026: 685 commits since the first one on 16 September 2026; 5 apps, 18 packages and 6 first-party packs; in docs/conformance-traceability.md, 10 of 18 scope items are PASS and 8 are PARTIAL, and all 77 acceptance tests (T01 to T77) are PASS.

```json
[
  {
    "id": "commits",
    "label": "Commits",
    "value": 685,
    "hint": "16 Sep to 5 Oct 2026"
  },
  {
    "id": "apps",
    "label": "Apps",
    "value": 5
  },
  {
    "id": "packages",
    "label": "Packages",
    "value": 18
  },
  {
    "id": "packs",
    "label": "First-party packs",
    "value": 6
  },
  {
    "id": "scope",
    "label": "Scope items at PASS",
    "value": 10,
    "unit": "/ 18",
    "hint": "the other 8 are PARTIAL"
  },
  {
    "id": "tests",
    "label": "Acceptance tests at PASS",
    "value": 77,
    "unit": "/ 77"
  }
]
```

## Following one message

You type something and press Enter. The web client posts it to /conversations/\{id\}/messages/stream, and the node answers over Server-Sent Events. Every route except /health, /openapi.json and the discovery document needs the node's bearer token, and a missing token gets the same 401 as a wrong one.

The first thing the node does with your sentence is try not to call a model at all. A slash command like /new is the host's to answer. A typed app intent, such as asking to open settings, is matched and answered by the host too. Anything you pointed at with @, a project, a file, another conversation, is checked again, and if it no longer holds, the whole message is refused by name instead of starting a turn without the thing you asked about.

Then a question most chat apps never ask: is Clark already answering in this conversation? If so, your message isn't just queued. A small decision, decideTurnAction, picks between steer (join the running answer), interrupt (replace it) and background (run it beside the current work). When it can't decide with confidence, it falls back to interrupt, which a comment in the code calls the recoverable direction rather than the silent one.

Flow of one message through a ClarkCant node. You type into the web or desktop client, which posts to the messages/stream route; the gateway checks the bearer token. Slash commands, app intents and @ references are handled by the host first. If Clark is already answering, a decision picks steer, interrupt or background. The message then reaches the conductor (handleUserMessage); a spoken sentence reaches the same conductor as its final transcript from the voice socket. If an installed capability fits, the conductor creates a task and dispatches a worker process; otherwise the model answers in a Pi session through pi-adapter, briefed by the context planner (recap and memory brief). The model calls node tools; tools with effects, and a worker's commands, pass preflight, the execution policy, the guardrail and the secret broker, and are written to audit\_log and the effect ledger. The show\_view tool goes to the view catalog, which builds a surface block placed in the reply. The reply is appended to SQLite (messages and search index), read aloud when the turn was spoken, and the client renders surface blocks with the shared widget renderer.



Nodes and edges follow apps/runtime/src/routes/conversations.ts, packages/core/src/conductor.ts, apps/runtime/src/model-turn.ts and apps/runtime/src/voice-session.ts.

Next the message reaches the conductor in packages/core. Its job is narrow, and it's the only place that makes this call: does this sentence become an answered message or a durable task? The order is deliberate. A real installed capability first. Then, only on the demo path, a scripted sample that is labelled as a sample. Then the model's own words.

Excerpt of handleUserMessage in packages/core/src/conductor.ts, lines 730 to 758. It lists usable capabilities and chooses an execution node; only when none is chosen are sample recipes considered, and a recipe runs only for a message from the demo path, with catch-all recipes dropped once a model exists. Then, with no execution node and a model available, runModelTurn answers. The order is real capability, then scripted demo, then the model's own words.



packages/core/src/conductor.ts, lines 730–758, inside handleUserMessage.

If a capability can do the work, the conductor creates a task, moves it through a state machine and hands it to a worker process. In that state machine, success is reachable only through verification with recorded evidence. An idle worker, a model ending its turn or a closed socket is not success. And if no model is configured and nothing installed can help, the task parks and Clark says so in plain words instead of improvising.

Most of the time, though, the model answers. The turn runs in the node's own process on a Pi session. Pi is the agent loop, and ClarkCant doesn't fork it. packages/pi-adapter is the only package that imports the Pi SDK; everything else depends on a PiAdapter interface. There's a RealPiAdapter typed against the SDK and a deterministic FakePiAdapter, which is how the whole app runs in CI without a provider account.

Before the prompt goes out, the context planner decides what the model is told. A fresh session for an existing conversation gets a recap: the newest 12 of the newest 40 messages, plus up to 4 older ones that share terms with what you just asked. Each turn also gets a memory brief built from records you can see. Every block the host adds is labelled public, internal, confidential or secret from its text alone, and a model profile can be limited to the classes it is allowed to receive.

## Widgets are a vocabulary, not a canvas

This is the part I'm proudest of. A reply can contain more than prose, but the model never writes UI. It calls a tool named show\_view with a view name from a catalog the host supplies, plus some values. The host's view catalog (apps/runtime/src/view-catalog.ts) checks those values and builds the block itself: a surface block that points at a widget definition and carries a snapshot with a text alternative.

The comment at the top of that file says the important bit: the catalog is what makes a forged approval unrepresentable rather than merely rejected. There's no parameter called type, owner or status in the tool's schema. The model has no words for a host-owned card. Approval cards, credential cards and task status are built by the node from its own records, and a model has no records to build one from.

Widgets come in trust lanes. Built-in catalog widgets are trusted React code rendering schema-checked props. Declarative compositions arrange existing components with no executable payload. Third-party UI runs in an isolated iframe on an opaque origin, never on the chat's origin, and talks to the host through a bridge. The actions a widget offers are compiled into bindings the host re-authorizes on every press; a widget can't call a tool just because it knows the name.

The catalog lives in packages/widget-catalog and the renderers live in packages/conversation-client. That second detail matters more than it looks. The same renderers are exported through a public entry, public-widget.tsx, which mounts one widget into a shadow root with local view state only: no actions, no runtime connection, no approvals, no credentials. If you're reading this with JavaScript on, the diagram above was drawn by that entry. This blog bundles it straight from the ClarkCant source instead of copying it.

## Autonomy, and who holds the brakes

ClarkCant is autonomous by default. In code the execution modes are autonomous, guarded and ask, and the default is autonomous. That sounds reckless until you see where authority actually sits.

Every effect, whether a command, a file write or a widget action that does more than change the view, goes through the same order. First host preflight (apps/runtime/src/preflight.ts): deterministic, no model call, impossible to argue with. It checks that the effect lands inside a folder this node owns, that the target and the capability exist, and it attaches a deadline and an output ceiling. Then the execution policy in packages/core decides who gets asked, in a fixed order: a prohibition you set beats an OS or OAuth consent boundary, which beats your own rules, which beat the mode. Then a guardrail judgment from Jev, which can allow, deny, constrain or ask a clarifying question, and can only ever narrow. Then the secret broker injects a credential just in time.

That last one is my favourite boring feature. The agent learns that github\_token exists, what it's for and who may use it. The value goes from a backend into one invocation, a child process's environment or a request header, and never back to a model. If a command that received a secret prints it anyway, every injected value of eight or more characters is replaced with \[redacted\] before the output reaches the conversation.

Everything with an effect lands in audit\_log, an append-only table that is shallow on purpose: what ran, when, how it ended. Never arguments, never values. And there's a big red button: POST /stop kills the process group of every running command, worker and terminal, interrupts running turns and cancels queued background work.

Who decides, in order

A prohibition you set beats an OS or OAuth consent boundary. That boundary beats your own rules. Your rules beat the execution mode. The guardrail can only narrow what is left.

## Memory you can read and delete

Memory isn't a mystery vector blob. When Clark decides something is worth keeping, the remember tool writes a sentence to memory\_records, redacted before it's written, along with the conversation it came from. The brief that goes into a turn is read fresh from the database every turn, not cached in the process. Delete a memory in the Memory tab and it stops reaching the prompt on the next turn, not after a restart. Your original message stays in the conversation: deleting a memory removes its way back into the prompt, not your history.

Search over history is lexical by default, SQLite FTS5 with BM25. We built the semantic path too, sqlite-vec with a quantized E5-small model and reciprocal rank fusion, and then measured it. On our labelled corpus, plain FTS got the top result right on 31 of 34 queries. Hybrid also got 31 at its best setting and dropped to 25 at looser ones, because nearest-neighbour search always returns something, even when the honest answer is nothing. So semantic search is off by default and the code waits behind a flag until a bigger corpus says otherwise. I like decisions that come with a number attached.

## Voice is the voice, not the mind

Voice runs on Gemini Live, default model gemini-3.8-live, behind a WebSocket proxy on the node. The browser never holds a provider credential, not even a short-lived token. It sends 16 kHz PCM16 audio to the node and gets audio back. The socket's first frame has to be an auth message, because a token in a WebSocket URL ends up in access logs.

The part that took a while to get right: the live model isn't allowed to answer. Your spoken sentence is transcribed, and the final transcript goes through handleUserMessage, the same conductor a typed message reaches. The agent answers with its tools, memory and context, and the live session's only job is to read that answer aloud. Otherwise you'd have two answers in the room, and only one of them read your files.

```typescript
/**
 * What the live session is for.
 *
 * It is the voice, not the mind. Left to itself the model answers whatever it hears, and then the
 * same question has two answers in the room: the model's guess, made without any tool and without
 * the conversation, and the agent's, which is the only one of the two that read the files, ran the
 * command or knows what was said five minutes ago. So the session is told to transcribe and to read
 * back, and to leave answering to the agent.
 */
const VOICE_INSTRUCTION = [
  "Bạn là giọng nói của trợ lý, không phải bộ não của nó.",
  "Có hai loại đầu vào và hai việc khác nhau, đừng lẫn chúng với nhau.",
  /*
   * Silence while the person speaks.
   *
   * This used to ask the model to transcribe what it heard, and the transcription was already being made
   * without it: both transcription configs are on, and the provider's own reading of the input is what the
   * transcript is built from. So the only thing that instruction added was the model saying those words out
   * loud - the person's own sentence, read back to them by their assistant before it had any answer to give.
   */
  "Khi nghe tiếng người dùng nói: giữ im lặng, không nói gì, không chép lại, không trả lời, không hỏi lại, không bình luận.",
  "Khi nhận được một lượt văn bản: đó là câu trả lời của trợ lý, và việc của bạn là đọc nguyên văn đoạn văn đó ngay lập tức, không thêm bớt chữ nào.",
].join(" ");

/**
```

apps/runtime/src/voice-session.ts, lines 445–469. The four instruction strings are Vietnamese in the source. They say: you are the assistant's voice, not its brain; there are two kinds of input and two different jobs; while the user speaks, stay silent, don't transcribe, answer, ask or comment; when a text turn arrives, it is the assistant's answer, read it verbatim at once.

Voice also knows what you're looking at. The client tells the node which widget has focus, so “move that to Friday” has a target, and a spoken answer to a question card goes through the exact function a click uses. Voice isn't a second product with its own code path.

## The machinery you're not supposed to see

There's no session picker in ClarkCant, on purpose. Underneath, a conversation does hold a Pi session, and sometimes several over its life. Switch models with Cmd/Ctrl+\] and nothing mutates the live session: the node writes your preference and, at the next turn boundary, hands off to a new generation briefed with a recap. A session that was evicted after going idle is rebuilt the same way. From your side it's still one conversation.

Background work is the same story. Background requests, run\_command processes, task workers and terminals all register with one WorkSupervisor. By default at most 3 background jobs run at once (1, 3 or 5 in Settings) with a queue of 10. Every child gets its own process group, so stopping a command also stops the sleep 60 & it left behind. After a crash, each unfinished job is reported once in the conversation that asked for it, and a command with effects is never re-run automatically. You can ask Clark what's running or open the process panel, but you never have to manage it.

Paired nodes follow the same rule. Today a home node can hand a scoped task to a confirmed peer through an automation that names it as executor. Grants intersect rather than compose, so A pairing with B and B pairing with C gives A nothing on C, and the receiving node's own policy still applies to every effect. You shouldn't need to know the topology to ask for status, but the target is always shown when an action matters.

## Where everything lives

Folder tree of the ClarkCant monorepo. apps/runtime: Headless node: gateway over HTTP and WebSocket, identity, lifecycle, health; apps/web: Browser client, same conversation components as desktop; apps/desktop: Thin Electron shell: typed IPC only, hosts no scheduler; apps/worker: Runs one app-managed Pi session per run, reports evidence; apps/cli: The clarkcant command: ask, read conversations, MCP over stdio; packages/core: Invariants: task/run state machine, policy, grants, leases, capability registry; packages/contracts: Runtime-validated contracts: commands, events, tasks, effects, widgets, NodeLink; packages/storage: SQLite WAL, migrations, outbox/inbox, effect ledger, backups; packages/pi-adapter: The only package that imports the Pi SDK; packages/conversation-client: Shared React UI: timeline, composer, cards, widget renderers; packages/widget-catalog: Canonical widget catalog, fixtures, prop validation; packages/widget-host: Built-in registry, mini-app sandbox policy, action-proposal compiler; packages/widget-sdk: Widget author contracts: props, state, events, actions, semantics; packages/widget-cli: Scaffold, conformance-test and pack a widget; packages/design-tokens: Design tokens, appearance compiler, WCAG contrast utilities; packages/node-link: Peer protocol: envelopes, version negotiation, dedup, delegation grants; packages/capability-host: Capability discovery, staged activation, rollback to previous generation; packages/execution-supervisor: Process profiles with minimal env, leases with epoch fencing; packages/host-adapters: Credential vault backends, secure input, OS capability probes; packages/integration-sdk: OAuth PKCE helpers, API-key custody, scope checks, readiness probes; packages/mcp-adapters: MCP tool/auth adapter and MCP Apps host bridge seam; packages/voice-adapters: Provider-neutral voice events, Gemini Live and TTS adapters; packages/signal-sources: Verify and normalize incoming signals, GitHub first; packs/browser-playwright: Browser Use: target, observation and action contract; needs a downloaded browser engine; packs/computer-linux-desktop: Linux virtual desktop: session and lease contract; needs a display server; packs/computer-macos: macOS native driver: input lease and consent; needs Accessibility permission; packs/data-canvas: Chart, table, map and timeline widget descriptors over dataset refs; packs/google-calendar: Reference integration; live access needs a registered OAuth client; packs/project-work: Read-only file questions and a controlled code task.



Responsibilities condensed from each app, package and pack package.json description.

Five apps, eighteen packages, and six first-party packs. The line I keep defending is the one between core and everything else. packages/core holds the invariants: the task state machine, effect reconciliation, policy and grant intersection, leases, the capability registry. Domain features live in packs (browser, calendar, project work, desktop drivers) and reach the core through contracts, never around it. A pack can contribute tools, widgets and setup recipes. It can't bring its own scheduler, its own conversation store or a hidden root account.

## What doesn't work yet

ClarkCant is a bootstrap, not a release, and the README says that before anything else. It has a section literally called “What does not work yet”. Live OAuth, the Google Calendar integration, the macOS and Linux native drivers, downloading a real package from git or npm, and a real consent flow that grants an installed package its capabilities are all on it. Their contracts, state machines and refusals are implemented and tested. The transports and the end-to-end flows aren't.

docs/conformance-traceability.md is the ledger I trust more than any roadmap. Each of the 18 scope items has a status backed by a named test, and pnpm invariants fails the build if the ledger and the machine-readable registry disagree. Today 10 scope items are PASS and 8 are PARTIAL, and every PARTIAL row says which layer is missing. All 77 acceptance tests are PASS, with the doc's own warning that having a test doesn't prove the current run passed.

Status board from docs/conformance-traceability.md. PASS (10): V01 Conversation client, V02 Portable runtime, V03 Persistent task/session runtime, V04 Trusted node linking, V05 Remote collaboration, V06 Workspace registry, V07 Capability platform, V11 Rich built-ins, V16 Onboarding/personalisation, V18 Operations/security. PARTIAL (8): V08 Conversational install, V09 Credential/auth setup, V10 Reference integration (Google Calendar), V12 Custom widgets, V13 Pins, V14 Browser Use, V15 Computer Use, V17 Live voice.



Status column of docs/conformance-traceability.md, 5 October 2026.

I'd rather you read that board than discover it. The repository's first commit is dated 16 September 2026, and it's 5 October as I write this. It moves fast, and I'd rather it move fast in the open.

If you take one idea away: in ClarkCant the model proposes and the host decides. The model picks words, tools and views. The host owns identity, permission, secrets, task state, what counts as success and what gets drawn as a trusted card. Everything in this post is that sentence applied over and over.

I might be wrong about where some of those lines sit. The code is Apache-2.0 at github.com/digitopvn/clarkcant. Read it, break it, tell me.
