Actually you don't need to read our docs, just ask Clark questions

Ask from a terminal:

clarkcant ask "how do I connect Cursor to you?"

The CLI runs from a ClarkCant checkout for now, so without an alias the same question is:

# From a ClarkCant checkout (the CLI is not published to npm yet)
node apps/cli/src/main.ts ask "how do I connect Cursor to you?"

Or, from any AI tool that speaks MCP (Claude Desktop, Cursor, …), call the ask_clark tool:

{ "jsonrpc": "2.0", "id": 1, "method": "tools/call",
  "params": { "name": "ask_clark", "arguments": { "text": "how do I connect Cursor to you?" } } }

The rest of these pages are for when you want the exact routes, frames and flags.

Built on open standards

ClarkCant is one conversation with one Clark. The same Clark is reachable through ordinary, open interfaces, so any third-party app, script or AI tool can drive it without a custom SDK:

One token, one gateway, the same semantics everywhere. Every surface is served by the same runtime node, on the same origin, with the same bearer token. There is no second business-logic stack: an MCP tool, a WebSocket request and a CLI command all end up at the same HTTP gateway handler, so they behave exactly like the REST call underneath.

Approvals are deliberately not exposed to MCP. Deciding an approval is a human surface. An AI client must not be able to approve its own guarded action, so there is no MCP tool for it.

Slash commands

Type / in the composer to see the commands, listed ahead of skills. Each one is answered by Clark in the conversation, as a message with a small card you act on right there, so nothing opens a screen of its own.

A command typed in the middle of a sentence is just text. Commands are not sent to the model. See provider sign-in routes.

Inbox and notifications

Open the inbox from its header badge, with the keyboard, or by asking Clark in text or voice: “open the inbox”. It brings together command approvals, package permissions, approvals a running task raised, open questions and stored notices about background work, expired requests and available updates. Clark's control_app tool reaches the same inbox.

Granting an approval raised by a running task lets the node run it again; if the node cannot accept the run, Clark says nothing has run and you can ask again. Denial or expiry ends the task. That decision belongs to you; machine interfaces refuse it. See task approval decisions.

Settings → Control offers desktop OS notifications and browser Web Notifications, with per-group toggles and quiet hours. The browser must grant notification permission. Notifications appear only while the app is hidden, unfocused or in orb/compact mode, and their redacted text never includes command lines. Clicking one opens the inbox on its item; an item already dealt with is reported as such.

If the OS refuses or fails to show a notification, an inline status beside the OS notification toggle explains why and says the item remains in the inbox. It clears after a notification is shown again. Update checks offer only versions the installer accepts on your host and stay silent while the node is offline.

The desktop window

The desktop app reopens the conversation window where you left it: same position and size, maximized or full screen if it was. If that display is no longer attached, or the window would open out of reach, it opens at the default place and size instead. A window remembered from a larger screen is shrunk to fit.

Labels in your interface language

Task cards name their status and evidence in the interface language, for example “needs your decision” or “outcome unknown”; a status this version does not know is shown as sent. In the widget library, each widget's source (Built-in, Installed package, Local development package) and status (Stable, Experimental) are worded the same way. Settings → Extensions lists each tool by what it does, with its name beside it, under “On this node” and “The agent's own (pi)”; the instructions the model is given for a tool are behind “What the model is told”.

Terminal in the conversation

Ask Clark in text or voice to “open a terminal in the project folder” to open a real shell in the conversation. Full-screen programs such as vim and htop work there. A command Clark proposes follows your execution policy; when it needs your approval it is prefilled and runs only when you press Enter. Clark does not type over a line you are editing. Press F6 to move focus out of the terminal.

Send the selection, the last finished command's result or the screen back to the conversation; the button says which will be sent, and credential-like strings are redacted before the model sees them. A terminal has one driver at a time: other cards observe it and offer “Drive here”. Its Progress panel lists other terminals, command invocations and background work. The card reports connecting, disconnected with reconnect, ended with an exit code, gone from the node or unavailable because the node has no PTY.

Settings → Widgets → Browse lists Terminal under System cards with a description, a static illustration labelled as an illustration and a hint to ask Clark. It has no live preview or open button. Search works with or without Vietnamese tone marks; choosing a family other than All hides System cards. The /terminal socket belongs to this app card.

Kanban boards in the conversation

When a reply has a board, move a card with a mouse, touch or keyboard. The built-in board is a view of work Clark already knows; it does not connect to a project service by itself. If a board is bound to change data elsewhere, that request goes through Clark's normal action and policy path. See the widget developer standard (§8.10).

Tree view

When Clark already has hierarchical information, it can appear as a nested outline using canvas.tree@1. Each item has a stable id and label, with optional secondary text, a closed-set icon and children. This is a view of known information, not a file browser: the tree reads no service and changing selection or expanding a branch does not edit source data. Its selection and expanded branches are saved with the widget and restored with its pin; saved ids absent from a later tree are ignored.

Arrow keys move through visible items and open or close branches; Home and End move to either end, type-ahead finds a label, and Enter or Space selects an item. Screen readers receive each item's level, position, set size, selected state and expanded state. The node validates a bounded hierarchy before it is shown. See the tree contract and bounds and the internal widget development guide.

Diagrams and graphs

When Clark has a flow, a dependency graph or a small tree to show, it can draw it with canvas.diagram@1. Your app draws the SVG itself and nothing in it runs: labels are text, shapes and lines are numbers from a shared layout, and there is no HTML, foreignObject, link, image, style or handler a label could name.

A diagram holds at most 60 nodes and 120 edges. A node id is at most 64 ASCII letters, digits, _ and -. A label is one line of at most 80 characters, an edge label at most 40, a group name at most 40 and the title at most 200. A node's shape is box, round, diamond or circle. An edge's direction is forward, both or none, and it can carry a label. Groups are one level deep. The node refuses a repeated node id, an edge to a node that is not there, an edge from a node to itself, the same edge twice, too many nodes or edges, an unknown field and a hidden character, and each refusal says why in the host's own sentence.

layout is layered (the default) or tree, and direction is TB (top to bottom) or LR (left to right). The layered layout breaks cycles, layers nodes by longest path and bends long edges through the layers they cross. The tree layout centres children under their parent and refuses a graph that is not a forest. A gap that holds an edge label is widened to fit it, and edges joining the same two nodes are drawn apart. The layout is deterministic: the same props give the same drawing on the node and in every client, and the largest accepted graph is laid out in bounded time.

Mermaid flowcharts

Clark can also hand show_view a Mermaid flowchart: { "mermaid": "flowchart LR ...", "title"?, "layout"? }. The node reads a documented subset of the flowchart syntax itself and keeps only the resulting diagram model. It never stores the Mermaid source and never loads Mermaid's renderer. The subset is:

A source naming more than 60 nodes or 120 links is refused as soon as it passes the limit. Everything that configures or extends Mermaid's renderer is refused by line, with the reason: click, href, call, style, classDef, class, :::, linkStyle, %%{init}%% directives, front matter, HTML or entity codes in labels, Markdown strings, fa: icons, other node shapes, other link styles, & chains, nested subgraphs and other diagram types.

Selection and keyboard

Choosing a node sends diagram.select with { selectedId }. The node checks it against the current diagram and keeps it as widget state, so the selection survives a reload; a restored id the diagram no longer has is ignored.

Each node is a button, and one node is in the tab order. The key along the flow (Down for TB, Right for LR) follows an edge forward and the opposite key follows one back. The cross keys step within a layer, and Home and End go to the first and last node. Enter or Space selects the focused node, or clears it if it is already selected, and Escape clears the selection. A node's accessible name gives its label, shape and group, and the nodes it leads to, comes from and is linked with, each with the label of the edge between them (“leads to Ship (yes)”). The selected node marks its edges and neighbours with stroke width and dashes as well as colour, and a live region announces the change. A wide drawing keeps its size and scrolls inside its card, so the page never scrolls sideways. The diagram has no motion of its own and follows both themes.

The semantic document gives the node and edge counts, the layout and direction, and the selected node with its neighbours and their edge labels. A text alternative, also shown under the drawing, lists every node with its outgoing edges. See the Mermaid subset parser and the widget developer standard (§8.12).

Current limits. When two labelled edges leave the same node into the same gap, their labels can sit close together; labels always stay clear of nodes. Like the tree, a diagram cannot yet be placed as one part of a view Clark lays out.

Maps

When Clark has places to show, it can draw them on a map with canvas.map@1: points with labels, lines such as routes, and areas. The map is drawn over an offline basemap that ships with the app, so with the default settings a map makes no request beyond your node and nobody learns which part of the world you are looking at.

The basemap is Natural Earth 1:110m land, version 5.1.2, which is in the public domain. It is generated by a build script that pins the source file's SHA-256, and the map credits Natural Earth under every drawing.

What a map holds

Positions are WGS84 [longitude, latitude] in a strict subset of GeoJSON geometry: Point, LineString and Polygon, an outline with up to 15 holes, each ring closed. A map holds at most 200 features and 5,000 positions in total. An id is at most 120 characters, a label 120, a description 300 and the title 200. The map fits its features, or an optional view of { center, zoom }, with a whole zoom from 0 to 18, replaces the fitted view. Geocoding, routing, your device's location, vector tiles, 3D, clustering and editing are not part of this widget, and the map never asks for your location.

The node refuses a map before it is shown, and says why: coordinates outside the globe, an unknown geometry type, too many features or positions, over-long or repeated text, hidden characters, and any URL. A key that names a link (url, href, src, tile, endpoint and the like) or a value that is one (https://, //, data:, javascript:, blob:) is refused, so a map's props can never name a host. Places Clark puts on a map are Clark's statement, and the map carries no “live” badge.

Map tiles, only if you choose a provider

Map tiles. Maps draw an offline basemap and work fully without tiles, which are off by default. To add provider tiles, open Settings → Extensions → Map tiles and enter the provider's origin, tile path, attribution and maximum zoom. This sets your node's tile policy, the preference maps.tilePolicy, which names one provider. If the provider needs a key, enter it with the header or query parameter it goes in. The key is bound to the origin you entered it for: ClarkCant sends it only to that origin and never shows it again. You can also ask Clark to turn tiles on or off. Clark follows your execution policy: in Autonomous mode the change happens at once and can be undone in Settings; in Ask every time mode you approve it on a card. In Guarded mode, naming a provider asks on a card and turning tiles off happens at once, and Deny everything refuses both. Clark can never move your key. If Clark sets a provider that needs a key at a different origin, the map stays offline until you enter the key again for that provider. GET /map-tiles reports why maps are offline in its offline field: no-provider, key-unavailable (no usable key saved) or key-origin-mismatch (the saved key belongs to another origin).

AI clients over MCP, the WebSocket relay and clarkcant api cannot write the policy or the key directly: those routes answer 403 PERSON_ONLY, and a client that wants tiles changed asks Clark. The routes are listed in the node's API.

Selection, keyboard and the table

Choosing a feature sends map.select with { selectedId }, empty to clear it, and moving the map sends map.view with { center, zoom }. The node checks both against the current map and keeps them as widget state, so the selection and the view survive a reload. A pan is saved 400 ms after it settles, so a run of key presses is one write. In a view Clark lays out, map.select is also an event with a selectedId field that other parts can follow.

The map can be focused and used by keyboard:

Dragging with a pointer pans too. On a touch screen, a swipe over a map you have not tapped scrolls the conversation; once you tap the map, a drag pans it. The zoom-in, zoom-out and reset buttons are 44 px. A live region states the view once it settles, and the selection. A pan slides only when motion is allowed; with reduced motion the view moves at once. The map follows both themes, and a narrow map moves its controls under the picture and draws its labels larger, without overflowing at 390 px.

Below the map, a table lists every feature with its kind and position. Its Select button selects the feature on the map and brings it into view, and selecting on the map highlights its row. That table is also the map's text alternative.

The semantic document gives the feature count and the count of each kind, the visible bounds and zoom, the selected feature's label and coordinates, and whether tiles are offline or from the policy's origin. See the map contract and bounds and the widget developer standard (§8.13).

Current limits. The map is one world that does not repeat: the view stops at the antimeridian, and a map does not wrap across the 180° meridian. The basemap, features and tiles are each drawn once, so they always agree.

Audio player and document preview

Clark can play a sound with canvas.audio@1 and show the text of a PDF or text file with canvas.document@1. The page never fetches anything for either widget. Clark names a source, and your node reads it, checks it under its media content policy and stores only what it checked.

Where the sound comes from

Clark names exactly one source, a title and, if it has one, a transcript of at most 4,000 characters:

Your node fills in the reference the page plays, the type, the length and the size, and for a fetched file the origin it came from. Clark cannot supply any of these, so a player never claims a type or a length nobody checked. The page only ever sees a reference to a file your node holds, never the outside URL. The same checks of type, size and length apply to an audio file you already hold.

Playing

The player uses the browser's own controls, keyboard included: focus it and press Space to play or pause. It never plays by itself, neither when it first appears nor when it comes back. It keeps whether it was playing, where it was and how long it is, written the same way as a video's, so after a reload it opens paused where you left it. A transcript, when there is one, is plain text in a section you can open with the keyboard, and a hidden character in it is drawn as a visible marker. Voice, the note added to your next turn and inspect_ui know the title, whether it is playing, its position and its length, and whether it has a transcript.

Previewing a document

Clark names a PDF or a text file in this conversation (plain text, Markdown, CSV, tab-separated values or JSON), by its artifact or attachment id. Your node reads its text, a PDF's through your node's own reader, and keeps at most 20,000 characters in at most 10 pages of at most 2,000 characters each, breaking at a line end or a space. A notice says when the preview was cut short and how much of it is shown. The preview is text only: pictures and layout are not shown, and nothing in the file is treated as markup or run. A hidden character is drawn as a marker, with a warning.

The page's text sits in a region a keyboard can reach. Tab moves to the Previous and Next buttons and Enter turns the page. At the first or last page the button stays focused and is announced as unavailable, a new page opens at its top, and the page number out of the total is announced politely. The page you are on is kept with the widget, so it is still open after a reload, and voice, the next-turn note and inspect_ui know it. A WAV is not a document, and the preview refuses it.

Both widgets follow both themes and fit at 390 px. With reduced motion they add no animation or transition. Neither can yet be placed as one part of a view Clark lays out: each is shown on its own, where your node checks its source. The Widget Library shows sample cards; the audio sample says it cannot be played rather than pretend.

The media content policy

Your node fetches audio from the web only from origins you list in CC_MEDIA_ORIGINS, a comma-separated list of bare https origins, at most 32. It is empty by default, so your node fetches nothing until you name an origin. A list with any entry that is not a bare https origin is ignored as a whole: your node allows nothing and says why once at startup. Each fetch follows these rules:

Every refusal ends with the rule it broke, as (media policy rule: <rule>), so Clark can tell you what to change. A file that passes counts against your storage. The page's own content policy is unchanged. See the rule names, the media content contract and the widget developer standard (§8.14 and §14.5).

Current limits. When a conversation loads, the page reads the whole file of every placed player, even though it plays nothing until you press play; reading it only when needed is tracked in #403. Each saved playback position is a full widget-state write, a cost tracked in #380.

Pictures and video in the conversation

A gallery or carousel keeps the picture you chose. The choice is stored with the widget, so it is still selected after a reload or a pin restore. A local video keeps whether it was playing, where it was and how long it is. When it comes back it opens paused at that position and never starts playing by itself.

Clark knows what you are looking at. Voice, the note added to your next turn and Clark's inspect_ui tool describe the selected picture (“showing picture 2 of 3”, with its alt text) or the video's playback state. While a video plays, its position is saved at most once every three seconds, and at once when you pause, seek or it ends. Saving when you leave the page is best-effort, so a “playing” state older than eight seconds is reported as paused at the last saved position. If your machine refuses a save, the widget says why beside it: a gallery or carousel shows the stored picture again, and a video stays where it is.

A view Clark lays out can include a gallery or carousel of your own imported pictures, newest first. Clark names only the widget and how it is wired; pictures it names are ignored. Choosing a picture sends media.select with { selectedIndex }, counted from 0, so another part of the view wired to the same value follows it and the view's state reports it.

The app's page and desktop window policy allow media-src 'self' blob:, only so a video can play from an object URL the app creates from bytes it fetched with your node's token. No remote media origin and no data: media are allowed.

Current limits. Your node cannot import video files yet, so a local video plays only from a reference your node already serves. A composed gallery or carousel keeps the pictures that were there when the view was laid out; pictures imported later appear only in a newly composed view. Each playback save is still a full widget action; reducing that cost is tracked in #380. See the widget developer standard (§8.11).

Quickstart

  1. Run a node. The installer sets one up; from a checkout it is node apps/runtime/src/main.ts. It listens on http://127.0.0.1:8765, loopback only by default.
  2. Find the token: localToken in <data-dir>/identity.json. The default data dir is ~/.clarkcant; in Docker the data dir is /data (token in /data/identity.json).
  3. Check the node and discover its surfaces. These three routes need no token:
export CLARKCANT_URL="http://127.0.0.1:8765"
export CLARKCANT_TOKEN="<token>"   # localToken from <data-dir>/identity.json
# Public: no token needed
curl -s "$CLARKCANT_URL/health"
curl -s "$CLARKCANT_URL/.well-known/clarkcant.json"
curl -s "$CLARKCANT_URL/openapi.json"

/health reports liveness, the runtime platform and the negotiated protocol, and no node identity. /.well-known/clarkcant.json lists every surface (api, mcp, websocket, cli) and its endpoints. /openapi.json describes the stable REST surface.

Then send Clark a first message:

# 1. Create a conversation → 201 { "conversationId": "…", "homeNodeId": "…" }
curl -s -X POST "$CLARKCANT_URL/conversations" \
  -H "Authorization: Bearer $CLARKCANT_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "title": "From my app" }'

# 2. Talk to Clark in it
curl -s -X POST "$CLARKCANT_URL/conversations/<conversationId>/messages" \
  -H "Authorization: Bearer $CLARKCANT_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "text": "What can you do on this machine?" }'

Authentication and errors

{ "code": "UNAUTHENTICATED", "message": "…" }

Keep the token private. It gives full access to your Clark. The node binds to loopback by default; --allow-public-bind exposes it and needs TLS in front. Sign-in with OAuth is not available: the bearer token is the only credential.

How Clark picks context

These environment variables on the node choose what context Clark sends the model with each message:

What goes to the Jev provider (only with CLARKCANT_CONTEXT_DECIDER=jev, and only when it is asked): the message being answered, redacted, at most 300 characters (400 for a tool-family choice), plus each candidate memory note or earlier message, redacted, at most 200 characters. Nothing else leaves the node, and Jev grants nothing.

What a model may be sent

Every block your node adds to a model's context is labelled public, internal, confidential or secret, from its text alone and before anything is ranked or sent:

A model profile may say what it receives with trustClass (local, first-party, approved-third-party or untrusted) and allowedDataClasses. By trust class, local may receive everything, including secret; first-party and approved-third-party up to confidential; untrusted only public. With both set, the profile receives only what both permit. A list with no trust class is taken as written, so it can add secret to an unlabelled profile. A profile with neither receives everything but secret, and a model several profiles name receives only what all of them permit.

With CLARKCANT_CONTEXT_PLANNER on, a block above the answering model's ceiling is left out of the recap, the memory brief, the earlier messages, the context a background run or task worker retrieves (including its read_context reads), project instructions, and the results of search_history, each judged on the whole stored entry, not only the snippet shown. Each says how many blocks were withheld, never what they said, and the recap does not point at a tool to read them back. Jev is offered only public and internal candidates.

Background work leaves out a model profile that may not receive the work's class (the class of the request or the task goal), whether the planner is on or off. When no profile qualifies, the work still runs on the node's configured model: the request or goal text itself is sent to it whatever its class, and only the retrieved context is narrowed to what that model may receive. When the class is what left every profile out, the node writes one JSON line to stderr (model-route, with the class and a count only), and the task's audit record says the work ran on the node's model because no model in the pool may receive that class; it names the class, never the goal. When routing itself fails, the work also runs on the configured model, with a model-route line marked route-failed that carries no error text.

Scope. The ceiling covers context your node adds, and only that. It does not cover tool results the model reads itself (a file, a command's output, a web page), attachments the person sends, or what a session that continues already holds from earlier turns; it is not a boundary against those, and a model that must never see a class of data should not be given tools that can read it. CLARKCANT_CONTEXT_PLANNER=off turns withholding off, search_history included, so secret-shaped recap, memory and history text reaches whichever model answers; routing by data class still applies with the planner off.

Project instructions

A project can keep instructions that apply only to part of it. Your node reads them only from a project inside a folder you already granted: .clarkcant/instructions.json holds up to 32 rules, and each included name is the file .clarkcant/instructions/<name>.md.

{
  "version": 1,
  "rules": [
    {
      "when": { "path": "src/payments/**", "operation": "write" },
      "include": ["payments"],
      "pin": false
    }
  ]
}

Instructions reach the model as project data: a header says they come from the project's .clarkcant files, are neither your words nor your node's, and grant nothing. Each is wrapped in a tag carrying a code your node draws once per session (a rebuilt or handed-off session gets a new one) and states once, in its own guidance at the start of that session, as the only code that marks project guidance; a task worker's brief names its own code. An instruction cannot close its own block, and a block copied into a file or a tool's output carries a code your node never stated to that session, unless the session itself repeated it. This is a signal the model reads, not an enforced boundary: the execution policy decides what any action may do, and instructions grant nothing either way. There is no opt-in step: a project's instructions apply because you granted its folder, and every effect still goes through the execution policy. Instructions are labelled like any other block and withheld above the model's ceiling. Limits: 8 names per rule, 4,000 characters per instruction, 6,000 per turn and a 64 KB rules file. A file your node cannot use is reported once on stderr (instructions-invalid, with the folder name and a reason: too-large, not-json, shape or unknown-version) and changes nothing. CLARKCANT_CONDITIONAL_INSTRUCTIONS=off (default on) turns this off.

When a conversation gets a fresh session

A conversation keeps one model session until it fails, is evicted or changes model. CLARKCANT_SESSION_POLICY (default off) decides at each turn whether to keep it:

A rebuild starts a fresh session briefed by the planned recap, restates instructions and tools, and leaves the conversation itself untouched. If the fresh session cannot be created, or the policy step fails, the turn continues on the old one. A message sent during a rebuild waits for it and lands in the new session, and Stop during a rebuild discards the new session and sends nothing. Any other value is off. The thresholds are judgements from a scripted cost simulation, not live measurements.

When you change the model

Changing the model never cuts off a reply Clark is writing: the change applies from your next message, which starts a new session on the new model, briefed with a recap of the conversation. Clark never answers on the previous model instead. If the switch fails, the turn says so in your language and the conversation is left as it was, for example: “Could not switch this conversation to {provider/model}: {reason}. Your message is saved and the conversation is unchanged. Retry, or choose another model.” Local file paths are removed from the reason. If getting ready to answer takes longer than the turn's time budget, the turn stops with “Could not get ready to answer within {N} s, so this turn was stopped. Your message is saved. Retry, or choose another model if this keeps happening.”

How long a turn may run

By default a turn has no time limit: it runs until it is done, and you can press Stop at any time. To cap it, choose 5, 15, 30 or 60 minutes in Settings → AI & Routing → Thinking and time. An operator can also set CC_MODEL_MAX_WALL_CLOCK_MS. A turn that reaches its cap stops and says so in minutes; your message is kept.

Where to next