Skip to main content

What is a Published Agent?

A Published Agent is a tenant-owned route that exposes one Agent Profile over a public HTTPS surface. Publishing does not change the profile itself: it adds a row that maps (tenant, slug) to the profile and tracks the public knobs that the Chat Gateway enforces on every request. Published agents are:
  • Slug-addressed: the public URL is https://<gateway>/v1/chat/{tenant}/{slug}
  • Profile-backed: every chat turn runs against the same profile that conductor users invoke privately
  • API-key authenticated: bearers are HMAC-hashed with a pepper; plaintext is shown exactly once
  • Soft-deleted: unpublishing keeps the row for history, conditional unique indexes on unpublished_at IS NULL
  • Tenant-scoped: Postgres row-level security partitions every read and write

Public API Access

Manage published agents through the Conductor API. Send chat messages to the agent’s publicUrl using a key issued for that published agent. See the Chat Gateway API for requests and responses.

Data Model

Six Postgres tables back every public chat turn. All six are RLS-scoped by current_setting('app.tenant_id').

Soft-Delete and Uniqueness

published_agents uses one partial unique index, published_agents_tenant_slug_active, on (tenant_id, slug) WHERE unpublished_at IS NULL. There is no database-level unique constraint on (tenant_id, profile_name): at most one active row per profile is an application-level invariant only, not enforced by the schema. Unpublishing stamps unpublished_at so the slug can be re-used immediately while history stays intact.

LISTEN / NOTIFY

published_agents and agent_api_keys carry triggers that emit published_agents_changed and agent_api_keys_changed notifications. The chat gateway currently reads on every request; the route cache wired through these channels is a planned follow-up.

API Key Identity

The gateway never sees plaintext keys after issuance. The full lifecycle: Tokens follow the format ao_<env>_<base32-22> where <env> is an operator-configured deployment tag (1-16 lowercase alphanumeric characters, read from AGENT_ORC_ENV, defaulting to dev when unset) rather than a fixed set of values. The base32 tail is 22 chars of random entropy. Reaction surface:
  • Revoke: DELETE /api/profiles/{name}/keys/{id} stamps revoked_at; the gateway returns 401 unauthorized on the next request without leaking the key id
  • Rotation: issue a new key, hand it to the client, then revoke the old one
  • Lost pepper: every existing key becomes unverifiable; treat this as a full rotation event

Conversations and Public Runs

A conversation_id is the only continuity handle on the public API. The first request mints one; subsequent requests pass it back to reuse the runtime session.

Why Conversation, Not Session?

The public surface uses conversation_id instead of exposing session_id for three reasons:
  1. Encapsulation: Conversations are the public concept; sessions are a runtime detail that may be recycled
  2. Cross-agent isolation: published_id is stamped on the conversation row at creation, and a stolen conversation_id cannot be reused against a different published agent
  3. One identity surface: Surfacing both ids would let callers desync them; the gateway resolves session reuse via the conversation row, not via wire input

Watcher and Public Run IDs

Every chat turn opens one in-process watcher goroutine inside the gateway. The watcher is the single source of truth for terminal conversation_messages writes: sync and stream handlers subscribe to it via an in-process channel, but never write the terminal row themselves. Public run ids (prun_...) are durable: even if the client disconnects, the watcher finishes the run and persists the terminal message. Callers re-attach with GET .../runs/{publicRunId} for the final answer or GET .../runs/{publicRunId}/stream to subscribe to a still-running watcher.

Ingress Metering

After a public chat request is accepted and dispatched to the conductor, the gateway increments an in-memory counter for that (tenant, published agent). The counter flushes to usage_records once per interval as kind='ingress', with published_id and the number of accepted requests since the previous flush. Ingress rows are append-only increments, not cumulative upserts. Tenant-wide usage sums ingress_requests over the selected window, while per-agent metrics filter by published_id. Phase 1 records ingress for stats only, so these rows have cost_usd = 0.

Crash Recovery Sweep

On boot, the chat gateway runs a one-shot sweep:
Rows that already have a conductor internal_run_id get a re-attached watcher so the terminal write happens even after a gateway crash. Rows that never made it past dispatch do not have an internal run to resume; the sweep finalizes them as lost. The sweep is logged and continues; a sweep failure never blocks boot.

Public Knobs

Every published row carries the request-shaping knobs the gateway enforces. The rate-limit key is the API key id, not the tenant: an abusive bearer cannot poison the whole tenant.

Tenant Boundary

A published-agent key is scoped to one tenant and agent. Requests must use the matching tenant and slug in the URL. Revoked keys are rejected.

Failure Modes

See Publishing an Agent for the publishing walkthrough, or the Chat Gateway API and Conductor publish endpoints for wire-level reference.
Last modified on September 6, 2026