Context & knowledge¶
A context document is a file or URL an agent draws on: product docs, an FAQ, a persona brief, pricing information. Documents come in two scopes: per-agent (visible to one agent only) and global (merged into every agent's system prompt automatically). Every write operation on either scope is an admin action. See Concepts overview for how context fits into the rest of AiFlow's model; this page is the endpoint reference for managing it.
Admin role required, and it differs by scope
Reading either scope (GET) needs any logged-in admin. Writing
needs more, and the two scopes are gated differently: know this
before you build against one and assume the other works the same
way.
| Scope | Write role required (upload, import, patch, delete, resync, reindex) |
|---|---|
Per-agent (/api/v1/agents/{agent_id}/context-docs) |
Owner, Admin, or Editor |
Global (/api/v1/context-docs) |
Owner or Admin only |
Global documents land in the system prompt of every agent in the deployment at once. That blast radius is why they're held to a higher bar than a single agent's own documents.
Supported sources¶
| Source | How | Notes |
|---|---|---|
.txt / .md upload |
POST .../context-docs (multipart) |
2 MB limit. |
.pdf upload |
POST .../context-docs (multipart) |
10 MB limit. Text is extracted at upload time; the original binary is discarded, only the extracted text is stored. Password-protected PDFs and scanned images with no text layer are rejected with a 422. |
.docx upload |
POST .../context-docs (multipart) |
10 MB limit, same text-extraction behavior as PDF. |
| Google Doc | POST .../context-docs/import-url |
Must be shared "Anyone with the link can view"; AiFlow fetches its plain-text export, no OAuth or download/re-upload round trip. |
| Any public web page | POST .../context-docs/import-url |
Main content is extracted automatically (navigation, footer, and ad chrome stripped); falls back to a plain text dump if extraction finds nothing usable. Fetch is capped at 10 MB. |
| Bulk URLs or a sitemap | POST .../context-docs/import-urls |
See Bulk import below. |
Licence caps¶
Two caps apply on top of the per-file size limits above: how many documents the deployment may hold, and how many megabytes they may total. Both count global and per-agent documents together against one deployment-wide budget, so where a document lives makes no difference to either.
Exceeding a cap returns 402 naming the limit. The size cap is checked
against what the upload would bring the total to rather than against the
total alone, since a deployment comfortably under its budget can still
cross it with a single large file. A bulk import is checked per document,
so a batch that crosses a cap partway imports what fits and reports the
rest against the URLs that did not.
Re-syncing a URL-sourced document is checked differently again: it replaces something that already exists, so it never counts against the document cap, and only the change in size counts against the megabyte cap. A deployment at its document cap can still refresh what it already has.
Legacy binary .doc (pre-2007 Word) is rejected outright, with a message
telling you to re-save as .docx or export as PDF; no parser for the old
binary format is bundled. A generic URL import also refuses to fetch a
private, loopback, link-local, or otherwise internal address (http/https
only), so it can't be pointed at your own internal network.
Uploading a document¶
Multipart, one file per request, using the field name file:
Returns 201 Created with the document, in the shape shown under
Response shape below. Indexing for search starts in the
background straight away, and the response does not wait for it.
The global (shared) equivalent¶
Identical multipart shape and response shape, just without {agent_id} in
the path, and gated to Owner or Admin instead of Editor and above (see the
note above). A global document is merged into every agent's system
prompt, ahead of that agent's own documents. Use it for facts every agent
should know, like company name, support hours, or brand voice, instead of
re-uploading the same file for each agent.
Importing from a URL¶
const response = await fetch(
"https://api.your-domain.com/api/v1/agents/1/context-docs/import-url",
{
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${adminToken}`,
},
body: JSON.stringify({
url: "https://docs.google.com/document/d/1AbCdEfGhIjKlMnOpQrSt/edit",
}),
},
);
const document = await response.json();
The same endpoint accepts any other public web page URL, not just Google Docs; main content is extracted the same way described in Supported sources above. AiFlow detects a Google Docs URL automatically, by pattern rather than a separate flag, and routes it through the export path instead of a raw page fetch.
Bulk import: a list or a sitemap¶
| Field | Type | Description |
|---|---|---|
urls |
list of string or null | An explicit list of pages to import. |
sitemap_url |
string or null | A sitemap.xml URL; every <loc> entry in it is imported. |
Provide at least one of the two. Sending both is fine; they're combined.
Omitting both returns 400.
const response = await fetch(
"https://api.your-domain.com/api/v1/agents/1/context-docs/import-urls",
{
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${adminToken}`,
},
body: JSON.stringify({ sitemap_url: "https://example.com/sitemap.xml" }),
},
);
const results = await response.json();
Returns 200 OK with one result per URL. Successes and failures appear
together in the same list, so one bad page doesn't fail the whole request:
[
{
"url": "https://example.com/help/shipping",
"status": "ok",
"detail": "Imported",
"document": { "id": 12, "filename": "shipping.txt", "...": "..." }
},
{
"url": "https://example.com/help/broken-page",
"status": "error",
"detail": "Could not fetch https://example.com/help/broken-page (HTTP 404).",
"document": null
}
]
Refreshing a document: resync vs. reindex¶
These solve two different problems and are easy to mix up:
| Endpoint | What it does | When to use it |
|---|---|---|
POST .../context-docs/{doc_id}/resync |
Re-fetches the document from its original source_url and overwrites the stored text in place. |
The source page changed and you want the agent's copy to catch up. Only works on a URL-sourced document; a plain file upload has no source_url and gets a 400 if you try. |
POST .../context-docs/{doc_id}/reindex |
Re-runs chunking/embedding on the text already on disk, no re-fetch. | A document's chunking_status came back failed and you want to retry indexing without touching the content itself. |
Both endpoints kick off re-indexing in the background and return
immediately with the document's current, pre-reindex state. Poll GET
.../context-docs and watch chunking_status if you need to confirm
completion.
Listing and updating documents¶
import httpx
response = httpx.patch(
"https://api.your-domain.com/api/v1/agents/1/context-docs/12",
headers={"Authorization": f"Bearer {admin_token}"},
json={
"is_active": True,
"always_include": True,
"summarize": True,
"summary_sentence_count": 15,
},
)
response.raise_for_status()
document = response.json()
const response = await fetch(
"https://api.your-domain.com/api/v1/agents/1/context-docs/12",
{
method: "PATCH",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${adminToken}`,
},
body: JSON.stringify({
is_active: true,
always_include: true,
summarize: true,
summary_sentence_count: 15,
}),
},
);
const document = await response.json();
| Field | Type | Behavior on PATCH |
|---|---|---|
is_active |
boolean | Required on every PATCH, unlike the three fields below. Send the value you want, including the current one, if you're only changing always_include/summarize/summary_sentence_count; omitting it is a 422, and there is no "leave unchanged" default for this specific field. Toggles the document in or out of the agent's context entirely. |
always_include |
boolean or null | Omit to leave unchanged. When true, the document's full text (or its summary, see below) is always present in the system prompt rather than only surfaced via search_knowledge_base. |
summarize |
boolean or null | Omit to leave unchanged. Meaningful only alongside always_include: true, condenses a long document instead of inlining it in full. |
summary_sentence_count |
integer or null (1-500) | Omit to leave unchanged. How many sentences summarize condenses down to. |
is_active is not optional on this endpoint
Every other field here follows the usual PATCH convention: omit
it and it stays unchanged. is_active doesn't. The schema requires
it on every request, and the handler sets it unconditionally. A
client that always sends {"always_include": true} without
re-sending the document's current is_active value won't silently
deactivate it: validation rejects the request outright for the
missing required field instead. It's still an easy 422 to trigger
if you build the request body from only the fields you think you're
changing.
Response shape¶
{
"id": 12,
"filename": "acme-faq.md",
"content_type": "text/markdown",
"char_count": 812,
"is_active": true,
"chunking_status": "chunked",
"chunked_at": "2026-01-15T10:31:00Z",
"always_include": false,
"summarize": false,
"summary_sentence_count": 3,
"source_url": null,
"last_synced_at": null
}
chunking_status is one of pending, chunked, or failed; see
Knowledge base search below for what drives
it. source_url and last_synced_at are only populated for a
URL-imported document; both stay null for a direct file upload.
Deleting a document¶
Returns 204 No Content. This is a hard delete: the stored text is
removed from disk and the row is gone. There is no soft-delete or
recovery path for a context document, unlike an agent.
Knowledge base search (RAG)¶
Every uploaded or imported document is chunked and embedded automatically
in the background; there's no separate "index this" step to remember.
Upload finishes, chunking_status starts at pending, and moves to
chunked once indexing completes, or to failed (retry with reindex,
see above).
Indexing alone doesn't make a document usable mid-conversation. Two distinct paths pull a document into what the model actually sees:
always_include: truefolds the document, or its summary, directly into the system prompt on every session, with no search involved.- The
search_knowledge_basetool, enabled per agent, lets the model issue a query mid-conversation and retrieve the most relevant chunks from everychunked, active document, drawing on both this agent's own documents and the global set. A document that is active but notalways_includeis only ever surfaced this way. If the tool isn't enabled on the agent, that document is effectively invisible to it, regardless ofis_active.
Skills¶
Capped by your licence
This deployment's licence sets a skills limit, counted across native
Agents and Orchestrators together. Creating one more past that cap
returns 402; see
Licence limits.
A skill is a different kind of context: not a fact to recall, a
step-by-step procedure to follow, matched by a short trigger
description ("use this when...") instead of a filename or a search
query. It's a separate resource from context documents, Skill, not
ContextDocument, and its content lives directly in the row, not on disk,
since a skill is always plain Markdown an admin writes, never a scanned
PDF or a fetched web page.
A skill belongs to exactly one native Agent, or to one Orchestrator, depending on which of the two endpoint families below created it.
The API cannot create a skill owned by neither. Only a row written straight to the database ends up that way, and such a skill matches for every agent and Orchestrator alike. See "How a skill gets used" below.
GET /api/v1/agents/{agent_id}/skills
POST /api/v1/agents/{agent_id}/skills
PATCH /api/v1/agents/{agent_id}/skills/{skill_id}
DELETE /api/v1/agents/{agent_id}/skills/{skill_id}
An Orchestrator has the identical set, same request and response shapes, just under its own resource path:
GET /api/v1/orchestrators/{orchestrator_id}/skills
POST /api/v1/orchestrators/{orchestrator_id}/skills
PATCH /api/v1/orchestrators/{orchestrator_id}/skills/{skill_id}
DELETE /api/v1/orchestrators/{orchestrator_id}/skills/{skill_id}
Owner, Admin, or Editor for writes, the same tier as a per-agent context document.
Creating a skill¶
curl -X POST https://api.your-domain.com/api/v1/agents/1/skills \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-d '{
"name": "Refund dispute escalation",
"trigger_description": "The caller is disputing a charge on their account.",
"body": "1. Confirm the charge.\n2. Offer a refund.\n3. Log the outcome."
}'
import httpx
response = httpx.post(
"https://api.your-domain.com/api/v1/agents/1/skills",
headers={"Authorization": f"Bearer {admin_token}"},
json={
"name": "Refund dispute escalation",
"trigger_description": "The caller is disputing a charge on their account.",
"body": "1. Confirm the charge.\n2. Offer a refund.\n3. Log the outcome.",
},
)
response.raise_for_status()
skill = response.json()
const response = await fetch(
"https://api.your-domain.com/api/v1/agents/1/skills",
{
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${adminToken}`,
},
body: JSON.stringify({
name: "Refund dispute escalation",
trigger_description: "The caller is disputing a charge on their account.",
body: "1. Confirm the charge.\n2. Offer a refund.\n3. Log the outcome.",
}),
},
);
const skill = await response.json();
Returns 201 Created. Embedding the trigger description starts
immediately in the background, same fire-and-forget shape as a context
document's indexing; the response doesn't wait for it to finish.
Response shape¶
{
"id": 4,
"name": "Refund dispute escalation",
"trigger_description": "The caller is disputing a charge on their account.",
"body": "1. Confirm the charge.\n2. Offer a refund.\n3. Log the outcome.",
"is_active": true,
"index_status": "indexed",
"indexed_at": "2026-01-15T10:31:00Z"
}
index_status is one of pending, indexed, or failed. Unlike a
context document, there is no reindex endpoint: PATCHing a new
trigger_description re-embeds it automatically, since a skill has no
separate "re-chunk the same text" case a document's failed indexing run
does.
Updating and deleting¶
PATCH follows the same "is_active is required, everything else is
optional and left unchanged if omitted" convention as a context document
(see the warning above). Changing trigger_description resets
index_status to pending and schedules a fresh embedding; changing
name or body alone does not, only what a skill is matched by
invalidates the existing embedding. DELETE is a hard delete, same as a
context document.
Writing a skill that actually matches¶
Only trigger_description is embedded. The name and the body are
never searched, so a trigger has to carry the whole job of being
recognised.
Write the trigger as the situation, in the words the agent would reach for on finding itself in it, rather than as the name of the procedure:
- Matches well: "A message arrived by email and needs a reply, or the task in front of you started as an inbound email rather than a live chat."
- Matches poorly: "Email handling procedure."
One body loads in full, the closest match, and the rest of the library comes back as a list of names and triggers. That listing is what makes a wrong guess survivable: the model can see a better-named procedure and ask for it, but only if the name and trigger tell it apart from the one it just read. Give each skill a distinctly different moment, and split one that covers two.
Do not lean on the ranking to be right. Similarity between a situation and a trigger lands in a narrow band, close enough that an unrelated procedure routinely outscores the one you meant, so treat the first body as a guess and the listing as the real index. There is no similarity floor either: if an agent has any skills at all, the closest is returned however far off the query was.
Watch for catch-all endings especially. A trigger closing with something like "or any related question" attracts queries that belong to another skill, and buries what it displaced one line down in a list. Describe the act being performed rather than the territory covered. A small library of sharply-triggered skills beats a large library of vague ones.
Put the steps in the body, in order, with the reason for anything counter-intuitive. A body is read by a model that has already decided it is in this situation, so it needs the procedure rather than the justification for having one.
How a skill gets used¶
The search_skills tool, enabled per agent, lets the model
describe the current situation and get back the closest skill's full
body, the names and triggers of its other skills, or a clear "no
matching skill" message when it has none. Passing a skill's exact name
instead of a situation loads that one alone, which is how a model acts on
the listing: read the names, recognise the right procedure, ask for it by
name. This is
different from search_knowledge_base: a knowledge-base
query returns a short relevant passage a model can quote, search_skills
returns a whole procedure a model is meant to follow. The result says so
in as many words, because a tool result otherwise reads as information to
absorb rather than steps to take, and a procedure that arrives looking
like a document tends to be read and not followed. Write the body as
instructions to the agent, in the imperative, rather than as a description
of what the procedure is. There is no
always_include equivalent for a skill: it only ever surfaces through the
tool, on demand, never folded into the base system prompt, so a growing
skill library costs nothing in every session until one actually matches.
A delegated agent is the exception, and it needs no tool. A
delegate_to_agent run gives the target no tools at all, by
design, so it could never look a skill up for itself. Its active skills
are written into the delegated system prompt in full instead. That keeps a
delegation target working when its procedures live in skills, at the cost
of prompt size on every delegated call, where a live agent pays only for
the one skill that matched.
An Orchestrator has its own search_skills, enabled the same way as any
other Orchestrator tool, matching against that
Orchestrator's own skills plus any that truly are global, exactly the
same matching rule as a native Agent, just never a different agent's or
Orchestrator's own skills, each owner's library stays its own.