NEWPromptev MCP Server: one URL for every tool, every agent and your team’s knowledge. Connect Claude Code, Cursor, or any MCP client.→

Knowledge tool reference

The one tool an agent uses to reach the corpus: search_knowledge_base in Python, searchKnowledgeBase in TypeScript. Every action, argument and response shape, how the host scopes it, how to serve it over MCP, and a side by side map of the two SDKs.

Overview

The knowledge tool is one tool with actions. Instead of a separate tool for search, reading, listing and computing, the model gets one name, search_knowledge_base, and picks what it does with the action argument. One description carries the decision rules once, and discover tells the model what exists before it guesses.

There is one implementation behind every door. engine.search_knowledge_base, call_knowledge_tool and the MCP tool served by create_mcp_app all call the same function, so an answer cannot differ between them.

The model sets

action and the arguments in the table below: what to look for, which sources or documents to narrow to, how many rows.

The host sets

principals, scope, redaction and secret_key, plus optional compute and map_reduce runners. None of them is a tool argument, so a model cannot cross a tenant or turn masking off.

Every answer is a result, not an exception. An unknown action, an action this deployment cannot run, or a missing argument returns success: false with an error sentence the model can relay. A malformed document id or cursor is the exception: it raises ValueError in Python.

Calling it

From your own code, call the engine method. principals works as it does everywhere in the engine: a list of the caller's principals, [] for anonymous, or TRUSTED for a trusted in-process job. principals=None is deprecated: it means trusted, warns, and a later release refuses it. scope has no default and must be given.

Python
from context_engine import ContextEngine, TRUSTED

caller = ["group:hr"]          # from your authenticated session; [] is anonymous

answer = await engine.search_knowledge_base(
    action="search",
    query="annual leave",
    principals=caller,
    scope=["hr"],              # required: the host's ceiling of source ids
)
if not answer["success"]:
    print(answer["error"])     # relay it to the model; it is written for one
TypeScript
import { ContextEngine, TRUSTED } from "@promptev/context-engine";

const caller = ["group:hr"]; // from your authenticated session; [] is anonymous

const answer = await engine.searchKnowledgeBase({
  action: "search",
  query: "annual leave",
  principals: caller,
  scope: ["hr"], // required: the host's ceiling of source ids
});
if (!answer.success) console.log(answer.error); // relay it to the model

In your own agent loop

knowledge_tool_definition() (knowledgeToolDefinition()) returns the tool as data: name, description, input_schema as JSON Schema with required: ["action"], and annotations. Hand that to your model, then run what it asks for with call_knowledge_tool (callKnowledgeTool). readOnlyHint is false on purpose: most actions only read, but compute runs generated code.

call_knowledge_tool takes the engine, then keyword arguments: action, principals (already resolved: it never invents a caller), scope, the model arguments listed below, and the host-only redaction, secret_key, compute and map_reduce. The TypeScript form takes one object with the same snake_case keys.

Python
from context_engine import call_knowledge_tool, knowledge_tool_definition

tool = knowledge_tool_definition()
# {"name": "search_knowledge_base", "description": "...",
#  "input_schema": {"type": "object", "properties": {...}, "required": ["action"]},
#  "annotations": {"readOnlyHint": False, "openWorldHint": False}}

# Your agent loop: hand "tool" to your model, then run what it asked for.
# Take ONLY the model's arguments from the model; principals and scope are yours.
result = await call_knowledge_tool(
    engine,
    **model_args,                 # action, query, source_ids, ...
    principals=caller,            # already resolved, never defaulted here
    scope=["hr", "finance"],
)
TypeScript
import { callKnowledgeTool, knowledgeToolDefinition } from "@promptev/context-engine";

const tool = knowledgeToolDefinition();
// { name: "search_knowledge_base", description, input_schema, annotations }

// Take ONLY the model's arguments from the model; principals and scope are yours.
const result = await callKnowledgeTool(engine, {
  ...modelArgs, // action, query, source_ids, ... (snake_case)
  principals: caller,
  scope: ["hr", "finance"],
});

Spread the model's arguments first, then set yours. In the examples above the host values come after the model's, so a principals or scope key the model invented cannot win.

Arguments

Every argument the model may set. The exported schema and the MCP tool are both built from this one property table, so they cannot disagree. Only action is required by the schema; each action checks its own required arguments.

ArgumentTypeUsed byMeaning
actionenum of 13allWhich action to run. Required.
querystringsearch, query_meta, compute, map_reduce, community_summaryWhat to look for, in words, or a bare content identifier. For compute and map_reduce it is the instruction.
document_idstring (uuid)get_doc, get_chunksOne document's system id, from discover, list or a search hit.
source_idsstring[]discover, list, search, query_meta, compute, map_reduce, graph actionsNarrow to these sources. Intersected with the host ceiling: ids outside it are ignored, not refused.
document_idsstring[] (uuids)search, compute, get_docs, map_reduce, discover, listNarrow to these documents. Intersects with source_ids. Any other action refuses it.
entitystringtraverse, get_neighborsThe start: an entity id from an earlier graph answer (preferred, and the only handle that survives masking) or its name.
depthintegertraverseHops, 1 to 5. Default 2.
categoryenumtraverse, find_relatedHIERARCHICAL, MEMBERSHIP, CREATION, TEMPORAL, SPATIAL, REFERENCE, FUNCTIONAL, QUANTITATIVE.
labelstringfind_relatedExact relationship wording as extracted, for example reports_to. Narrower than category.
entity_typeenumfind_relatedPERSON, ORG, PRODUCT, LOCATION, REFERENCE, TEMPORAL, CONCEPT.
top_kintegersearchPassages to return. Default 10.
mode"hybrid" | "graph"searchHow to rank. Omit it and the engine decides from the documents in scope.
limitintegerdiscover, list, query_meta, map_reduce, graph actionsHow many rows. Defaults per action are listed with each action below.
cursorstringdiscover, listA previous page's next_cursor, sent back verbatim. An object is also accepted where a transport pre-parses JSON.
startintegerget_chunksFirst chunk position, 0-based.
endintegerget_chunksLast chunk position, inclusive. At most 25 chunks per call.
max_charsintegerget_docsTotal text budget. Default 200,000.

A search hit's document_id is a system id. Content identifiers written inside documents, such as invoice numbers, belong in query, never in document_id.

The 13 actions

The order below is the order of KNOWLEDGE_ACTIONS. Examples show the engine call; the model sends the same arguments as JSON. Response shapes are abridged: ... marks elided content, and // comments are explanations, not part of the payload.

discover

Needs: Always available

When to use it. First, whenever the agent does not already know what is in scope. One call returns the documents, what is inside each one and which actions this deployment can run, so later calls can use real sheet, column, section and field names.

ArgumentTypeNotes
actionstring"discover"
source_idsstring[]Narrow to these sources (intersected with the ceiling).
document_idsstring[]Narrow to these documents. Also asks for their structure in full, past the cap a whole-page discover applies.
limitintegerDocuments per page. Default 50, clamped to 1 to 200.
cursorstringThe next_cursor of a previous page, sent back verbatim.
Python
answer = await engine.search_knowledge_base(
    action="discover",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "discover",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "documents": [
    {"id": "...", "name": "costs.xlsx", "source_id": "finance",
     "kind": "spreadsheet",            // "spreadsheet" or "text", decided by MIME
     "document_type": "...", "mode": "hybrid", "failure_reason": null,
     "structure": {...}}               // sheets / sections + last_page / keys / chunks
  ],
  "count": 50, "has_more": true,
  "next_cursor": "...", "next_page": {"action": "discover", "cursor": "..."},
  "document_types": [...],             // census of the whole scope
  "fields_by_type": {...},             // extracted field names, types, counts
  "available_actions": {"search": "...", "compute": "... (OFF)", ...},
  "next_action": {"for a clause or a fact": {"action": "search", "query": "..."}, ...}
}

A spreadsheet structure lists sheets with name, column headers, row count and frame_key (the name compute will give that sheet). A document with headings lists sections and last_page, JSON lists keys, anything else falls back to chunks.

Lists inside a structure are capped, with the rest reported as more_columns, more_sections, more_keys or more_sheets. Naming the documents in document_ids returns them whole.

fields_by_type and document_types are computed over the whole scope, not only the current page.

available_actions covers the other twelve actions. An action this deployment cannot run is listed with " (OFF)" appended, never hidden.

Needs: Always available

When to use it. Questions answered by reading text. To look up an ID, code, SKU or invoice number, send the bare identifier alone (for example "2525"), not the whole question.

ArgumentTypeNotes
actionstring"search"
querystringRequired. Words, or a bare content identifier.
source_idsstring[]Narrow to these sources.
document_idsstring[]Narrow to these documents (intersects with source_ids).
top_kintegerHow many passages. Default 10.
mode"hybrid" | "graph"Omit to let the engine decide from the documents in scope. "hybrid" forces wording only, "graph" forces the connection signal on.
Python
answer = await engine.search_knowledge_base(
    action="search",
    query="annual leave carry over",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "search",
  query: "annual leave carry over",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "hits": [
    {"document_id": "...", "document_name": "HR leave policy", "chunk_text": "...",
     "score": 0.83, "source_id": "hr", "chunk_idx": 4}
  ],
  "usage": {...},
  "hint": {"action": "compute", "reason": "aggregate question over a spreadsheet"}  // only on a mixed scope
}

chunk_idx is where the passage sits in its document. Pass it to get_chunks as start to read around it.

A query that names a file (for example "what is in report_2024.xlsx") searches document names instead and adds "matched_by": "filename" to the same hit shape. mode is ignored on that path. If no visible document has that name, the ordinary search runs. Turn it off with filename_search=False (filenameSearch: false).

An aggregate question (sum, total, average, count, top, per, a "by column" grouping, a numeric comparison and so on) over a scope where every document is a spreadsheet is refused with a next_action pointing at compute. See refusals below. Turn it off with redirect_aggregates_to_compute=False (redirectAggregatesToCompute: false).

get_doc

Needs: Always available

When to use it. Reading one whole document by its system id, taken from discover, list or a search hit. Never pass an invoice number or other content identifier here.

ArgumentTypeNotes
actionstring"get_doc"
document_idstring (uuid)Required.
Python
answer = await engine.search_knowledge_base(
    action="get_doc",
    document_id="3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "get_doc",
  document_id: "3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "document": {
    "id": "...", "name": "...", "text": "...", "acl": [...], "status": "completed",
    ...,
    "failure_reason": null,      // one FAILURE_REASONS code, or null
    "failure_message": null      // the fixed sentence for that code
  }
}

The raw error string of a failed ingest is never returned here; the code and its fixed sentence are. The operator surfaces (engine.get_document, list_documents, DocumentReport.error) keep the raw string.

A document that is absent, not visible to the caller, or outside the host ceiling all answer the same "document not found: <id>".

get_docs

Needs: Always available

When to use it. Several whole documents at once, by id.

ArgumentTypeNotes
actionstring"get_docs"
document_idsstring[]Required.
max_charsintegerTotal document text to return before stopping. Default 200,000, floor 1,000.
Python
answer = await engine.search_knowledge_base(
    action="get_docs",
    document_ids=["3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10", "8d1e7f02-6b3c-4a59-8e2d-1f0c9b7a6e35"],
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "get_docs",
  document_ids: ["3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10", "8d1e7f02-6b3c-4a59-8e2d-1f0c9b7a6e35"],
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "documents": [{...}, {...}],   // same per-document shape as get_doc
  "count": 2,
  "remaining_document_ids": ["..."],                               // only when the budget ran out
  "next_page": {"action": "get_docs", "document_ids": ["..."]}
}

An id that is absent or not visible is simply missing from documents, never an error.

The first document is always returned, even if it alone exceeds max_chars.

get_chunks

Needs: Always available

When to use it. Walking one long document in order, a range at a time, or reading around a search hit by starting at its chunk_idx.

ArgumentTypeNotes
actionstring"get_chunks"
document_idstring (uuid)Required.
startintegerFirst chunk position, 0-based. Default 0.
endintegerLast position, inclusive. At most 25 chunks come back per call whatever is asked.
Python
answer = await engine.search_knowledge_base(
    action="get_chunks",
    document_id="3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10",
    start=4,
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "get_chunks",
  document_id: "3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10",
  start: 4,
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "document_id": "...", "total_chunks": 120, "start": 4, "end": 28,
  "chunks": [{"position": 4, "text": "...", "language": "en"}, ...],
  "has_more": true, "next_start": 29,
  "next_page": {"action": "get_chunks", "document_id": "...", "start": 29}
}

list

Needs: Always available

When to use it. Browsing documents without searching. Cheaper than discover: it does not fetch structure, fields or the census.

ArgumentTypeNotes
actionstring"list"
source_idsstring[]Narrow to these sources.
document_idsstring[]Narrow to these documents.
limitintegerDefault 50, clamped to 1 to 200.
cursorstringThe next_cursor of a previous page.
Python
answer = await engine.search_knowledge_base(
    action="list",
    source_ids=["hr"],
    limit=20,
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "list",
  source_ids: ["hr"],
  limit: 20,
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "documents": [{"id": "...", "name": "...", "source_id": "hr", "kind": "text",
                 "document_type": null, "mode": "hybrid", "failure_reason": null,
                 "status": "completed", "parked": null}],
  "count": 20, "has_more": true,
  "next_cursor": "...", "next_page": {"action": "list", "cursor": "..."}
}

A malformed cursor is refused with "invalid cursor", never read as page one.

A document the host model deferred reads status "waiting_model", not failed, and parked says where it stands: stage, parked_at, retry_at, rounds, request_keys, last_skip_reason, last_skip_at.

query_meta

Needs: An LLM configured

When to use it. The structured fields a question names (dates, amounts, parties) and each document&rsquo;s value for them. It does not filter or compare by value: read the values and compare them yourself.

ArgumentTypeNotes
actionstring"query_meta"
querystringRequired. A question naming a field from fields_by_type.
source_idsstring[]Narrow to these sources.
limitintegerCandidates to return. Default 20, clamped to 1 to 100.
Python
answer = await engine.search_knowledge_base(
    action="query_meta",
    query="invoice amount and invoice date",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "query_meta",
  query: "invoice amount and invoice date",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "question": "...",
  "resolved_fields": [...],
  "documents": [{"id": "...", "source_id": "...", "name": "...",
                 "document_type": "...", "structured_data": {...}, ...}],
  "count": 3
}

Refused for document_ids. The same aggregate refusal as search applies over an all-spreadsheet scope.

compute

Needs: An LLM and enable_code_execution, or a host compute runner

When to use it. Any figure derived from spreadsheets (total, average, count, ranking, margin, comparison) and finding the exact row that matches one id or value. It runs over every row of the sheets in scope, not a sample.

ArgumentTypeNotes
actionstring"compute"
querystringRequired. What to compute, in plain words, with columns named as discover spelled them.
source_idsstring[]Narrow to these sources.
document_idsstring[]Narrow to these documents.
Python
answer = await engine.search_knowledge_base(
    action="compute",
    query="total amount by region in costs.xlsx",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "compute",
  query: "total amount by region in costs.xlsx",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,              // whether the generated code ran
  "result": ...,
  "code": "...", "stdout": "", "error": null,
  "execution_time": 0.41,       // executionTime in TypeScript
  "documents_used": ["..."],    // documentsUsed in TypeScript
  "provider_tokens": {"llm_input": 812, "llm_output": 164},  // providerTokens in TypeScript
  "attempts": 1                 // 1 or 2 code-generation calls
}

Nothing tabular in scope comes back as a result, for example {"success": false, "error": "no tabular (CSV/TSV/XLSX) documents found in scope for compute()"}, not as a transport error.

The whole answer goes through the redaction policy except the top-level list of document ids.

map_reduce

Needs: An LLM configured, or a host map_reduce runner

When to use it. Asking the same question of every document in scope and getting one answer per document, for example "which contracts mention X", where search would return a handful of passages and miss the rest. Not for figures: that is compute.

ArgumentTypeNotes
actionstring"map_reduce"
querystringRequired. The question to ask of each document.
source_idsstring[]Narrow to these sources.
document_idsstring[]Narrow to these documents.
limitintegerDocuments to read, newest first. Default 25, at most 200.
Python
answer = await engine.search_knowledge_base(
    action="map_reduce",
    query="Does this contract have an auto-renewal clause?",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "map_reduce",
  query: "Does this contract have an auto-renewal clause?",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "results": [
    {"document_id": "...", "document_name": "...", "data": {...}},
    {"document_id": "...", "document_name": "...", "error": "..."}
  ],
  "processed": 24, "failed": 1, "considered": 25,
  "hint": {"action": "compute", "reason": "..."}   // only when some targets are spreadsheets
}

Refused for any instruction when every document the call would read is a spreadsheet, because it reads a slice of each document and would merge partial totals. The check runs before any model call and before a host runner.

get_neighbors

Needs: graph.enabled

When to use it. What is one step from a named thing, and which way each link points.

ArgumentTypeNotes
actionstring"get_neighbors"
entitystringRequired. An entity id from an earlier graph answer (preferred), or its name as written in the documents.
source_idsstring[]Narrow to these sources.
limitintegerDefault 20, clamped to 1 to 200.
Python
answer = await engine.search_knowledge_base(
    action="get_neighbors",
    entity="Acme Corp",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "get_neighbors",
  entity: "Acme Corp",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "entity": {"id": "...", "name": "acme corp", "type": "ORG"},
  "neighbors": [{"id": "...", "name": "...", "type": "PERSON", "category": "MEMBERSHIP",
                 "label": "works_for", "evidence": "...", "direction": "incoming"}],
  "count": 7
}

traverse

Needs: graph.enabled

When to use it. Everything within a few steps of a named thing.

ArgumentTypeNotes
actionstring"traverse"
entitystringRequired. An entity id (preferred) or name.
depthintegerHops, 1 to 5. Default 2.
categoryenumOnly follow this kind of connection.
source_idsstring[]Narrow to these sources.
limitintegerPaths. Default 50, clamped to 1 to 200.
Python
answer = await engine.search_knowledge_base(
    action="traverse",
    entity="Acme Corp",
    depth=2,
    category="HIERARCHICAL",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "traverse",
  entity: "Acme Corp",
  depth: 2,
  category: "HIERARCHICAL",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "entity": {"id": "...", "name": "...", "type": "ORG"},
  "paths": [{"target_id": "...", "target": "...", "target_type": "PERSON",
             "category": "HIERARCHICAL", "label": "reports_to", "evidence": "...", "hops": 2}],
  "count": 12
}

Needs: graph.enabled

When to use it. Connections of a kind across the corpus, with no starting entity (an entity is refused; for how two named things connect, traverse from one). Also the way in when every name in an answer is masked: take a source_id or target_id from its result.

ArgumentTypeNotes
actionstring"find_related"
categoryenumThe kind of connection. At least one of category, label or entity_type is required.
labelstringThe exact relationship wording as extracted, for example reports_to.
entity_typeenumKeep only relationships with an entity of this type on one end.
source_idsstring[]Narrow to these sources.
limitintegerRelationships. Default 50, clamped to 1 to 200.
Python
answer = await engine.search_knowledge_base(
    action="find_related",
    category="CREATION",
    entity_type="PRODUCT",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "find_related",
  category: "CREATION",
  entity_type: "PRODUCT",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "relationships": [{"source_id": "...", "source": "...", "source_type": "ORG",
                     "target_id": "...", "target": "...", "target_type": "PRODUCT",
                     "category": "CREATION", "label": "makes", "evidence": "..."}],
  "count": 9
}

community_summary

Needs: graph.enabled and graph.community_summaries

When to use it. The themes each source groups into, ranked against a question.

ArgumentTypeNotes
actionstring"community_summary"
querystringOptional. Without one, the main themes. A question no summary clears the relevance floor for gets the closest themes, marked loose_match.
source_idsstring[]Narrow to these sources.
limitintegerCommunities. Default 3, clamped to 1 to 10.
Python
answer = await engine.search_knowledge_base(
    action="community_summary",
    query="supplier risk",
    principals=caller,
    scope=["hr", "finance"],
)
TypeScript
const answer = await engine.searchKnowledgeBase({
  action: "community_summary",
  query: "supplier risk",
  principals: caller,
  scope: ["hr", "finance"],
});
Returns
{
  "success": true,
  "communities": [{"summary": "...", "entity_count": 14, "relationship_count": 22,
                   "hierarchy_level": 0, "relevance": 0.7121}],
  "count": 3
}

Community summaries are opt-in (graph.community_summaries / graph.communitySummaries, off by default). With them off this action is advertised as off: discover marks it (OFF) and never suggests it, and a call is refused with an error naming the setting.

A community is built and read inside one source. Summaries are withheld (success false, with an error saying why) from a caller who cannot see every chunk in scope, and refused when the host&rsquo;s scope names documents, since a summary covers a whole source.

Graph answers and masking. Entity names, types, labels, evidence and community summaries come from document text, so every string in a graph answer goes through the redaction policy. The entity ids (id, source_id, target_id) are exempt: they carry no text, and they are the only handle a model can pass back as the next entity when the names are masked. An id the caller cannot see answers entity not found, the same as an unknown name. Rebuilding the graph mints new ids.

Refusals and off actions

Which actions can run is decided per call, not when the tool is built, so an engine whose config resolves per request stays correct:

query_meta: an LLM is configured.

compute: an LLM is configured and enable_code_execution (enableCodeExecution) is true, or the host passed a compute runner.

map_reduce: an LLM is configured, or the host passed a map_reduce runner.

traverse, find_related, get_neighbors, community_summary: graph.enabled.

Everything else is always available.

An off action is advertised as off by discover and, if called anyway, returns a fixed sentence. The other common refusals:

Result shapes
// Aggregate question over a scope that is all spreadsheets (search, query_meta)
{"success": false,
 "error": "these documents are spreadsheets; searching their rows would answer from a sample. Use compute over the whole sheet.",
 "next_action": {"action": "compute", "query": "total spend by region",
                 "suggestion": {"frame_key": "...", "measure": "amount", "group_by": "region"}}}

// An action this deployment cannot run
{"success": false, "error": "compute is off: code execution is not enabled on this deployment. Quote the rows you found instead."}

// document_ids sent to an action that does not honour it
{"success": false, "error": "document_ids is not honored by get_doc ..."}

In the aggregate refusal, next_action.query is the question handed back so the redirect is a call the model can make. suggestion is vocabulary, not code: frame_key is the name compute will give the newest in-scope sheet, measure its first numeric column, group_by its first text column. Any of them can be null. A scope that mixes prose and spreadsheets is still searched and only gets an additive hint.

Answers may also carry next_page or next_action: the next call, already filled in. The tool description tells the model to follow it rather than guess.

Scope and ceilings

scope is the ceiling the host sets: the documents this tool may ever reach. It uses only nouns the engine already knows, source ids and optionally document ids within them. Map your own idea onto it: a project to its source ids, a matter to its source ids, an agent to the documents it was handed.

What a host may pass

A list of source ids; a Scope with source_ids and document_ids (a sourceIds / documentIds object in TypeScript); or UNSCOPED for the whole corpus on purpose. None / null raises: that is what a host that forgot looks like.

Narrow, never widen

The model's source_ids and document_ids are intersected with the ceiling. An id outside it is dropped silently, because an error would reveal which ids are real. A request left with nothing in scope returns nothing.

Python
from context_engine import Scope, UNSCOPED

scope = ["hr", "finance"]                         # a list: these source ids
scope = Scope(source_ids=("hr",),                 # sources, and specific documents in them
              document_ids=("3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10",))
scope = UNSCOPED                                  # the whole corpus, on purpose
scope = None                                      # raises ValueError: scope is required

from context_engine.knowledge_tool import narrow_to_ceiling, resolve_scope
narrow_to_ceiling(["hr", "legal"], ("hr", "finance"))   # -> ["hr"]   ("legal" dropped)
narrow_to_ceiling(["legal"], ("hr", "finance"))         # -> []       (nothing in scope)
narrow_to_ceiling(None, ("hr", "finance"))              # -> ["hr", "finance"]
narrow_to_ceiling(None, None)                           # -> None     (no filter at all)
TypeScript
import { narrowToCeiling, resolveScope, UNSCOPED, type ScopeInput } from "@promptev/context-engine";

let scope: ScopeInput = ["hr", "finance"];        // an array: these source ids
scope = { sourceIds: ["hr"], documentIds: ["3f2b8c1e-5a4d-4e8f-9b61-2c7d0a9e4f10"] };
scope = UNSCOPED;                                 // the whole corpus, on purpose
resolveScope(null);                               // throws: scope is required

narrowToCeiling(["hr", "legal"], ["hr", "finance"]); // -> ["hr"]
narrowToCeiling(["legal"], ["hr", "finance"]);       // -> []
narrowToCeiling(null, ["hr", "finance"]);            // -> ["hr", "finance"]
narrowToCeiling(null, null);                         // -> null (no filter at all)

Who may name documents

A document-id ceiling is a whitelist. get_doc and get_chunks check it and the source ceiling too, and a document outside either reads exactly like one that does not exist. discover and list drop listed documents outside a document-id ceiling. Within the ceiling, the caller's principals still apply: the ACL is enforced inside the SQL of every read.

Only six actions honour document_ids (search, compute, get_docs, map_reduce, discover, list). The others refuse it rather than answer over the whole scope as if it had been honoured.

resolve_scope and narrow_to_ceiling live in context_engine.knowledge_tool; only Scope is exported at the Python package root. TypeScript exports resolveScope and narrowToCeiling at the root.

Redaction and host runners

redaction overrides the deployment's configured policy for one call, with the same contract engine.search has. Without it, a host with a per-project or per-customer policy would get masked text from every other read path and unmasked text from this one. secret_key is the other half: the key hash rules use for this call. A per-tenant policy hashed with the engine-wide key gives every tenant the same token for the same value, so tokens would join across tenants. Omitted, the engine's own policy and key apply.

Every action applies them: document text, chunks, hits, computed results and graph answers.

Replacing compute and map_reduce

A host can pass its own compute and map_reduce callables. They replace the built-in ones, which is where a host applies its own rules for permission, billing and approval, and they make the action available whatever the deployment settings say. A runner is never handed the redaction policy or the hash key. What it returns is masked on the engine side, the whole value including any documents_used. What the runner's own model saw is the host's responsibility.

Python
async def my_compute(instruction, *, source_ids, document_ids, principals) -> dict:
    ...   # your sandbox, your billing, your approval

async def my_map_reduce(instruction, *, source_ids, document_ids, principals, limit) -> dict:
    ...

app.mount("/mcp", create_mcp_app(engine, principals=principals, scope=scope,
                                 compute=my_compute, map_reduce=my_map_reduce))

# or per call, without MCP:
await engine.search_knowledge_base(action="compute", query="...", principals=caller,
                                   scope=["finance"], compute=my_compute)
TypeScript
const myCompute = async (
  instruction: string,
  opts: { sourceIds: string[] | null; documentIds: string[] | null; principals: unknown },
) => ({ /* your sandbox, your billing, your approval */ });

const myMapReduce = async (instruction: string, opts: Record<string, unknown>) =>
  ({ /* opts: sourceIds, documentIds, principals, limit */ });

const handler = await createMcpApp(engine, {
  principals, scope, compute: myCompute, mapReduce: myMapReduce,
});

The spreadsheet refusal for map_reduce runs before a host runner is called. When the engine resolves the documents a run would read, the runner receives those ids; if that set is empty, it receives the caller's original document_ids instead of an empty list.

Serving over MCP

create_mcp_app (createMcpApp) serves the knowledge tool over MCP streamable HTTP, with no stdio. It needs the [mcp] extra in Python and the @modelcontextprotocol/server and @modelcontextprotocol/node peers in TypeScript. Beside search_knowledge_base it always registers the connector-tool pair search_tools and execute_tool. The served description adds one sentence saying results are scoped to the caller's permissions.

Python
from fastapi import FastAPI
from context_engine import create_mcp_app

app = FastAPI()

def principals():            # zero-argument, sync or async, called on EVERY tool call
    return current_user_groups()           # [] for anonymous; never None

def scope():                 # same contract; or pass a plain value
    return sources_for_current_project()   # e.g. ["hr", "finance"]

app.mount("/mcp", create_mcp_app(
    engine,
    principals=principals,   # required, keyword-only
    scope=scope,             # required in practice: None fails every call
    redaction=policy_for_current_tenant,   # optional: a policy, or a callable returning one
    secret_key=key_for_current_tenant,     # optional: hash key for this call
))
TypeScript
import { createServer } from "node:http";
import { createMcpApp } from "@promptev/context-engine";

// createMcpApp is async: await it. It returns a (req, res) handler.
const handler = await createMcpApp(engine, {
  principals: () => currentUserGroups(),       // required; [] is anonymous
  scope: () => sourcesForCurrentProject(),     // required; throws TypeError if missing
  redaction: () => policyForCurrentTenant(),   // optional, per call
  secretKey: () => keyForCurrentTenant(),      // optional, per call
});

createServer((req, res) => {
  void handler(req, res);
}).listen(8080);
OptionRequiredWhat it is
principalsYesZero-argument callable, sync or async, called fresh on every tool call. lambda: [] serves everyone as anonymous; it has no default because caller identity has no safe one.
scopeYesThe ceiling, as a value or a zero-argument callable resolved per call. TypeScript throws TypeError when the app is built without it. Python's parameter defaults to None, which makes every call fail with "scope is required".
redactionNoA RedactionPolicy, or a callable returning one, per call.
secret_key / secretKeyNoA string, or a callable returning one, per call.
compute, map_reduce / mapReduceNoHost runners, as above.
approval_scope / approvalScopeNoZero-argument callable returning the approval scope (a run id is the expected shape) for execute_tool.

Python returns

A Starlette ASGI app to mount wherever you like. The underlying FastMCP instance is on app.state.mcp.

TypeScript returns

A promise of a Node (req, res) handler, with the McpServer on handler.mcp. It is exported at the root and at @promptev/context-engine/mcp.

The context-engine mcp command serves one database this way with anonymous principals and UNSCOPED. See the CLI reference. For HTTP routes see the HTTP API reference.

Python and TypeScript parity

Both packages run the same engine over the same Postgres schema; a corpus ingested by one is searchable by the other. Names follow each language: snake_case in Python, camelCase in TypeScript. Wire payloads the model or an MCP host reads (knowledge tool results, failure_reason) are snake_case in both. The tables list what each package exports from its root; a name marked for one side often exists on the other in a submodule, which the table says.

Engine methods

PythonTypeScriptNotes
ContextEngine(config)new ContextEngine(config)
aclose()aclose()
with_redaction(...)withRedaction(...)
embedder (property)embedder (getter)
ingestingest
resume_embeddingsresumeEmbeddings
searchsearchResult usage keys follow each port: mode_used / modeUsed.
get_documentgetDocumentPython raises KeyError when absent or not visible; TypeScript has DocumentNotFoundError.
get_documentsgetDocuments
get_document_textgetDocumentText
get_chunksgetChunks
update_documentupdateDocument
delete_documentdeleteDocument
statsstats
list_documentslistDocuments
documents_by_namedocumentsByName
document_structuredocumentStructure
document_typesdocumentTypes
field_summaryfieldSummary
spreadsheet_schemaspreadsheetSchema
tabular_scopetabularScope
query_structuredqueryStructured
computecomputeResult keys: execution_time, documents_used, provider_tokens vs executionTime, documentsUsed, providerTokens.
map_reducemapReduce
map_reduce_targetsmapReduceTargets
search_knowledge_base(**kwargs)searchKnowledgeBase({...})The TypeScript argument object uses the tool's snake_case names.
register_toolregisterTool
register_function_toolregisterFunctionTool
update_toolupdateTool
delete_tooldeleteTool
list_toolslistTools
search_toolssearchTools
test_tooltestTool
execute_toolexecuteTool

Exported by both roots

PythonTypeScriptNotes
ContextEngineConfigContextEngineConfigPython reads CE_ env vars on construction (pydantic-settings); TypeScript has ContextEngineConfig.fromEnv().
EmbeddingConfig, LLMConfig, GraphConfig, SearchConfig, FusionConfig, RerankerConfig, ExtractionConfig, StorageConfigsame names (types; plus ContextEngineConfigInit, SearchConfigInit)Python classes, TypeScript types.
TRUSTED, UNSET, UNSCOPEDTRUSTED, UNSET, UNSCOPED
Scope (frozen dataclass)Scope, ScopeInput (types)TypeScript passes a plain object: { sourceIds, documentIds }.
call_knowledge_toolcallKnowledgeTool
knowledge_tool_definitionknowledgeToolDefinition
create_mcp_appcreateMcpAppTypeScript is async (await it) and is also at the ./mcp subpath.
RedactionPolicy, RedactionRule, apply_redactionRedactionPolicy, RedactionRule, applyRedaction
ToolConfig, ToolKind, config_schemaToolConfig, ToolKind, configSchema
resolve_approval, ApprovalRecord, ApprovalExpired, ApprovalNotPendingresolveApproval, ApprovalRecord, ApprovalExpired, ApprovalNotPending
format_result, rows_to_tsvformatResult, rowsToTsv
Hooks, emit_usage, emit_progress, emit_errorHooks, emitUsage, emitProgress, emitError
UsageEvent, DocumentReport, IngestReportsame (types)
FAILURE_REASONS, failure_message, classify_failureFAILURE_REASONS, failureMessage, classifyFailure
graph_units, units_for_filegraphUnits, unitsForFile
Hit, SearchResult, run_search, GRAPH_REASONSHit, SearchResult, runSearch, GRAPH_REASONS
query_type, QueryType, is_spaceless, trigram_limit, vector_floorqueryType, QueryType, isSpaceless, trigramLimit, vectorFloor
rrf_fuserrfFuse
compute, compute_over_frames, list_documents, get_document_text, query_structuredcompute, computeOverFrames, listDocuments, getDocumentText, queryStructuredModule-level functions beside the engine methods.
extract, Extractedextract, Extracted
extract_structured_data, resolve_fields, upsert_registry, ExtractionResultextractStructuredData, resolveFields, upsertRegistry, ExtractionResult
Embedder, LLMClient, build_embedder, build_llm_client, call_llm, host_rerankEmbedder, LLMClient, buildEmbedder, buildLlmClient, callLlm, hostRerankThe built-in OpenAI-compatible client and the host-reranker helper.
PostgresBackend, StorageBackend, ChunkRow, SearchScopePostgresBackend, StorageBackend, ChunkRow, SearchScopeStorageBackend, ChunkRow and SearchScope are types in TypeScript.
CodeRunnerCodeRunner (type)
HttpClientFactory, HttpClientPurposeFetchFactory, FetchPurpose (types)Python hands back an httpx client per purpose; TypeScript a fetch.
encrypt_dict, decrypt_dict, get_secret_keyencryptDict, decryptDict, getSecretKey
EngineActionError, AmbiguousToolError, IngestTooLarge, ExtractionFailed, GraphLegUnavailable, InputTooLarge, EmbedBatchFailed, EmbeddingBatchRejectedsame names
__version____version__

Only at the TypeScript root

Python equivalentTypeScriptNotes
context_engine.cli.run_migraterunMigratePython has it in the cli module, not at the root.
context_engine.runner (CeleryRunner, InProcessRunner, TaskRunner)CeleryRunner, InProcessRunner, TaskRunner, TaskStatus
KeyErrorDocumentNotFoundError
ImportError naming the extraExtraMissingError
context_engine.sandboxCodeExecutionError, CodeExecutionTimeout
context_engine.llm_schemasloadSchema, validate, strictPortable, SchemaError, SchemaValidationError
context_engine.providers.structuredstructuredCall, StructuredCallError, StructuredRefusal, structuredRepairCount, resetStructuredRepairCount
context_engine.providers.llm (LLM_MAX_RETRIES)LLM_MAX_RETRIES, resolveMaxRetries, resetNativeSchemaSupport
context_engine.knowledge_toolKNOWLEDGE_ACTIONS, KNOWLEDGE_TOOL_DESCRIPTION, resolveScope, narrowToCeiling, KnowledgeAction, KnowledgeComputeFn
context_engine.actionsdocumentStructure, documentTypes, fieldSummary, spreadsheetSchema, spreadsheetSchemaFromTextThe engine methods exist in both.
context_engine.config (redaction_key, check_redaction_key)redactionKey, checkRedactionKey, graphNeo4jConfiguredPython: GraphConfig.neo4j_configured is a property.
context_engine.intakecheckIntakeSize, checkRowLimit, checkTabularRows, countCsvRows, countTabularRows, countXlsxRows, DEFAULT_MAX_FILE_BYTES
context_engine.routing_corecheckContentLength, uploadLimitBytes
context_engine.search.redact_hitsredactHits
context_engine.hooks.emit_tool_callemitToolCall
context_engine.tools.approval.should_require_approvalshouldRequireApproval
context_engine.tools.executors.function.function_toolfunctionTool
context_engine.fusion.DEFAULT_LEG_WEIGHTDEFAULT_LEG_WEIGHT
noneresolvePrincipalsPython resolves internally.
nonetypes: Principals, Trusted, Unset, ProgressEvent, GraphReason, GraphReport, Mode, StructuredReply, StructuredCallOpts, StructuredResult, EmbedKind, JsonSchema, SchemaName, HostLookup, ResponseMode, RedactionRuleInit, ComputeDocument, ComputeFrames

Only in Python, or different by design

PythonTypeScriptNotes
create_router (FastAPI)createHonoRouter / createHonoApp (./hono), createExpressRouter (./express), createFastifyPlugin (./fastify)Different frameworks per language; TypeScript routers are subpath imports.
create_tools_router (FastAPI)noneIn TypeScript the tool admin and execute routes are part of each router.
create_flask_blueprint, create_django_urlpatternsnoneNo tool routes on Flask or Django.
presidio_detector, presidio_detectors (root, lazy)presidioDetector, presidioDetectors (./presidio subpath)Python runs Presidio in process ([presidio] extra, analyzer= injectable). TypeScript calls a Presidio Analyzer over HTTP at CE_PRESIDIO_URL (or PRESIDIO_URL) and fails open.
StorageConfig: pool_size 5, max_overflow 10, pool_timeout 30.0, pool_recycle 1800, pool_pre_ping TrueStorageConfig: poolMax 10, poolIdleTimeoutMs 10000, poolConnectionTimeoutMs 30000SQLAlchemy pool vs pg.Pool. pg has no pre-ping.

Configuration fields are covered field by field in the configuration reference; full signatures are in the Python API and TypeScript API pages.

Further reading

Cookbooks

Each package ships a runnable multi-tenant RAG recipe: cookbook/multi_tenant_rag.py and js/cookbook/multi_tenant_rag.ts. Three documents share one corpus: an HR policy with acl=["group:hr"], an engineering runbook with acl=["group:eng"], and a holidays page with no ACL. One question is asked four ways, and the recipe asserts the answers: the HR employee sees holidays and the HR policy, the engineer sees holidays and the runbook, an anonymous visitor sees only holidays, and a trusted job sees all three. Each recipe reads DATABASE_URL and embedding settings from the environment, runs migrations and writes documents, so point it at a scratch database. Get the packages from PyPI and npm.

The access-control model

principals has three meanings. TRUSTED disables ACL filtering; [] is anonymous and sees only documents with no ACL; a list sees unrestricted documents plus any whose ACL overlaps it. None is deprecated and means trusted, so never use it for a request with no user: unauthenticated is []. A value that is not a list raises.

Enforcement is in the SQL of every retrieval leg (full-text, trigram, vector and the graph entity filter), before ranking. A document the caller may not see is never fetched, scored or sent to a model. acl IS NULL means unrestricted, and matching is overlap, not containment.

Remote surfaces take principals by injection only, never from a body, query string or header. auth and principals are required on every router factory and on create_mcp_app. A router's principals returning None is rejected; a genuinely trusted mount writes lambda: TRUSTED.

Writes are authorised separately. Filing or patching a document under an ACL the caller does not hold returns 403. A tool with no ACL is refused unless the caller is trusted.

Updating with acl=None unrestricts the document. Leave the argument out to keep the ACL.

The vector index. An ACL on an approximate index is a post-filter that can silently return short. The engine counts eligible rows, scans exactly below StorageConfig.ann_exact_threshold (default 50,000), and above it sizes pgvector's iterative scan to the table. pgvector 0.8 or later is required. context-engine check-acl-exposure measures this on your own data, read-only.

Mistakes worth naming: passing TRUSTED and filtering afterwards (the ranking already reflects documents the caller cannot see), caching a search result across users, and reading principals from the request.

Thresholds follow the query's shape

The trigram leg's similarity threshold moves with the query, from 0.45 for a very short query down to 0.15 for a long one, measured in characters for spaceless scripts (Chinese, Japanese, Thai, Khmer, Burmese, Lao, Tibetan) and in words otherwise. A matching minimum cosine similarity for the vector leg exists but is off by default: search.vector_floor="adaptive" turns it on (0.45 for one or two words, 0.40 for a code, an acronym or an exact-match query, down to 0.25 for a sentence). With it on, the vector leg can return fewer candidates and nothing backfills them. Measure on your own embedding model before enabling it. search.trgm_limit pins the trigram threshold to one value.

Keyword search settings

The full-text leg uses Postgres text search over the simple configuration, which keeps every word. By default (search.lexical_match="all") a chunk must contain every word of the query, stopwords included: precise for keyword queries, but a question such as “what is the notice period for contractors” usually matches nothing. "any" matches a chunk containing any of the query’s words after dropping stopwords (search.lexical_stopwords, the built-in English list by default) and leaves the order to the rank function. Quoted phrases and -term keep their meaning in both modes. For "notice period" contractor -draft, "all" matches chunks with the exact phrase AND the word contractor and without the word draft; "any" matches chunks with the phrase OR the word contractor, still without draft. The simple configuration does not stem, so -draft does not exclude “drafts”. The query is split into words, phrases and terms on ASCII whitespace. search.lexical_rank picks ts_rank_cd (words close together, the default), ts_rank (frequency anywhere) or the experimental, opt-in bm25 (Okapi BM25 in SQL, turned on per database with context-engine bm25-enable, not a default; with search.bm25.k1 1.2, search.bm25.b 0.75 and search.bm25.idf_scope "source", "global" or "document"; its statistics include documents the caller may not see, though no content or id is exposed), and search.lexical_normalization passes Postgres’s 0 to 63 length-normalization bitmask through. Query text is always bound as data, and every variant applies the same scope predicate as the other legs. A leg weighted 0 in fusion.weights is switched off and its query is not sent: the vector leg also skips the query embedding unless compression needs it, and the graph leg is never started, billed as hybrid, and reported as graph_weight_zero. The default weights are 1.0 for the full-text, vector and graph legs and 0.4 for the trigram leg, which the retrieval benchmark measured ahead of 0.8 on four of five public datasets. search.trgm_max_query_words skips the trigram leg for queries longer than that many words, and search.lexical_candidates sets how many candidates the full-text and trigram legs hand to fusion; both are off by default, and measured on five public datasets neither improved ranking everywhere. Fusion uses reciprocal rank fusion with fusion.k 20. The TypeScript names are lexicalMatch, lexicalRank, bm25.idfScope, lexicalNormalization, lexicalStopwords, trgmMaxQueryWords and lexicalCandidates.

Why a document failed

A failed document carries error, the raw exception text for operators, and failure_reason, one fixed code, with failure_message as its fixed sentence. Model-facing doors (this tool's get_doc, get_docs and list, and the routers' document read for non-trusted callers) never return error. A later success clears both.

failure_reasonfailure_message
input_too_largethe file is too large to process in one go. Split it into smaller files and upload those
provider_rejectedthe embedding provider refused the request. Check the model name and key
provider_unavailablethe embedding provider was unavailable. Try again later
extraction_failedthe file could not be read
chunk_set_changedthe document was re-ingested while it was being processed. Re-ingest or resume
model_deferred_expiredthe host model did not answer within ingest.max_park_seconds. Re-ingest the document
embedding_dimension_mismatchthe embedder returned vectors of a different size than this database stores. Check the embedding model and its dimension
unknownprocessing failed

An unrecognised code reads as unknown rather than as no failure.

The ingest usage event

hooks.on_usage is the metering seam; the package never handles billing itself. Every UsageEvent(kind="ingest") carries mode (the mode the document was actually processed at) and mode_reason (why that differs from the request, currently only "tabular documents stay hybrid", else None) in its detail, on every path that emits one. A DOCX is counted by the page count the file reports, with pages_source saying where it came from: app_xml, rendered_breaks, page_breaks or estimate. A declared count is only believed as far as the body could hold it.

2,500 free credits · No card required · No subscription