How configuration works
ContextEngineConfig is the one settings object every part of the engine reads. You build it once and pass it to ContextEngine. Both packages have the same sections and the same defaults, with snake_case names in Python and camelCase names in TypeScript.
Two required fields
database_url and embedding (a provider and a model). Everything else has a default, and every optional subsystem is off until you turn it on.
Checked at construction
A subsystem turned on without what it needs raises when the config is built, not on the first call. See Validation. Out-of-range values are refused the same way in both packages when the configuration is built, not at first use: the ranges are in the Type column of each table.
Where the names come from: Python is a pydantic-settings model; TypeScript is a class over zod 4 schemas. In Python, IngestConfig is imported from context_engine.config; the other section classes are exported from context_engine. In TypeScript each section is a plain object literal.
Loading from the environment
Every field can come from an environment variable: prefix CE_, the field name in upper case, and __ between nesting levels. embedding.api_key is CE_EMBEDDING__API_KEY, and graph.extraction_llm.model is CE_GRAPH__EXTRACTION_LLM__MODEL. The same variable names work in both packages; TypeScript converts each segment to camelCase.
from context_engine import ContextEngine, ContextEngineConfig
# No arguments: every field comes from CE_* environment variables.
config = ContextEngineConfig()
# Keyword arguments win over the environment, field by field.
config = ContextEngineConfig(database_url="postgresql://user:pass@db:5432/app")
engine = ContextEngine(config)
import { ContextEngine, ContextEngineConfig } from "@promptev/context-engine";
// Reads CE_* from process.env. Throws unless CE_DATABASE_URL and
// CE_EMBEDDING__PROVIDER / CE_EMBEDDING__MODEL (or overrides) are present.
const config = ContextEngineConfig.fromEnv();
// Overrides are deep-merged over the environment.
const withOverride = ContextEngineConfig.fromEnv({
databaseUrl: "postgresql://user:pass@db:5432/app",
});
const engine = new ContextEngine(config);CE_DATABASE_URL=postgresql://user:pass@localhost:5432/mydb
CE_EMBEDDING__PROVIDER=openai
CE_EMBEDDING__MODEL=text-embedding-3-small
CE_EMBEDDING__API_KEY=sk-...
CE_LLM__PROVIDER=openai
CE_LLM__MODEL=gpt-4o-mini
CE_LLM__API_KEY=sk-...
CE_GRAPH__ENABLED=true
CE_GRAPH__EXTRACTION_LLM__PROVIDER=openai
CE_GRAPH__EXTRACTION_LLM__MODEL=gpt-4o-mini
CE_GRAPH__EXTRACTION_LLM__API_KEY=sk-...
CE_INGEST__MAX_ROWS=500000
CE_SECRET_KEY=<base64url 32-byte key>
Python reads the environment whenever ContextEngineConfig is constructed. Values passed as keyword arguments take priority. Dict and list fields such as fusion.weights are given as JSON.
TypeScript reads the environment only through ContextEngineConfig.fromEnv(overrides); new ContextEngineConfig(init) never looks at it. The loader turns true / false into booleans and integer or decimal strings into numbers, and leaves everything else a string. It does not parse JSON. A dictionary setting is set one key at a time, and its keys keep their spelling, as in Python: CE_GRAPH__RERANK_WEIGHTS__VECTOR_SCORE sets vector_score, and CE_FUSION__WEIGHTS__MY_LEG sets the leg my_leg. Field names are still camel-cased. Set the redaction policy in code.
Neither package reads a .env file. Export the variables, or load the file with your own tooling first.
A variable only reaches the port that has the field. CE_STORAGE__POOL_SIZE is Python only and CE_STORAGE__POOL_MAX is TypeScript only; each port ignores the other’s.
Full example
Most sections at once. Anything you leave out keeps the default in the tables below.
from context_engine import (
ContextEngine, ContextEngineConfig, EmbeddingConfig, LLMConfig,
GraphConfig, RerankerConfig, SearchConfig, StorageConfig,
)
from context_engine.config import IngestConfig
config = ContextEngineConfig(
database_url="postgresql://user:pass@localhost:5432/mydb",
embedding=EmbeddingConfig(provider="openai", model="text-embedding-3-small",
api_key="sk-...", max_retries=1),
llm=LLMConfig(provider="openai", model="gpt-4o-mini", api_key="sk-..."),
vision_llm=LLMConfig(provider="openai", model="gpt-4o-mini", api_key="sk-..."),
graph=GraphConfig(
enabled=True, # Postgres-only graph mode: no neo4j_uri
extraction_llm=LLMConfig(provider="openai", model="gpt-4o-mini", api_key="sk-..."),
rerank_weights={"entity_match": 0.4}, # the other four keep their defaults
),
reranker=RerankerConfig(candidates=50), # the window a host reranker= sees
search=SearchConfig(vector_floor="adaptive"),
ingest=IngestConfig(max_rows=500_000),
storage=StorageConfig(pool_size=10, max_overflow=20, pool_recycle=300),
secret_key="...", # base64url, 32 bytes
)
engine = ContextEngine(config)import { ContextEngine, ContextEngineConfig } from "@promptev/context-engine";
const config = new ContextEngineConfig({
databaseUrl: "postgresql://user:pass@localhost:5432/mydb",
embedding: { provider: "openai", model: "text-embedding-3-small",
apiKey: "sk-...", maxRetries: 1 },
llm: { provider: "openai", model: "gpt-4o-mini", apiKey: "sk-..." },
visionLlm: { provider: "openai", model: "gpt-4o-mini", apiKey: "sk-..." },
graph: {
enabled: true, // Postgres-only graph mode: no neo4jUri
extractionLlm: { provider: "openai", model: "gpt-4o-mini", apiKey: "sk-..." },
rerankWeights: { entity_match: 0.4 }, // the other four keep their defaults
},
reranker: { candidates: 50 }, // the window a host reranker sees
search: { vectorFloor: "adaptive" },
ingest: { maxRows: 500_000 },
storage: { poolMax: 20, poolIdleTimeoutMs: 30_000 },
secretKey: "...", // base64url, 32 bytes
});
const engine = new ContextEngine(config);Top-level fields
Fields that sit directly on ContextEngineConfig.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| database_url | databaseUrl | SecretStr / Secret | required | CE_DATABASE_URL | Postgres connection URL. Python hands it to SQLAlchemy (postgresql:// uses psycopg2 from the [postgres] extra); TypeScript hands it to pg.Pool as connectionString. It carries the password, so it is a secret field. |
| embedding | embedding | EmbeddingConfig | required | CE_EMBEDDING__* | The embedding provider. See Embedding. |
| default_mode | defaultMode | "hybrid" | "graph" | "hybrid" | CE_DEFAULT_MODE | The mode ingest() uses when a call names none. "graph" also runs entity and community extraction after the chunks are committed, and requires graph.enabled. |
| llm | llm | LLMConfig | None | None | CE_LLM__* | The model for structured extraction, compute, map_reduce and query_meta. Without it those features are off. See LLM. |
| vision_llm | visionLlm | LLMConfig | None | None | CE_VISION_LLM__* | The model that transcribes scanned PDF pages and images. Without it, scanned PDF pages fall back to Tesseract OCR when the OCR dependency is installed, and a standalone image is treated as media with no text. |
| storage | storage | StorageConfig | defaults | CE_STORAGE__* | Backend, exact-scan threshold and connection pool. See Storage and pools. |
| graph | graph | GraphConfig | disabled | CE_GRAPH__* | Entity graph. See Graph. |
| reranker | reranker | RerankerConfig | candidates 50 | CE_RERANKER__* | The widened candidate window a host reranker sees. There is no built-in reranker. See Reranker. |
| fusion | fusion | FusionConfig | RRF, k=20 | CE_FUSION__* | How the retrieval legs are merged. See Fusion. |
| search | search | SearchConfig | both off | CE_SEARCH__* | Query-shape knobs. See Search. |
| extraction | extraction | ExtractionConfig | defaults | CE_EXTRACTION__* | Rendering, vision and OCR limits. See Extraction. |
| ingest | ingest | IngestConfig | defaults | CE_INGEST__* | Intake caps, progress events, the resume fence and the park limit. See Ingest. |
| enable_code_execution | enableCodeExecution | bool | False | CE_ENABLE_CODE_EXECUTION | Allows compute() to run LLM-generated code. Off, compute() raises EngineActionError before any LLM call or sandbox run. The in-process sandbox is not a security boundary on its own: turn this on only inside OS-level isolation (a container, gVisor or similar). |
| redirect_aggregates_to_compute | redirectAggregatesToCompute | bool | True | CE_REDIRECT_AGGREGATES_TO_COMPUTE | The knowledge tool refuses an aggregate question (a total, an average, a ranking) over spreadsheets only, and answers with the compute call to make instead. A mixed scope still searches and carries a hint. In TypeScript an explicit null counts as off. |
| filename_search | filenameSearch | bool | True | CE_FILENAME_SEARCH | The knowledge tool answers a question that names a file by searching document names first, under the same scope. When no name matches, the ordinary search runs. In TypeScript an explicit null counts as off. |
| secret_key | secretKey | SecretStr / Secret | None | None | CE_SECRET_KEY | Base64url-encoded 32-byte AES-256-GCM key for tool configs that hold secrets. See Secret keys. |
| redaction_secret_key | redactionSecretKey | SecretStr / Secret | None | None | CE_REDACTION_SECRET_KEY | HMAC key for hash redaction rules. Falls back to secret_key when unset. |
| allow_private_egress | allowPrivateEgress | bool | False | CE_ALLOW_PRIVATE_EGRESS | Lets http and mcp tools and their probes reach loopback, private, link-local and metadata addresses. Off by default because a private destination registered by a caller is server-side request forgery. |
| tools.mcp.protocol | tools.mcp.protocol | "auto" | "modern" | "legacy" | "auto" | CE_TOOLS__MCP__PROTOCOL | How a connection to a third-party MCP server picks its protocol version. auto: the MCP SDK asks the server and falls back to the initialize handshake when the server does not know the question. modern: the 2026-07-28 revision only. legacy: the handshake only. A tool’s own config may set protocol for its connection, and that value wins. See Connecting MCP servers. |
| tools.mcp.redirects | tools.mcp.redirects | "refuse" | "same_origin" | "any" | "refuse" | CE_TOOLS__MCP__REDIRECTS | What to do when an MCP server answers with a redirect. refuse fails the request. same_origin follows one that keeps the scheme, host and port. any also follows one to another origin, without credentials, the session id or the tool’s headers. A followed redirect is egress-checked at every hop, capped at 3 hops and never taken from https to http. Per-connection key: redirects. |
| tools.mcp.elicitation_url | tools.mcp.elicitationUrl | bool | False | CE_TOOLS__MCP__ELICITATION_URL | Pass a third-party MCP server’s URL-mode question (open a sign-in or payment page) to on_elicitation. Off, such a question is declined without asking the hook. On, the URL goes to the hook only; the engine never opens or fetches it. Needs the hook. Per-connection key: elicitation_url or elicitationUrl. |
| tools.mcp.elicitation_timeout_s | tools.mcp.elicitationTimeoutS | float > 0 | 300 | CE_TOOLS__MCP__ELICITATION_TIMEOUT_S | Seconds on_elicitation may take per question. On expiry the server gets cancel and on_error fires. Not counted against the tool call’s own timeout. Per-connection key: elicitation_timeout_s or elicitationTimeoutS. |
| tools.mcp.redact_logs | tools.mcp.redactLogs | bool | True | CE_TOOLS__MCP__REDACT_LOGS | Keeps secrets out of the MCP client’s log lines: no URL query string, userinfo, session id or header value, and a URL is logged as scheme, host, port and path. False lets those through for debugging. What a user answered to a server’s question is never logged with either value. Global only: a tool config cannot change it. |
| tools.mcp.max_response_bytes | tools.mcp.maxResponseBytes | int > 0 | None | 10000000 | CE_TOOLS__MCP__MAX_RESPONSE_BYTES | The most decoded bytes one MCP response may carry, counted as it is read, so a compressed body that unpacks to more is stopped too. Past it the request fails with an error naming this setting. Nonenull lifts it. Global only. Connecting to a server and listing its tools before a call is bounded at 60 seconds as a whole. |
| tools.audit_on_failure | tools.auditOnFailure | "fail" | "warn" | "fail" | CE_TOOLS__AUDIT_ON_FAILURE | What a tool call does when its audit row cannot be written at all, neither the full row nor the last-resort row (the database is down). The tool has already run by then, and on_error is called with stage tool_audit either way. "fail" raises ToolAuditFailed, which says the tool ran and carries its result, so nothing runs unrecorded and the caller can still avoid running it twice; the routers answer 502 with tool_ran: true, and the call is metered, because the tool ran. "warn" returns the tool’s result as usual, for a deployment that would rather keep serving than keep a complete audit trail. It is global: one value for the engine, not per tool or per call. In TypeScript an unset value means "fail". See the audit guarantee. |
| tools.audit_cancel_wait_s | none | float ≥ 0 | 10 | CE_TOOLS__AUDIT_CANCEL_WAIT_S | Python only. A call cancelled after its tool started (a timeout around execute_tool, an MCP client that gave up) writes its audit row before the cancellation is raised: the full row and, if that fails, the fallback rows. This is how many seconds that wait may last. When the database does not answer in that time, on_error is called with stage tool_audit and a TimeoutError, the cancellation is raised, and the write is left to finish in the background, so the row may still arrive. 0 does not wait. A call that is not cancelled is not bounded by this setting. JavaScript cannot cancel a promise, so the TypeScript package has no counterpart. |
| tools.db.query_timeout_s | tools.db.queryTimeoutS | float > 0 | None | 30 | CE_TOOLS__DB__QUERY_TIMEOUT_S | Seconds a db tool query may run on the database server before it is cancelled there (Postgres statement_timeout, MySQL max_execution_time, MariaDB max_statement_time, MSSQL the connection timeout, ClickHouse max_execution_time, Oracle call_timeout; SQLite has none). A timed-out query is a failed call. A tool’s own query_timeout_s overrides it; Nonenull lifts it. |
| redaction | redaction | RedactionPolicy | empty (no-op) | code only | Span-level masking at ingest and/or output. Set in code. See Redaction. |
Embedding
EmbeddingConfig, required. Environment prefix CE_EMBEDDING__.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| provider | provider | "openai" | "azure_openai" | "custom" | required | CE_EMBEDDING__PROVIDER | Which OpenAI-compatible endpoint the built-in client calls. See Providers. With a host embedder it is the identity the database records and nothing is called. |
| model | model | str | required | CE_EMBEDDING__MODEL | The provider’s model name. |
| dim | dim | int ≥ 1 | None | None | CE_EMBEDDING__DIM | Vector width. None is detected on the first embed. It must match the --dim the database was migrated with. |
| api_key | apiKey | SecretStr / Secret | None | None | CE_EMBEDDING__API_KEY | Provider API key. A secret field: see Secret fields. |
| base_url | baseUrl | str | None | None | CE_EMBEDDING__BASE_URL | The endpoint root. Required for azure_openai (the Azure v1 path) and custom (any OpenAI-compatible /embeddings); optional for openai. |
| max_retries | maxRetries | int ≥ 0 | 2 | CE_EMBEDDING__MAX_RETRIES | Retries on top of the first attempt, applied by the built-in client (408, 409, 429, 5xx and transport failures, Retry-After honoured). 0 disables retrying. A host embedder is never retried. |
| tokenizer | tokenizer | "auto" | "exact" | "estimate" | "auto" | CE_EMBEDDING__TOKENIZER | How tokens are counted before requests are packed. Python: auto uses tiktoken cl100k only if its encoding is already cached (never downloads), exact may download, estimate never imports tiktoken. TypeScript: the encoding ships with js-tiktoken, so auto and exact both count exactly. |
| max_batch_items | maxBatchItems | int > 0 | None | None | CE_EMBEDDING__MAX_BATCH_ITEMS | Items per request. None uses the provider default in the table below. |
| max_batch_tokens | maxBatchTokens | int > 0 | None | None | CE_EMBEDDING__MAX_BATCH_TOKENS | Tokens per request. None uses the provider default. |
| max_input_tokens | maxInputTokens | int > 0 | None | None | CE_EMBEDDING__MAX_INPUT_TOKENS | Largest single chunk. A chunk over it raises rather than being split. None uses the provider default. |
| token_bytes_ratio | tokenBytesRatio | float > 0 | 3.0 | CE_EMBEDDING__TOKEN_BYTES_RATIO | Estimator only: UTF-8 bytes per token. |
| safety_margin | safetyMargin | float in (0, 1] | 0.85 | CE_EMBEDDING__SAFETY_MARGIN | Estimator only: headroom applied to the batch budget (never to max_input_tokens). |
| max_concurrency | maxConcurrency | int ≥ 1 | 4 | CE_EMBEDDING__MAX_CONCURRENCY | Embedding requests in flight per document. |
Provider request limits
The defaults used when max_batch_items, max_batch_tokens or max_input_tokens is left unset. Set one to override it, for example against a self-hosted server with different limits.
| Provider | Items per request | Tokens per request | Tokens per input |
|---|
| openai, azure_openai, custom (and a host embedder) | 256 | 250,000 | 8,192 |
LLM
LLMConfig is used in three places, each with its own environment prefix: llm (CE_LLM__), vision_llm (CE_VISION_LLM__) and graph.extraction_llm (CE_GRAPH__EXTRACTION_LLM__). The table shows the llm names. All three are optional when the engine is given a host model (see Bring your own model), which then makes every model call.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| provider | provider | "openai" | "azure_openai" | "custom" | required | CE_LLM__PROVIDER | Which OpenAI-compatible endpoint the built-in client calls. See Providers. |
| model | model | str | required | CE_LLM__MODEL | The endpoint’s model name. |
| api_key | apiKey | SecretStr / Secret | None | None | CE_LLM__API_KEY | API key, sent as Authorization: Bearer, or in the api-key header on azure_openai. |
| base_url | baseUrl | str | None | None | CE_LLM__BASE_URL | The endpoint root: required for custom and azure_openai, optional for openai (a gateway). |
| max_retries | maxRetries | int ≥ 0 | 2 | CE_LLM__MAX_RETRIES | Retries on top of the first attempt, applied by the built-in client. The call timeout (240 seconds) is per attempt. A host model is never retried. |
| schema_enforced | schemaEnforced | bool | None | None | CE_LLM__SCHEMA_ENFORCED | Does this endpoint enforce response_format json_schema? Unset, only OpenAI proper and Azure are trusted to; every other endpoint is reported prompt and validated. True trusts the endpoint you point at; False validates everywhere. |
Providers
The built-in client is one OpenAI-compatible client over plain HTTP (POST {base_url}/chat/completions, POST {base_url}/embeddings); no provider SDK is installed. What each provider value reads from its config. Provider names are the same in both packages. For any other model or embedding API, pass model / embedder to the engine (see Bring your own model).
provider names the wire format, not the vendor. anthropic, gemini, vertex_ai and bedrock are refused at construction with a message naming the way forward. Every vendor still works, through a host model or through custom with its OpenAI-compatible root: the per-vendor table is under Use any model provider.
| Provider | Used for | Fields it reads | Notes |
|---|
| openai | embeddings, LLM | api_key; base_url optional | https://api.openai.com/v1 unless base_url says otherwise. The one endpoint, with Azure, on which the built-in client claims an enforced JSON schema. |
| azure_openai | embeddings, LLM | api_key (the Azure key), base_url (the Azure v1 path, https://<resource>.openai.azure.com/openai/v1/) | The key travels in the api-key header instead of Authorization: Bearer. |
| custom | embeddings, LLM | base_url; api_key if the server wants one | Any OpenAI-compatible root: OpenRouter, Groq, DeepSeek, Together, Fireworks, Mistral, Ollama, vLLM, LM Studio, a LiteLLM proxy, Gemini’s .../v1beta/openai/ (beta) and Anthropic’s .../v1/ (for trying it out; it ignores response_format). Reported prompt and validated unless schema_enforced says otherwise. |
Endpoints do not follow redirects. A base_url that answers with a 3xx fails the call rather than sending the key and the prompt to the redirect target. Point it at the final address.
Graph
GraphConfig, off by default. Environment prefix CE_GRAPH__. Migrate with context-engine migrate --graph before turning it on. Python needs the [graph] extra; TypeScript needs graphology and graphology-communities-louvain, plus neo4j-driver only with a Neo4j.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| enabled | enabled | bool | False | CE_GRAPH__ENABLED | Turns the graph subsystem on. Needs extraction_llm or a host model on the engine; the engine raises at construction when it has neither. |
| neo4j_uri | neo4jUri | str | None | None | CE_GRAPH__NEO4J_URI | Optional Neo4j backend for graph search (Neo4j 5.1 or later). Unset (the default), graph expansion and connectivity scoring run in Postgres; set, they run in Neo4j. Entities, relationships, communities and navigation run in Postgres either way. Each graph ingest then also writes that document’s own graph to Neo4j; a write that fails is reported and replayed by resync_graph() / resyncGraph(). See Neo4j is kept in step. |
| neo4j_user | neo4jUser | str | "neo4j" | CE_GRAPH__NEO4J_USER | Neo4j user. |
| neo4j_password | neo4jPassword | SecretStr / Secret | None | None | CE_GRAPH__NEO4J_PASSWORD | Neo4j password. URI and password are all-or-nothing. |
| neo4j_database | neo4jDatabase | str | "neo4j" | CE_GRAPH__NEO4J_DATABASE | Neo4j database name. |
| neo4j_connection_timeout_s | neo4jConnectionTimeoutS | float > 0 | 10 | CE_GRAPH__NEO4J_CONNECTION_TIMEOUT_S | Seconds the Neo4j driver waits to open one connection. A write that fails to connect is retried once, so an ingest or delete against an unreachable Neo4j costs seconds, not minutes. |
| neo4j_connection_acquisition_timeout_s | neo4jConnectionAcquisitionTimeoutS | float > 0 | 15 | CE_GRAPH__NEO4J_CONNECTION_ACQUISITION_TIMEOUT_S | Seconds the Neo4j driver waits to take a connection from its pool. |
| extraction_llm | extractionLlm | LLMConfig | None | None | CE_GRAPH__EXTRACTION_LLM__* | The built-in model that extracts entities and relationships. Required when enabled, unless the engine is given a host model, which then makes these calls too. |
| community_summaries | communitySummaries | bool | False | CE_GRAPH__COMMUNITY_SUMMARIES | Ask the model for a short summary of each detected community and embed it. Off by default: detection (Louvain, no model call) always runs, and summaries cost one model call plus one embedding batch per source on every graph ingest. Off, communities are stored without a summary, resume_documents skips its community pass, community_match contributes nothing, community_summary is advertised as off, no community unit is billed, and the next graph ingest or rebuild of a source drops its existing summaries. Turn it on later and the next resume_documents fills them. In TypeScript, unset or null means off. |
| rerank_weights | rerankWeights | dict[str, float] | None | None (built-in table) | CE_GRAPH__RERANK_WEIGHTS | Weights for the graph leg’s five signals: vector_score 0.3, entity_match 0.3, relationship_relevance 0.2, community_match 0.1, graph_connectivity 0.1. A partial map overrides only the keys it names; an unknown key raises. TypeScript also accepts the Python spelling rerank_weights; rerankWeights wins when both are set. |
| max_query_entities | maxQueryEntities | int > 0 | 30 | CE_GRAPH__MAX_QUERY_ENTITIES | Ceiling on the leg’s entity set (matched entities plus one hop). Matched entities are never evicted by hops. |
| entity_max_name_words | entityMaxNameWords | int > 0 | 5 | CE_GRAPH__ENTITY_MAX_NAME_WORDS | Longest entity name, in words, the query is probed for. |
| entity_min_name_chars | entityMinNameChars | int > 0 | 3 | CE_GRAPH__ENTITY_MIN_NAME_CHARS | Query n-grams shorter than this never become candidates. |
| entity_match_threshold | entityMatchThreshold | float in [0, 1] | None | 0.6 | CE_GRAPH__ENTITY_MATCH_THRESHOLD | word_similarity floor for the fuzzy fallback, which runs only when the exact pass found nothing. None turns the fallback off (it scans the entity table). |
| entity_merge_threshold | entityMergeThreshold | float in [0, 1] | 0.75 | CE_GRAPH__ENTITY_MERGE_THRESHOLD | Trigram similarity at which a new entity name merges into an existing one of any type, at ingest. Two names that differ only in their digits (“invoice 10234” and “invoice 10235”) are never merged, whatever the score. |
| entity_merge_fuzzy_threshold | entityMergeFuzzyThreshold | float in [0, 1] | 0.45 | CE_GRAPH__ENTITY_MERGE_FUZZY_THRESHOLD | The looser merge step: same entity type and a shared word of four or more characters. The same digit rule applies. |
| entity_fuzzy_merge | entityFuzzyMerge | "words" | "similarity" | "off" | "words" | CE_GRAPH__ENTITY_FUZZY_MERGE | How a name that is not an exact or normalised match may still join a stored entity. "words": every word of one name must appear in the other, so “Acme Corp” joins “Acme Corp.” and “Northgate Logistics” never joins “Northwind Logistics”. "off": exact or normalised names only. "similarity": the trigram scores alone. |
| entity_fuzzy_merge_across_sources | entityFuzzyMergeAcrossSources | bool | False | CE_GRAPH__ENTITY_FUZZY_MERGE_ACROSS_SOURCES | Whether a fuzzy match may join an entity mentioned only in other sources. Off: a fuzzy match is taken from the document’s own source. An exact match links across sources either way. |
To ask whether the Neo4j backend is configured, use config.graph.neo4j_configuredgraphNeo4jConfigured(config.graph).
Reranker
Reranking is bring-your-own: ContextEngine(config, reranker=fn) / new ContextEngine(config, { reranker }) with fn(query, docs) returning indices into docs, most relevant first, sync or async. There is no built-in rerank provider. RerankerConfig (environment prefix CE_RERANKER__) holds only the window; naming a provider field on it raises at construction. The engine drops out-of-range and duplicate indices and appends anything omitted, so a reranker can only reorder hits; one that raises ends that search on the fused order. Hits are masked by the redaction policy before the reranker sees them.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| candidates | candidates | int ≥ 1 | 50 | CE_RERANKER__CANDIDATES | How many fused hits are handed to the host reranker. Also widens how many candidates each leg is asked for. |
Fusion
FusionConfig. Environment prefix CE_FUSION__.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| method | method | "rrf" | "rrf" | CE_FUSION__METHOD | Reciprocal rank fusion. The only method. |
| k | k | int ≥ 1 | 20 | CE_FUSION__K | The RRF constant: a leg’s rank r adds its weight divided by k + r, so a smaller k gives the top of each list more say. Measured with the repository’s retrieval benchmark: against the textbook 60, 20 raised hybrid nDCG@10 by 0.04 to 0.05 on four of five public datasets, with no measurable change on the fifth. |
| weights | weights | dict[str, float] | fts 1.0, trgm 0.4, ann 1.0, graph 1.0 | CE_FUSION__WEIGHTS | Per-leg weight: full-text, trigram, vector and graph. Replacing the map replaces all four, so name every leg. The trigram weight was measured with the repository’s retrieval benchmark: against 0.8, 0.4 raised nDCG@10 by 0.04 to 0.07 on four of five public datasets, with no measurable change on the fifth. CE_FUSION__WEIGHTS__TRGM=0.8 sets one leg from the environment. A weight of 0 turns a leg off and its query is not sent. For the vector leg the query is not embedded either, unless compression needs the vector. For the graph leg the graph stage never starts, even with mode="graph": the search is billed as hybrid and reports the reason graph_weight_zero. A custom retrieval leg passed to the constructor is weighted under its own name: 1.0 when absent, and 0 means it is never called. |
Search
SearchConfig: how much retrieval reacts to the shape of the query. Environment prefix CE_SEARCH__. With both fields unset, retrieval uses no vector floor and the query-shaped trigram threshold.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| vector_floor | vectorFloor | "adaptive" | None | None (off) | CE_SEARCH__VECTOR_FLOOR | "adaptive" drops vector candidates under a cosine floor chosen from the query’s shape. Off by default because similarity bands differ between embedding models; measure on your own corpus first. Python validates assignment, so a typo raises instead of silently reading as off. |
| trgm_limit | trgmLimit | float in [0, 1] | None | None | CE_SEARCH__TRGM_LIMIT | Pins pg_trgm’s similarity threshold for every query. None uses the query-shaped rule (0.45 for a two-character CJK query down to 0.15 for a long sentence). |
| lexical_match | lexicalMatch | "all" | "any" | "all" | CE_SEARCH__LEXICAL_MATCH | How the full-text leg combines the query’s words. "all" requires every word, stopwords included, so a natural-language question often matches nothing. "any" matches a chunk containing any word left after dropping stopwords and lets the rank order the matches: more candidates, more weight on ranking and fusion. Measured on five public datasets, "any" raised nDCG@10 on multi-hop questions and lowered it on argument retrieval, so "all" stays the default; measure on your own queries before switching. Quoted phrases and -term keep their meaning in both: phrases are ANDed with the other words in "all" and ORed in "any", and a -term always excludes. Words are split on ASCII whitespace. |
| lexical_rank | lexicalRank | "ts_rank_cd" | "ts_rank" | "bm25" | "ts_rank_cd" | CE_SEARCH__LEXICAL_RANK | How full-text matches are ordered. ts_rank_cd (cover density) rewards query words that occur close together; ts_rank counts how often they occur anywhere in the chunk. bm25 (experimental) is Okapi BM25 computed in SQL: rare words count for more, repetition saturates, long chunks are normalised. It orders the same matches lexical_match selects, and the permission check applies to every scored row. It is opt-in per database: run context-engine bm25-enable once, and until it has finished a bm25 search ranks with ts_rank_cd and reports a LexicalRankFallback to the error hook. A database that never enables it sees no change to ingest; enabled, the statistics never make two writers wait on each other. Another BM25 engine plugs in as a custom retrieval leg. Measured with the repository’s retrieval benchmark, bm25 with lexical_match="any" raised hybrid nDCG@10 on three of five public datasets and lowered it on one, so it is not a default and does not improve default search; measure on your own queries first. |
| bm25.k1 | bm25.k1 | float ≥ 0 | 1.2 | CE_SEARCH__BM25__K1 | BM25 term-frequency saturation. 0 counts only whether a word is present; larger values let repetition matter longer. Read only when lexical_rank is bm25. |
| bm25.b | bm25.b | float, 0 to 1 | 0.75 | CE_SEARCH__BM25__B | BM25 length normalisation. 1 fully penalises a chunk longer than average, 0 ignores length. Replaces lexical_normalization under BM25. |
| bm25.idf_scope | bm25.idfScope | "source" | "global" | "document" | "source" | CE_SEARCH__BM25__IDF_SCOPE | Where BM25’s corpus statistics (how rare a word is, the average chunk length) come from. source: the chunk’s own source, so one source’s documents never change another’s scores. global: every source together, so other sources’ (other tenants’) documents shift the scores, and each query counts across the whole corpus. document: the chunk’s own document, counted at query time, with no enabling needed. Statistics include documents the caller may not see; no content or id is exposed, and a hidden chunk is never scored or returned. |
| bm25.compact_after_deltas | bm25.compactAfterDeltas | int ≥ 1 | None | 1000 | CE_SEARCH__BM25__COMPACT_AFTER_DELTAS | After a chunk write commits, a source whose BM25 statistics hold more than this many rows is folded into one. It never makes the write wait: a source another maintenance call holds is skipped until a later write. None (TypeScript null) turns it off; then run context-engine bm25-compact now and then. |
| lexical_normalization | lexicalNormalization | int, 0 to 63 | 0 | CE_SEARCH__LEXICAL_NORMALIZATION | The rank function’s normalization bitmask, passed to Postgres unchanged. 0 ignores chunk length; 1 divides by 1 + log(length), 2 by length, 4 by the mean distance between matches (ts_rank_cd only), 8 by unique words, 16 by 1 + log(unique words), 32 maps the rank into 0 to 1. |
| lexical_stopwords | lexicalStopwords | list[str] | None | None (built-in English list) | CE_SEARCH__LEXICAL_STOPWORDS | Words "any" mode drops from the query. None uses LEXICAL_STOPWORDS; an empty list keeps every word. Parsed by the same text-search configuration as the query, so case does not matter. A query made only of stopwords, with no phrase or -term, keeps its words. Ignored in "all" mode. |
| trgm_max_query_words | trgmMaxQueryWords | int ≥ 1 | None | None | CE_SEARCH__TRGM_MAX_QUERY_WORDS | Skips the trigram leg for a query of more than this many words, exactly as a fusion weight of 0 would for that query: it is not sent or fused, and it is absent from the leg hits and timings. A word is a run of characters that are not Unicode White_Space (query_words / queryWords), so a query in a script written without spaces counts as one word. None runs the leg for every query, and that default is unchanged. The trigram leg sets the search time of a long question: on sentence-length questions (16 words at the median) over 100,646 chunks, for a caller who sees 10% of the corpus, the median hybrid search took 2,563 ms unset, 2,160 ms at 16 (53% of the questions still ran the leg) and 113 ms at 8, on one laptop with one query at a time. The results change with it: at 8, 92.5% of the default’s results were still in the top ten. Permissions are unaffected. Measured on five public datasets, 8 or 16 raised hybrid nDCG@10 on three and lowered it by 0.025 on one corpus of long, name-heavy questions, so it ships unset. The trigram leg is for names, codes and misspellings, which are short queries. See the benchmarks. |
| lexical_candidates | lexicalCandidates | int ≥ 1 | None | None | CE_SEARCH__LEXICAL_CANDIDATES | How many candidates the full-text and trigram legs each hand to fusion. None uses the depth every leg uses: ten times the requested result count, between 50 and 500. A larger value is capped at 500. The vector and graph legs are unaffected. |
| plugin_timeout_s | pluginTimeoutS | float | None | None (10 seconds) | CE_SEARCH__PLUGIN_TIMEOUT_S | Seconds every plugin call may take: a custom retrieval leg, a custom fusion function, and each call to a custom graph backend. A leg that times out is a leg outage, a fusion function falls back to RRF, and a graph backend leaves the graph leg with its vector seeds. The built-in Neo4j backend is bounded by the same setting. Must be greater than 0. |
| query_embed_retry_after_cap_s | queryEmbedRetryAfterCapS | float ≥ 0 | 5 | CE_SEARCH__QUERY_EMBED_RETRY_AFTER_CAP_S | The longest the built-in model client waits on a Retry-After while embedding a search query. A rate-limited key then makes the search answer from its lexical legs in seconds. Ingest keeps its full patience. |
| query_embed_max_retries | queryEmbedMaxRetries | int ≥ 0 | 1 | CE_SEARCH__QUERY_EMBED_MAX_RETRIES | Retries for the embedding of a search query. |
Ingest
IngestConfig. Environment prefix CE_INGEST__. The intake caps apply before a document is read or parsed, and they never sample or truncate: over a cap is a refusal.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| max_file_bytes | maxFileBytes | int > 0 | None | 268435456 (256 MiB) | CE_INGEST__MAX_FILE_BYTES | Largest accepted upload. Over it, ingest refuses with IngestTooLarge; nothing is truncated. None / null lifts the limit; 0 is rejected. |
| max_rows | maxRows | int > 0 | None | None | CE_INGEST__MAX_ROWS | Largest tabular document (csv, tsv, xlsx), in rows summed across sheets. None means no row ceiling. |
| max_document_text_chars | maxDocumentTextChars | int > 0 | None | None | CE_INGEST__MAX_DOCUMENT_TEXT_CHARS | Cap on the stored document text. Every ingest records meta_data.text_truncated, and compute and discover refuse a truncated document rather than answer from a partial frame. |
| progress_batches | progressBatches | bool | True | CE_INGEST__PROGRESS_BATCHES | Emit a "progress" event per completed embedding batch, besides the stage’s started / done pair. |
| resume_stale_seconds | resumeStaleSeconds | int ≥ 0 | None | 600 | CE_INGEST__RESUME_STALE_SECONDS | How long a part-embedded document must sit still before resume_embeddings() with no document id takes it over. None lifts the fence. Naming a document id is never fenced. |
| max_park_seconds | maxParkSeconds | int > 0 | None | None | CE_INGEST__MAX_PARK_SECONDS | How long one round of a document parked as waiting_model may wait. It counts the current round only: a document whose latest park is older than this is finalized failed with failure_reason="model_deferred_expired" on the next resume. None waits for ever. See Batch mode. |
Progress events: "progress" is a third state beside started and done. A consumer that switches on the state should ignore one it does not know; progress_batches=False is the escape hatch for one that cannot change.
Storage and pools
StorageConfig. Environment prefix CE_STORAGE__. The pool fields differ by port because they are passed straight through: to sqlalchemy.create_engine in Python and to pg.Pool in TypeScript. Size them for your workers times their concurrent searches, against what the database allows.
| Python | TypeScript | Type | Default | Env var | Meaning |
|---|
| backend | backend | "postgres" | "postgres" | CE_STORAGE__BACKEND | The only backend. |
| ann_exact_threshold | annExactThreshold | int ≥ 0 | 50000 | CE_STORAGE__ANN_EXACT_THRESHOLD | 0 switches the exact vector route off. At or below this many eligible rows (the caller’s permissions together with any source or document filter) the vector leg sorts them exactly; above it, it scans the index, which is approximate. The exact route costs time per query: measured at 384 dimensions on one laptop, the leg took a median 114 ms over 10,056 visible chunks and 337 ms over 48,378, against 18 ms for the index on that last caller. Lower it for wide embeddings, modest hardware or a strict p99; raise it to keep more callers exact. See the benchmarks. |
| write_batch_rows | writeBatchRows | int > 0 | None | 32 | CE_STORAGE__WRITE_BATCH_ROWS | Most chunk rows one write statement carries: the chunk insert, the embedding write and the copy that reuses vectors on a re-ingest. No statement grows with the document or with the embedding batch. Writing a vector also writes its index entry, so on a small database a statement carrying a few hundred vectors can run past a statement_timeout: lower it there, or for wide embeddings; raise it on a large instance to save round trips. Each embedding statement commits on its own, so one that is cancelled loses only its own rows. None (null) puts no cap on a statement. |
| pool_size | none | int ≥ 0 | 5 | CE_STORAGE__POOL_SIZE | Python: connections kept open (SQLAlchemy pool_size). |
| max_overflow | none | int ≥ -1 | 10 | CE_STORAGE__MAX_OVERFLOW | Python: extra connections allowed during a burst, then discarded. |
| pool_timeout | none | float ≥ 0 | 30.0 | CE_STORAGE__POOL_TIMEOUT | Python: seconds a caller waits for a free connection before raising. |
| pool_recycle | none | int ≥ -1 | 1800 | CE_STORAGE__POOL_RECYCLE | Python: seconds after which a pooled connection is reopened. Keep it below the idle timeout of any pooler or proxy in front of Postgres. |
| pool_pre_ping | none | bool | True | CE_STORAGE__POOL_PRE_PING | Python: test each connection with a round trip before handing it out, so a restart or failover does not cost one failed query per pooled connection. |
| none | poolMax | int > 0 | 10 | CE_STORAGE__POOL_MAX | TypeScript: connections in the pg.Pool. |
| none | poolIdleTimeoutMs | int ≥ 0 | 10000 | CE_STORAGE__POOL_IDLE_TIMEOUT_MS | TypeScript: milliseconds an idle connection is kept before it is closed. |
| none | poolConnectionTimeoutMs | int ≥ 0 | 30000 | CE_STORAGE__POOL_CONNECTION_TIMEOUT_MS | TypeScript: milliseconds a caller waits for a connection before the attempt fails. pg has no pre-ping. |
Redaction
redaction takes a RedactionPolicy: a list of rules plus any custom detector functions they reference. An empty policy, the default, does nothing. It holds functions, so it is set in code rather than from the environment.
from context_engine import ContextEngineConfig, RedactionPolicy, RedactionRule
config = ContextEngineConfig(
# ... database_url, embedding ...
redaction=RedactionPolicy(rules=[
RedactionRule(name="email", detector="email"), # mask on output
RedactionRule(name="ssn", detector="ssn", apply_at="both"),
RedactionRule(name="customer", pattern=r"CUST-\d{6}", action="hash"),
]),
redaction_secret_key="...", # used by the hash rule; falls back to secret_key
)import { ContextEngineConfig, RedactionPolicy } from "@promptev/context-engine";
const config = new ContextEngineConfig({
// ... databaseUrl, embedding ...
redaction: new RedactionPolicy({
rules: [
{ name: "email", detector: "email" }, // mask on output
{ name: "ssn", detector: "ssn", applyAt: "both" },
{ name: "customer", pattern: "CUST-\\d{6}", action: "hash" },
],
}),
redactionSecretKey: "...", // used by the hash rule; falls back to secretKey
});Policy fields
rules (a list of RedactionRule) and custom_detectorscustomDetectors, a map from detector name to a function that takes text and returns (start, end) spans. Rule names must be unique, and every detector must be built in or present in the custom map.
Rule fields
| Python | TypeScript | Type | Default | Meaning |
|---|
| name | name | str | required | Unique within the policy. Also the default placeholder, [NAME]. |
| detector | detector | str | None | None | A built-in detector (email, phone, ssn, credit_card, iban, api_key) or a key of custom_detectors. Exactly one of detector / pattern. |
| pattern | pattern | str | None | None | A regular expression. An invalid one raises at construction. |
| field | field | str | None | None | Reserved. Setting it raises, because field-targeted rules are not implemented. |
| action | action | "mask" | "hash" | "remove" | "mask" | hash needs an effective redaction key. |
| placeholder | placeholder | str | None | None | Replacement text for mask. |
| apply_at | applyAt | "ingest" | "output" | "both" | "output" | Where the rule runs. |
| unless | unless | list[str] | [] | Principals exempt at output time. Not allowed with apply_at="ingest". |
Presidio detectors
Microsoft Presidio can supply detectors for the custom map. The two ports do it differently:
Python: presidio_detectors(entities, *, analyzer=None, language="en", score_threshold=0.5) from context_engine.redaction_presidio runs Presidio in process. Install the [presidio] extra and a spaCy model (python -m spacy download en_core_web_lg for English).
TypeScript: presidioDetectors(entities, { language, scoreThreshold }) from @promptev/context-engine/presidio calls a Presidio Analyzer HTTP service named by CE_PRESIDIO_URL (or PRESIDIO_URL). Building the detectors throws if neither is set. A failed analyzer call returns no spans.
Secret keys
Two keys, protecting two different things. Neither is required until a feature needs it.
secret_key (CE_SECRET_KEY)
A base64url-encoded 32-byte AES-256-GCM key. It encrypts tool configs that hold secrets: OAuth tokens, API keys, database credentials. Using a secret-bearing tool without it raises an error that names the setting; a key that does not decode to 32 bytes is reported as malformed.
redaction_secret_key (CE_REDACTION_SECRET_KEY)
The HMAC key for hash redaction rules. Unset, hash rules use secret_key. Set it to keep the two apart, for example to give each tenant its own hash key while tool encryption keeps one AES key.
One way to generate a secret_key that both packages accept:
python -c "import base64, secrets; print(base64.urlsafe_b64encode(secrets.token_bytes(32)).decode())"
Changing secret_key strands stored secrets. Tool configs encrypted under the old key do not decrypt. The redaction key is never used for encryption and the encryption key is only used for hashing as the fallback; read the effective hash key with redaction_key(config)redactionKey(config).
Secret fields
Every field that holds a credential is a secret type on the built config, in both packages: database_url, secret_key, redaction_secret_key, every api_key (embedding, llm, vision_llm, graph.extraction_llm) and neo4j_password. You pass plain strings, in code or through the CE_ variables; the config wraps them, and the value is masked in logs and serialisation.
from context_engine import ContextEngineConfig, EmbeddingConfig
config = ContextEngineConfig(
database_url="postgresql://user:pass@localhost:5432/mydb", # plain strings in
embedding=EmbeddingConfig(provider="openai", model="text-embedding-3-small", api_key="sk-..."),
)
print(config.embedding.api_key) # **********
config.embedding.api_key.get_secret_value() # "sk-..."
config.model_dump()["database_url"] # SecretStr('**********')
config.embedding.api_key == "sk-..." # False: compare the read valueimport { ContextEngineConfig, secretValue } from "@promptev/context-engine";
const config = new ContextEngineConfig({
databaseUrl: "postgresql://user:pass@localhost:5432/mydb", // plain strings in
embedding: { provider: "openai", model: "text-embedding-3-small", apiKey: "sk-..." },
});
console.log(config.embedding.apiKey); // **********
JSON.stringify(config); // every secret field is "**********"
config.embedding.apiKey?.get(); // "sk-..."
secretValue(config.embedding.apiKey); // "sk-...", also takes a plain string or nullPython: the fields are pydantic SecretStr. Read one with .get_secret_value(). repr, str, model_dump(), model_dump(mode="json") and model_dump_json() show '**********' in their place, so a host that serialises a config to rebuild it later must dump the secrets itself.
TypeScript: the fields are a Secret, exported from the package root beside secretValue(). Read one with .get() or secretValue(field), which also takes a plain string or null, so a read site works with a value assigned after construction. The value sits in a private field: JSON.stringify, util.inspect, console.log, String(), a template string, Object.keys and a spread never show it. Never String() or template a secret field into a header or a URL, since that hands the mask along as the key. A Secret passed in is kept as is, so a config can be rebuilt from another config’s fields. GraphConfigInit is the graph section’s input type.
Compare the read value, not the field. config.llm.api_key == "sk-..." and config.llm.apiKey === "sk-..." are false.
Validation
These combinations raise when the config is constructed (ValueError in Python, Error in TypeScript). TypeScript messages use the camelCase field names.
| Condition | Message |
|---|
| graph.enabled without extraction_llm and without a host model (raised when the engine is built, not the config: the model is a constructor argument) | graph enabled but neither graph.extraction_llm nor a host model is set |
| neo4j_uri without neo4j_password, or the reverse | graph neo4j_uri set but neo4j_password missing |
| An unknown key in rerank_weights | graph rerank_weights has unknown key(s): ... |
| A RerankerConfig naming enabled, provider, model, api_key or base_url | the built-in reranker was removed: pass reranker= (a function (query, docs) -> ranked indices). |
| An LLM or embedding provider the built-in client does not have (anthropic, gemini, vertex_ai, bedrock, voyage, cohere) | provider '...' was removed from the built-in client. Pass model= ... (each message names the way forward; the repository’s migration notes quote all of them) |
| An unknown field on LLMConfig or EmbeddingConfig (Python) | Extra inputs are not permitted |
| default_mode="graph" with the graph disabled | default_mode is 'graph' but graph enabled is False |
| A hash redaction rule with neither redaction key set | redaction rule 'name': action='hash' requires ... to be set |
Field bounds are enforced too: for example max_retries below 0, safety_margin above 1, or max_file_bytes of 0 are rejected. Redaction rules are checked when they are built: exactly one of detector / pattern, a valid regex, no field, and no unless on an ingest-only rule.
A config error never echoes its input, in either package, because the input is where the keys and passwords live. In Python the message leaves it out, but pydantic’s err.errors() still returns the raw input by design, so log err.errors(include_input=False):
import logging
from pydantic import ValidationError
from context_engine import ContextEngineConfig
log = logging.getLogger(__name__)
try:
config = ContextEngineConfig()
except ValidationError as err:
log.error("bad config: %s", err.errors(include_input=False)) # never the raw inputOther environment variables
Variables outside the ContextEngineConfig scheme that the code reads.
| Variable | Port | Read by |
|---|
| CE_PRESIDIO_URL, PRESIDIO_URL | TypeScript | The Presidio Analyzer endpoint for presidioDetectors. CE_PRESIDIO_URL is checked first. |
| CE_COMPAT_DATABASE_URL | Python (repository test) | Not read by the package. It points the managed-Postgres preflight test in the engine repository (tests/test_managed_postgres_compat.py) at a scratch database to check pgvector, pg_trgm and unaccent before you commit to a provider. It runs the real migrations. |
| TIKTOKEN_CACHE_DIR, DATA_GYM_CACHE_DIR | Python | Where tokenizer="auto" looks for a cached cl100k encoding. Pre-seed it for exact counts offline. |
The CLI’s mcp command builds its config from these same CE_ variables.