| extract | async (content, filename, mime, *, vision_llm, hooks, extraction=None, http_client_factory=None, max_rows=None) -> Extracted | Text extraction for one file. Degrades instead of raising, except IngestTooLarge over max_rows. Extracted carries text, pages, slides, pages_source and more. |
| run_search | async (query, *, config, hooks, backend, embedder, session_factory, source_ids=None, document_ids=None, principals=None, top_k=10, mode="hybrid", ...) -> SearchResult | The retrieval pipeline under engine.search. Here principals=None means trusted with no warning; prefer the engine method. |
| rrf_fuse | (ranked_lists: dict[str, list[str]], *, k: int, weights: dict[str, float]) -> list[tuple[str, float]] | Reciprocal rank fusion. A leg missing from weights counts 1.0; a leg weighted 0.0 is ignored. |
| build_embedder | (cfg: EmbeddingConfig, *, http_client_factory=None) -> Embedder | The built-in OpenAI-compatible embedding client. |
| build_llm_client | (cfg: LLMConfig, *, http_client_factory=None) -> LLMClient | The built-in OpenAI-compatible chat client. |
| call_llm | async (cfg, *, system, user, json_mode=False, images=None, max_tokens=None, thinking_budget=None, temperature=None, response_schema=None, schema_name="response", client=None, http_client_factory=None) | One LLM call. Returns (text, {"input", "output"}); with response_schema a reply that also carries mode. |
| host_rerank | async (fn: RerankFn, query, docs) -> list[int] | Calls a host reranker and checks the shape of its reply; RerankFn is the type ContextEngine(reranker=...) takes. |
| encrypt_dict / decrypt_dict | (data: dict, key: bytes) -> str / (token: str, key: bytes) -> dict | AES-256-GCM, base64 of nonce plus ciphertext. A wrong key fails closed. |
| get_secret_key | (config) -> bytes | Decodes config.secret_key (base64url, 32 bytes). RuntimeError when it is missing. |
| CodeRunner | Callable[[str, dict, int], dict | Awaitable[dict]] | Your isolation boundary for compute: called with (code, context, timeout); context["dfs"] holds live DataFrames. Its result is trusted only as far as JSON. |
| compute_over_frames | async (frames: dict[str, DataFrame], instruction, *, config, model_cfg=None, timeout=30, hooks=None, principals=None, documents=None, redaction=None, secret_key=None, code_runner=None, http_client_factory=None, model=None) -> dict | compute over DataFrames you already hold. Same guards and result shape; timeout clamped to 1 to 300 seconds. model is a host function for this call, beating model_cfg and config.llm. |
| compute, get_document_text, list_documents, query_structured | module functions taking session_factory=... | The functions under the engine methods of the same names. Prefer the methods; these take principals=None as trusted without a warning. |
| query_type, is_spaceless, trigram_limit, vector_floor, QueryType | query shape helpers | How the engine reads a query: "exact_match", "conceptual" or "hybrid", and the per-query trigram and vector thresholds. |
| extract_structured_data, resolve_fields, upsert_registry, ExtractionResult | structured extraction | The extraction behind extract_structured=True (never raises; failures come back as quality="failed") and the field registry. |
| StorageBackend, PostgresBackend, ChunkRow, SearchScope | storage | The chunk-plane protocol, its Postgres implementation, a chunk on its way in, and the filters every search leg applies. |
| LEXICAL_STOPWORDS | storage | The built-in English stopword list search.lexical_match="any" drops from a query when search.lexical_stopwords is unset. |
| units_for_file, graph_units, failure_message, FAILURE_REASONS, classify_failure | usage helpers | The unit formulas (PDF and DOCX per page, PPTX per slide, image 1, other per MB; graph ceil(chunks/8) plus the primary communities actually summarised, so none with graph.community_summaries off), and the failure codes with their fixed sentences. |
| emit_usage, emit_error, emit_progress, Hooks | hook plumbing | Report through hooks from your own code; none of them raise. |
| format_result, rows_to_tsv | (result, response_mode="json") / (rows: list[dict]) -> str | The TSV conversion execute_tool uses, for any JSON you already hold. |
| HttpClientFactory, HttpClientPurpose | Callable[[purpose], httpx.AsyncClient | None] | Purposes: llm, embeddings, reranker, tool_http, mcp_oauth. A client you supply is never closed by the engine. |
| Embedder, LLMClient, OpenAICompatClient, OpenAICompatError, OpenAICompatReplyError | classes | The built-in client; build the two wrappers with build_embedder and build_llm_client. |
| BUILTIN_LEGS, RetrievalLeg, RetrievalScope, FusionFn, GraphBackend, RetrievalFailed | plugin types | The plugin seams: the reserved leg names (fts, trgm, ann, graph), a leg's signature and the read-only scope it receives, the fusion signature, the graph backend protocol, and the error a search raises when every leg failed. SearchResult.timings holds each leg's wall time in milliseconds, never part of usage. |
| GRAPH_REASONS | tuple[str, ...] | Why the graph leg did or did not run: not_requested, graph_disabled, graph_unreachable, no_graph_documents, no_entities_in_query, graph_detection_failed, no_visible_graph_chunks, auto, requested, graph_weight_zero. |
| ModelRequest, ModelReply, ModelFn | dataclasses and the host model type | What a host model receives and returns. ModelRequest.key is the request key (below). ModelReply carries text, tokens, mode ("prompt" by default) and refusal. |
| request_key, embedding_request_key | (request: ModelRequest) -> str / (text: str, kind: str, model: str) -> str | The stable batch keys, identical in both ports and on every resume: sha256 hex over the canonical JSON of the request (text purposes post-redaction, vision over the page image bytes), and the per-text embedding key with model your config.embedding.model. |
| ContentSource | Callable[[ParkedDocument], Any] | The content_source callback of resume_documents: returns the original bytes, sync or async, or None. |
| __version__ | str | The installed package version. |