Context and knowledge
What is a context engine?
In short
A context engine is the layer that decides what information an AI model sees for each request: it finds relevant documents and facts, checks the asker may see them, ranks and trims them, and hands the model a compact, cited context. Context engineering is the wider practice of designing everything that goes into a model's context window, including instructions, retrieved knowledge, tool results and memory.
Every AI model answers from what is in its context window: the instructions, the conversation, and whatever documents or data were put in front of it for this request. A model that is given the wrong context gives the wrong answer, however capable it is.
A context engine is the part of an AI system whose job is to get that context right. It sits between your data and the model, and for each question it decides what the model gets to see.
What does a context engine do?
A context engine typically handles five jobs for every request:
- Find candidates. Search documents, records and other knowledge for material related to the question, often with more than one search method.
- Apply permissions. Remove anything the person asking is not allowed to read, ideally inside the search itself. See RAG security.
- Rank and fuse. Combine results from different methods into one ordered list, so the best passages come first.
- Trim to fit. Keep only what is useful, so the context stays small, fast and cheap.
- Keep sources. Track where each passage came from, so the answer can cite it and people can check it.
The output is a short, relevant, permission-checked bundle of text that the model reads alongside its instructions.
What is context engineering?
Context engineering is the practice of designing everything a model receives, not only the prompt wording. Anthropic describes it as curating the smallest set of high-signal information that lets the model do the job (Effective context engineering for AI agents).
For an AI agent, the context usually contains:
- Instructions: the agent’s role, rules and tone.
- Knowledge: passages retrieved from documents and data.
- Tools: descriptions of the actions the agent can take, and the results of the ones it already took.
- History and memory: the current conversation and, sometimes, relevant earlier ones.
- The request itself: including files the person attached.
Each of these competes for space. Too little and the agent guesses. Too much and it gets slower, costs more, and can lose track of what matters. A context engine is the tool that makes the knowledge part of this balance manageable.
How is a context engine different from RAG, a vector database or a knowledge graph?
These terms are often used as if they mean the same thing. They describe different layers.
| What it is | What it stores or does | Handles permissions? | Typical role | |
|---|---|---|---|---|
| Vector database | A store that finds items with similar embeddings (number vectors that represent meaning) | Vectors plus metadata; similarity search | Only as metadata filters you define | One search method |
| Knowledge graph | A store of entities and their relationships | Nodes and edges, such as “contract A is signed by company B” | Only if you model it | One search method, good for connected questions |
| RAG | A pattern: retrieve, then generate | Nothing by itself; it describes a flow (Lewis et al., 2020) | Depends on the retrieval step | The overall approach |
| Context engine | The system that runs retrieval for AI | Combines search methods, permissions, ranking, trimming and citations | Yes, as a core job | The layer that feeds the model |
A simple way to remember it: RAG is the recipe, vector databases and knowledge graphs are ingredients, and the context engine is the kitchen that puts them together for each order.
Why is vector search alone not enough?
Vector search finds text with similar meaning. That is powerful, but it has gaps:
- Exact terms. Product codes, invoice numbers and names are often found better by keyword search.
- Spelling and variants. Fuzzy (trigram) matching catches typos and partial names.
- Relationships. “Which suppliers are linked to contracts that expire this quarter” spans several documents; a graph helps.
- Permissions. Similar documents are often grouped by topic, and so are permissions. Filtering after the search can remove most of the nearest matches. In our benchmark on 100,000 documents with synthetic, team-shaped permissions, a hybrid search filtered afterwards returned nothing for 64% of searches when the asker could see 10% of the documents. The write-up explains why.
This is why many context engines use hybrid search: several methods run together and their results are fused into one ranking. Measure it against vector search alone on your own queries, because fusion can also push good matches down.
What should you look for in a context engine?
- Permission-aware retrieval that applies the asker’s access inside every search method.
- Hybrid search: keyword, fuzzy and vector, with an optional graph.
- Source sync from the places your documents live, with per-file control over what is indexed.
- Citations on every answer.
- Data control: where documents and embeddings are stored, and whether you can use your own database.
- Evaluation: a way to measure recall with realistic data and permissions, not only demos.
- Clear limits: no retrieval system is perfect; you should be able to see what was retrieved for a given answer.
For the security side of these choices, read RAG security and AI data leakage.
Further reading
- Context layer vs knowledge graph vs RAG vs agent memory: what each one stores and where each fails, with a decision table.
- RAG alternatives in 2026: long context, fine-tuning, graphs, memory and tool calls compared.
How promptev handles context
- The promptev context engine is open source under Apache-2.0, available as a Python package and an npm package that share one schema.
- It runs on PostgreSQL with pgvector, and fuses full-text, trigram and vector search (plus an optional graph) into one ranking.
- Every search leg runs under the same permission check, so the asker’s identity and groups are applied inside the search.
- In the promptev app, agents follow Google Drive, SharePoint, OneDrive and Dropbox sharing, and answers cite their sources.
- Workspaces can point documents, embeddings and retrieval at their own PostgreSQL with their own embedding model. See the context engine page.
Frequently asked questions
Is a context engine the same as RAG?
No. RAG is a pattern: retrieve documents, then generate an answer from them. A context engine is the system that does the retrieval part well, usually combining several search methods, permissions, ranking and source tracking.
What is the difference between prompt engineering and context engineering?
Prompt engineering focuses on how you word the instructions. Context engineering covers everything the model receives: instructions, retrieved documents, tool definitions and results, conversation history and memory, and how much of each fits in the context window.
Do I need a vector database to build a context engine?
Usually you need vector search, but it does not have to be a separate database. Many teams add vector search to a database they already run, such as PostgreSQL with the pgvector extension.
Why not just put all documents in a long context window?
Long windows cost more per request, are slower, and models tend to use very long contexts less reliably. It also ignores permissions: everything in the window is visible to the model, whoever is asking.
Where does a knowledge graph fit in?
A knowledge graph stores entities (people, products, contracts) and how they relate. A context engine can use it as one retrieval method, useful for questions that span several documents or ask about connections.

Faisal Saeed is Founder & CEO of Promptev, building next-gen context engineering infrastructure that enables teams to orchestrate, scale, and deploy production-ready generative AI systems with confidence.