Agent context
Context engineering for every agent: which sources are indexed, how they are chunked and retrieved, and how the window is actually spent on a turn.
Indexed chunks
3,530
across 5 sources
Window used
50%
7,936 of 16,000
Grounding rate
88%
answers with a citation
Recall @8
94%
on the eval set
Context window budget
Where the 16,000 token budget goes on a typical onboarding turn.
System prompt420 · 3%
Tool definitions1,180 · 7%
Retrieved policy chunks3,240 · 20%
Customer 360 record860 · 5%
Conversation history2,140 · 13%
Current user turn96 · 1%
Headroom8,064 tokens
Onboarding policy v9
nightlyDocument corpus · 1840 chunks · 2.1M tokens indexed
Retrieval config
Hybrid search, reranked before it reaches the model.
Embedding modelbge-m3 (multilingual)
Chunk size512 tokens, 64 overlap
Strategyhybrid BM25 + dense
Top k8, reranked to 4
Rerankerbge-reranker-v2-m3
Min score0.62
Retrieval quality
Measured on a 240-question eval set.
Recall @894%
Precision @481%
Grounding rate88%
Stale chunk rate6%
Context policy
Governance applied at retrieval time, not after.
Persona scoping
retrieval respects the caller’s access rights
PII masking at index
entities pseudonymised before embedding
Freshness ceiling
chunks older than 30 days are down-weighted
Citation required
answers without a source are blocked
Cross-domain reads
blocked unless explicitly granted