What this prompt does
This prompt turns a vague "my context window keeps overflowing" problem into a structured optimization plan for a specific [use_case] running on [model] with a [context_size] token budget. It feeds the model your [current_problem] and then walks through ten concrete strategies in order: a priority framework, retrieval-based dynamic selection, summarization chains, structured formatting via [format_strategy], cached system-prompt prefixes, few-shot rotation, compression, multi-turn truncation rules, an explicit token split, and an A/B test scored on [eval_metric].
The structure works because it forces the model to treat the context window as a budget rather than an unlimited bucket. The [sys_pct], [ctx_pct], and [out_pct] variables make that budget explicit as percentages that should sum to 100, so the output reasons about trade-offs instead of stuffing everything in. Naming the [model] and [context_size] keeps recommendations realistic for that model's actual limits, and [language] decides whether you get Python, TypeScript, or another implementation with benchmarks. The ten-strategy ordering also matters: it moves from cheap, high-leverage wins like priority ranking and cached prefixes toward more involved work like summarization chains and A/B testing, so you can stop early once your [current_problem] is solved rather than building the whole machine.
When to use it
- You are hitting context limits on large inputs like full PRs, long documents, or extended chat histories
- You are building a retrieval-augmented agent and need to decide what to include per query
- You want to move stable instructions into a cached prefix to cut cost and latency
- You need a defensible token budget split across system prompt, context, and output
- You are losing important detail from earlier turns in long conversations
- You want to A/B test two context strategies before committing one to production
Example output
Expect a structured plan rather than a single answer: a ranked priority framework, a description of the retrieval and summarization approach, a token budget table reflecting your [sys_pct]/[ctx_pct]/[out_pct] split, and runnable [language] code that implements the packing logic. It usually closes with a benchmark harness comparing strategies against your [eval_metric], so you can see which approach actually wins on your data. The code is a starting scaffold rather than a drop-in library, so plan to wire it into your real retrieval and tokenization stack.
Pro tips
- Make
[sys_pct],[ctx_pct], and[out_pct]add up to 100; if they don't, the budget reasoning gets muddy - Set
[context_size]to the model's real limit, not an aspirational one, so summarization triggers fire at the right point - Be specific in
[current_problem]— "losing context from earlier in the conversation" produces better truncation advice than "too slow" - Use
[format_strategy]to match how your model parses best; XML-style tags with priority attributes tend to be easy for it to weight - Pick an
[eval_metric]you can actually measure, like answer accuracy and completeness, or the A/B section stays theoretical - Iterate by re-running with a tighter
[context_size]once the first plan works — it surfaces which compression techniques matter most under real pressure