CertKeen

Claude Certified Architect — Foundations · Free practice question 9 of 10

Prompt caching for repeated context

A documentation Q&A app sends the same 90 KB knowledge-base prelude with every user query. Cost has grown linearly with query volume. Which is the most appropriate first-pass cost optimization?

  1. A.Switch to a smaller Claude model and accept some quality regression.
  2. B.Enable prompt caching on the static prelude so identical preludes don't re-tokenize on every call.
  3. C.Move the knowledge base into a vector store and retrieve only relevant chunks per query.
  4. D.Compress the knowledge base into a smaller summary that Claude reads each time.
Show answer and explanation

Correct answer: B. Enable prompt caching on the static prelude so identical preludes don't re-tokenize on every call.

Why: The prelude is stable across queries — prompt caching reduces both input cost and latency on hits while keeping full context. Retrieval is a larger architectural change with new correctness risks; summarization loses information; switching models trades quality you may not need to give up.

More free Claude Certified Architect — Foundations questions