Claude Certified Architect — Foundations · Free practice question 8 of 10
Execution-based evals for SQL generation
Your team is adding a new prompt for SQL generation and wants to set up an eval harness before shipping. The team can hand-label about 80 (question, expected-SQL) pairs over the next two weeks. Which eval setup gives the most useful signal?
- A.Score generated SQL with string-equality against the labeled expected SQL.
- B.Execute both the generated and expected SQL against a known dataset and compare result sets.
- C.Have Claude score its own generated SQL for correctness on a 1–10 scale.
- D.Track only latency and cost in production; trust the model to be correct.
Show answer and explanation
Correct answer: B. Execute both the generated and expected SQL against a known dataset and compare result sets.
Why: Two SQL statements can differ in syntax (alias names, join order, whitespace) but produce identical results — and result-correctness is what users care about. String equality is too brittle; self-scoring leaks the eval signal back into the system under test; production-only metrics don't measure correctness.
More free Claude Certified Architect — Foundations questions
- Narrowing label definitions to stop over-triggering
- Severity and confidence metadata for review findings
- Tool-based verification to prevent hallucinated APIs
- Precision vs recall trade-offs in extraction
- Evidence fields for auditable structured output
- Tool schemas for strict JSON conformance
- Routing scarce human review capacity
- Prompt caching for repeated context
- Shadow-mode rollout for prompt changes