CertKeen

Claude Certified Architect — Foundations · Free practice question 8 of 10

Execution-based evals for SQL generation

Your team is adding a new prompt for SQL generation and wants to set up an eval harness before shipping. The team can hand-label about 80 (question, expected-SQL) pairs over the next two weeks. Which eval setup gives the most useful signal?

  1. A.Score generated SQL with string-equality against the labeled expected SQL.
  2. B.Execute both the generated and expected SQL against a known dataset and compare result sets.
  3. C.Have Claude score its own generated SQL for correctness on a 1–10 scale.
  4. D.Track only latency and cost in production; trust the model to be correct.
Show answer and explanation

Correct answer: B. Execute both the generated and expected SQL against a known dataset and compare result sets.

Why: Two SQL statements can differ in syntax (alias names, join order, whitespace) but produce identical results — and result-correctness is what users care about. String equality is too brittle; self-scoring leaks the eval signal back into the system under test; production-only metrics don't measure correctness.

More free Claude Certified Architect — Foundations questions