Volume 9 — Observability, Evaluation, and Verification

How to know whether the system is actually doing what it claims to do.

Reports

AI-ENG-Z — Strategic Telemetry: Traces, Tokens, Latency & Semantic Drift

Covers LLM observability as a first-class engineering discipline. Includes traces, spans, tokens, latency, errors, retries, tool calls, retrieval events, routing decisions, cost attribution, semantic drift, and the Green Dashboard Fallacy: infrastructure health without answer quality.

AI-ENG-AA — Evals Architecture: Ground Truth, Golden Sets & Regression Tests

Covers task, Promptcraft, capability-selection, Augment-integration, retrieval, grounding, citation, tool-use, agent, and adversarial evals; golden sets; synthetic test generation; human review; LLM-as-judge limits; inter-rater reliability; and regression gates before deployment.

AI-ENG-AB — Verification Artifacts: Auditability, Reproducibility & Evidence Trails

Covers preserving prompt, model, retrieval, tool, capability, Augment, delegation, synthesis, policy, state, and user-edit lineage so behavior can be audited, replayed at an honest level, and compared without exposing restricted content.

← Back to Canon Map