Volume 9 — Observability, Evaluation, and Verification
How to know whether the system is actually doing what it claims to do.
Reports
AI-ENG-Z — Strategic Telemetry: Traces, Tokens, Latency & Semantic Drift
Covers LLM observability as a first-class engineering discipline. Includes traces, spans, tokens, latency, errors, retries, tool calls, retrieval events, routing decisions, cost attribution, semantic drift, and the Green Dashboard Fallacy: infrastructure health without answer quality.
AI-ENG-AA — Evals Architecture: Ground Truth, Golden Sets & Regression Tests
Covers task, Promptcraft, capability-selection, Augment-integration, retrieval, grounding, citation, tool-use, agent, and adversarial evals; golden sets; synthetic test generation; human review; LLM-as-judge limits; inter-rater reliability; and regression gates before deployment.
AI-ENG-AB — Verification Artifacts: Auditability, Reproducibility & Evidence Trails
Covers preserving prompt, model, retrieval, tool, capability, Augment, delegation, synthesis, policy, state, and user-edit lineage so behavior can be audited, replayed at an honest level, and compared without exposing restricted content.