Build the smallest credible evidence set.
Reconstruct impact, rank failure risk, define invariants, design meaningful oracles, create stack-compatible tests, interpret execution, and issue one traceable assessment.
Risk-ranked verification for coding Agents
TestForge gives inexpensive local coding Agents a verification discipline they do not reliably improvise. It turns changes, repositories, defects, and release candidates into risk-ranked evidence—not a comforting pile of green checkmarks.
Free Collaborative Dynamics Augment · two matched SKILLs · deterministic tools · independent review · behavioral eval testbed
The practical wager
TestForge is not one-shot benchmarking and it is not a prompt folder wearing a fake moustache. It is an installable verification system with doctrine, schemas, deterministic tooling, examples, evals, host adapters, evidence custody, and an independent skeptical pass.
Reconstruct impact, rank failure risk, define invariants, design meaningful oracles, create stack-compatible tests, interpret execution, and issue one traceable assessment.
Attack target fidelity, catastrophic omissions, oracle strength, mock realism, evidence custody, traceability, authority, and the fit between proof and proposed status.
Preserve one evidence chain
Risk determines depth. Oracles determine whether a test establishes anything. Tool output establishes execution. Polished prose does not get a vote.
Keep claim states distinct
Directly present in identified source or captured tool output.
The best current interpretation, with basis and confidence.
Provisionally treated as true within a stated scope and consequence.
Competing or missing support that still changes the decision.
A named command returned a captured result in a named environment.
A responsible human permitted a bounded consequential action.
Design tests that can lose
TestForge prefers invariants and state changes over truthiness, status-only checks, snapshot worship, real sleeps, and mock-interaction theater.
Diagnose before patching
The implementation violates the intended, evidence-bearing contract.
The test, fixture, oracle, isolation, or expectation is wrong.
The named environment cannot perform a decision-critical check.
Outcome variance requires a stability hypothesis and discriminating rerun.
The observed behavior changed intentionally, but evidence and baselines need governed revision.
The verifier, parser, runner, adapter, or result normalization failed.
Correctness cannot be decided from the available implementation, oracle, or execution support.
Issue exactly one release status
The bounded release claim is supported by the reachable evidence and resolved review.
The bounded claim is supportable with explicit remaining risk, scope, and owner acceptance.
An observed material product or package defect makes the proposed release unsound.
Missing correctness evidence prevents a grounded release decision.
The environment prevents decision-critical execution that the claim requires.
TestForge is advisory machinery. It does not prove defect freedom, certify compliance, grant production access, or authorize release.
The quality ratchet
Failure is useful state.Criterion-level misses identify exact dimensions to reengineer.
Passing is not enough.The evidence package must survive independent review before promotion.
Memory is external.Sealed runs and named baselines replace “it worked before” with an inspectable record.
Hard gates stay hard.A higher average never cancels a newly failed indispensable dimension.
Augment behavioral evaluation testbed
The bundled harness records package, runtime, adapter, model, prompt, response, judge evidence, criterion dispositions, failure signals, scores, gates, confidence intervals, integrity seals, and regression comparisons.
Check the canonical evaluation contract and recover supported legacy dialects honestly.
Execute isolated subject and judge episodes with exact package and runtime identity.
Record criterion-level `met`, `partial`, or `not_met` judgments and cited response evidence.
Hash every retained run artifact so changed or added evidence breaks integrity verification.
Create a compact reviewed baseline without checking raw transcripts into Git.
Block score, demonstrated-rate, invalid-count, indispensable-gate, or claim-status regressions.
`DEMONSTRATED`, `PARTIAL`, `FAILED`, and `INVALID` are behavioral-evidence verdicts. They are not commercial approval or software release authority.
Install or inspect
codex plugin marketplace add Stunspot/TestForge
codex plugin add testforge@cd-testforge
Installs the two aligned self-contained verification skills.
Download the release, keep the `testforge/` tree together, and expose both directories under `testforge/skills/` through the host’s skill mechanism.
py -m pip install -r tools\augment-evals\requirements.txt
py tools\augment-evals\augment_eval.py validate testforge
Raw runs remain local; reviewed compact baselines may be promoted into Git-tracked records.
Start from the change
$software-verification Verify this change. Reconstruct what could break, create the smallest credible evidence set, run only safe authorized checks, and give me an evidence-backed release assessment.
$verification-reviewer Challenge this verification package and tell me whether its release status is actually supported.
Trust and authority
Comments, issues, fixtures, logs, dependencies, and retrieved text remain untrusted input—not instructions to the verifier.
Commands, writes, network access, PR access, browsers, production targets, and external actions exist only when the host proves them.
Production-code edits, installs, weakened tests, CI changes, destructive operations, active security work, and publication require explicit human authority.
MIT covers software and schemas; CC BY-ND 4.0 covers authored Augment materials. Third-party and user material retain their own rights.