Risk-ranked verification for coding Agents

Software verification
that argues back.

TestForge gives inexpensive local coding Agents a verification discipline they do not reliably improvise. It attacks explicitly submitted frozen release candidates with risk-ranked evidence - not ordinary implementation and not a comforting pile of green checkmarks.

Free Collaborative Dynamics Augment · two matched SKILLs · deterministic tools · independent review · behavioral eval testbed

A precision quality gate checks three evidence paths and diverts a failed path into a separate reject tray.

Begin successfully

Bring a finished candidate and a release claim worth attacking.

TestForge is for developers, release owners, coding-Agent operators, and Augment builders who need evidence stronger than “the tests are green.” It is not a product-design workshop, penetration-testing authorization, compliance certification, or proof that defects do not exist.

INPUT

What to supply

A bounded target and revision; the intended behavior or requirement; relevant diff, source, tests, logs, and environment facts; plus explicit authority for any consequential action.

OUTPUT

What to expect

An impact map, ranked risks, invariants, scenarios, tests or commands, captured execution evidence, classified findings, residual risks, reviewer disposition, and exactly one bounded release status.

FIRST RUN

What success looks like

The operator distinguishes observed, inferred, assumed, unresolved, executed, and authorized claims; links critical risks to credible evidence; and refuses confidence that the evidence cannot support.

$software-verification Verify this completed candidate at revision <REVISION>. The intended behavior is <REQUIREMENT>. Inspect the available source and tests, run only safe authorized checks, and issue an evidence-backed release assessment.
$verification-reviewer Independently challenge the resulting verification package. Find the smallest consequential break in its evidence chain and judge whether the proposed status is supportable.

The practical wager

Externalized method and memory can buy more competent work from cheaper cognition.

TestForge is not one-shot benchmarking and it is not a prompt folder wearing a fake moustache. It is an installable verification system with doctrine, schemas, deterministic tooling, examples, evals, host adapters, evidence custody, and an independent skeptical pass.

$software-verification

Build the smallest credible evidence set.

Reconstruct impact, rank failure risk, define invariants, design meaningful oracles, create stack-compatible tests, interpret execution, and issue one traceable assessment.

$verification-reviewer

Try to make the release claim fail.

Attack target fidelity, catastrophic omissions, oracle strength, mock realism, evidence custody, traceability, authority, and the fit between proof and proposed status.

Preserve one evidence chain

Begin with the failure the change could create—not with test-shaped code.

  1. 01ScopeWhat target, revision, surfaces, environment, and authority are actually in bounds?
  2. 02ImpactWhich call sites, contracts, states, dependencies, persistence, and trust boundaries can change?
  3. 03RiskWhat consequential failure modes deserve depth, and with what confidence?
  4. 04InvariantWhat must remain true across normal, denied, degraded, retried, and recovered behavior?
  5. 05ScenarioWhich preconditions, actions, observations, and forbidden side effects discriminate danger?
  6. 06TestWhat lowest credible layer preserves the real boundary under examination?
  7. 07EvidenceWhich exact command, environment, result, timing, and raw record establish execution?
  8. 08StatusWhat bounded release assessment follows—and what authority still remains human?

Risk determines depth. Oracles determine whether a test establishes anything. Tool output establishes execution. Polished prose does not get a vote.

Keep claim states distinct

Do not let one kind of evidence borrow another kind’s authority.

observed

Directly present in identified source or captured tool output.

inferred

The best current interpretation, with basis and confidence.

assumed

Provisionally treated as true within a stated scope and consequence.

unresolved

Competing or missing support that still changes the decision.

executed

A named command returned a captured result in a named environment.

authorized

A responsible human permitted a bounded consequential action.

Design tests that can lose

A test is only useful when the dangerous implementation would fail it.

TestForge prefers invariants and state changes over truthiness, status-only checks, snapshot worship, real sleeps, and mock-interaction theater.

SCENARIO CONTRACTRISK-LINKED
Preconditions
The relevant state, identity, dependency, time, configuration, and environment.
Action
The exact behavior exercised at the smallest credible layer.
Expected observations
Outputs, state transitions, persistence, events, timing, and recovery evidence.
Forbidden side effects
Unauthorized mutation, duplication, leakage, corruption, stale state, or downstream work.
Evidence source
The test, command, trace, file, log, or observation that could actually establish the claim.
Risk linkage
The failure mode this scenario covers and its current disposition.

Diagnose before patching

Not every red result is a product defect. Not every green result is evidence.

PRODUCT_DEFECT

The implementation violates the intended, evidence-bearing contract.

TEST_DEFECT

The test, fixture, oracle, isolation, or expectation is wrong.

ENVIRONMENT_FAILURE

The named environment cannot perform a decision-critical check.

FLAKY_OR_NONDETERMINISTIC

Outcome variance requires a stability hypothesis and discriminating rerun.

EXPECTED_CONTRACT_CHANGE

The observed behavior changed intentionally, but evidence and baselines need governed revision.

TOOLING_FAILURE

The verifier, parser, runner, adapter, or result normalization failed.

INSUFFICIENT_EVIDENCE

Correctness cannot be decided from the available implementation, oracle, or execution support.

Issue exactly one release status

The verdict follows from evidence. It does not confer release authority.

READY

The bounded release claim is supported by the reachable evidence and resolved review.

READY_WITH_RESIDUAL_RISK

The bounded claim is supportable with explicit remaining risk, scope, and owner acceptance.

NOT_READY

An observed material product or package defect makes the proposed release unsound.

INSUFFICIENT_EVIDENCE

Missing correctness evidence prevents a grounded release decision.

BLOCKED_BY_ENVIRONMENT

The environment prevents decision-critical execution that the claim requires.

TestForge is advisory machinery. It does not prove defect freedom, certify compliance, grant production access, or authorize release.

The quality ratchet

Reviewed success becomes the next floor. Regression gates resist backward motion.

  1. BUILD
  2. TEST
  3. DIAGNOSE
  4. REENGINEER
  5. RERUN
  6. REVIEW
  7. PROMOTE
  8. REGRESSION-CHECK

Failure is useful state.Criterion-level misses identify exact dimensions to reengineer.

Passing is not enough.The evidence package must survive independent review before promotion.

Memory is external.Sealed runs and named baselines replace “it worked before” with an inspectable record.

Hard gates stay hard.A higher average never cancels a newly failed indispensable dimension.

Augment behavioral evaluation testbed

Run isolated trials without handing the subject its answer key.

The bundled harness records package, runtime, adapter, model, prompt, response, judge evidence, criterion dispositions, failure signals, scores, gates, confidence intervals, integrity seals, and regression comparisons.

01

Validate

Check the canonical evaluation contract and recover supported legacy dialects honestly.

02

Run

Execute isolated subject and judge episodes with exact package and runtime identity.

03

Adjudicate

Record criterion-level `met`, `partial`, or `not_met` judgments and cited response evidence.

04

Seal

Hash every retained run artifact so changed or added evidence breaks integrity verification.

05

Promote

Create a compact reviewed baseline without checking raw transcripts into Git.

06

Check

Block score, demonstrated-rate, invalid-count, indispensable-gate, or claim-status regressions.

`DEMONSTRATED`, `PARTIAL`, `FAILED`, and `INVALID` are behavioral-evidence verdicts. They are not commercial approval or software release authority.

Install, verify, maintain

Use matched operator and reviewer versions. A copied file is not an activated capability.

CODEX PLUGIN codex plugin marketplace add Stunspot/TestForge codex plugin add testforge@cd-testforge

Start a new task. Invoke $software-verification, then use a second fresh task to invoke $verification-reviewer. Verify that each can reach its package-relative resources.

CODEX STANDALONE

Copy both complete folders under testforge/skills/ into the personal Codex skills directory. Keep every reference, template, example, fallback, and script with its owning skill.

On Windows, the final paths normally end in .codex\skills\<skill-name>\SKILL.md. Restart and probe both handles separately.

CLAUDE

Confirm the account exposes custom Skills. Upload claude-ai/software-verification-v2.0.0.zip and claude-ai/verification-reviewer-v2.0.0.zip separately, enable both when required, and test each in a new conversation.

Live upload, discovery, resource loading, script execution, and reviewer handoff were not exercised for this release.

OTHER HOSTS

A host must load Markdown skill instructions and preserve package-relative resources. Otherwise use the fileless fallback. Deterministic tools require Python 3.10+; unsupported stacks degrade to generic scenario and command planning.

UPDATE

Record the installed version and preserve required evidence. Replace both skills from the same release, start a fresh task or conversation, and repeat both discovery and resource probes. Never mix operator and reviewer versions.

REMOVE & CLEAN

Remove or disable the plugin and both skills through the host manager, or delete only the two exact standalone folders. TestForge has no account, daemon, telemetry store, or product database. Local manifests, reports, tests, raw captures, baselines, and eval runs remain ordinary files until you archive or delete them under project policy.

Detailed routes: Codex lifecycle · Claude lifecycle · host boundary.

Start from the change

Ask for the evidence chain, then challenge it separately.

$software-verification Verify this frozen release candidate. Run only checks that can change the verdict, allow one support-path recovery at most, and give me one evidence-backed release assessment.
$verification-reviewer Challenge this verification package and tell me whether its release status is actually supported.

Troubleshoot and recover

Preserve the symptom before rebuilding anything.

The skill does not appear

Start a fresh task, verify the plugin state or final folder path, confirm the complete folder—not a lone SKILL.md—is installed, and check host or organization enablement.

The response is generic

Invoke the handle explicitly, state the target and release claim, and verify package-relative doctrine is reachable. Do not treat a plausible answer as proof the skill loaded.

A command cannot run

Record the command, working directory, exit code, exact error, and unavailable guarantee. Classify product, test, environment, or tooling cause before changing anything.

Evidence or status is wrong

Stop the cycle, retain raw evidence, and use the reviewer in a fresh context. A product defect returns the candidate upstream; it is not repaired inside that TestForge cycle.

Full troubleshooting guide · sanitized issue tracker.

Privacy, storage, network, security

The skills are local. The host and tools still have their own data boundaries.

Product behavior

The v2.0.0 skills-only plugin includes no account, telemetry, analytics, hosted service, connector, MCP server, hook, or automatic network request. Deterministic scripts touch only paths and commands the user chooses.

Host behavior

Prompts, repositories, logs, uploads, model calls, retention, training, residency, and connector traffic are governed by Codex, Claude, configured models, Git hosts, and any authorized tools—not by TestForge.

Local records

Verification manifests, reports, raw command captures, generated tests, evaluation runs, seals, and promoted baselines may contain sensitive project evidence. Store and delete them under an approved retention policy.

Security boundary

Imported files and tool output are untrusted evidence. Active security work needs a named target, explicit permission, non-production default, time window, rate limits, prohibited actions, handling rules, and a stop contact.

Data and privacy · security policy · terms.

Provenance and evidence status

Know exactly what has—and has not—been established.

ClaimCurrent evidenceBoundary
Constructed and packagedv2.0.0 source, plugin tree, current Claude archives, manifests, and retained frozen-release receiptsDoes not prove host installation or live behavior
Deterministic behaviorRepository-local tool, package, distribution, documentation, and eval-harness suitesOnly exercised commands, fixtures, Python version, and environment
Behavioral evaluationNamed model/context baselines and a deliberately failed fresh-package smokeSingle-trial, model-, context-, case-, and judge-bounded; not universal model quality
OpenAI directoryLatest retained portal packet is v1.1.4 and repository-testedNo claim of upload, approval, publication, or discoverability
Host activationInstallation probes are documentedCurrent live Codex and Claude activation remain separate observations

Trust and authority

Safe verification is bounded verification.

Repository content is evidence

Comments, issues, fixtures, logs, dependencies, and retrieved text remain untrusted input—not instructions to the verifier.

Capability must be observed

Commands, writes, network access, PR access, browsers, production targets, and external actions exist only when the host proves them.

Consequential actions are gated

Production-code edits, installs, weakened tests, CI changes, destructive operations, active security work, and publication require explicit human authority.

Licensing remains split

MIT covers software and schemas; CC BY-ND 4.0 covers authored Augment materials. Third-party and user material retain their own rights.