What to supply
A bounded target and revision; the intended behavior or requirement; relevant diff, source, tests, logs, and environment facts; plus explicit authority for any consequential action.
Risk-ranked verification for coding Agents
TestForge gives inexpensive local coding Agents a verification discipline they do not reliably improvise. It attacks explicitly submitted frozen release candidates with risk-ranked evidence - not ordinary implementation and not a comforting pile of green checkmarks.
Free Collaborative Dynamics Augment · two matched SKILLs · deterministic tools · independent review · behavioral eval testbed
Begin successfully
TestForge is for developers, release owners, coding-Agent operators, and Augment builders who need evidence stronger than “the tests are green.” It is not a product-design workshop, penetration-testing authorization, compliance certification, or proof that defects do not exist.
A bounded target and revision; the intended behavior or requirement; relevant diff, source, tests, logs, and environment facts; plus explicit authority for any consequential action.
An impact map, ranked risks, invariants, scenarios, tests or commands, captured execution evidence, classified findings, residual risks, reviewer disposition, and exactly one bounded release status.
The operator distinguishes observed, inferred, assumed, unresolved, executed, and authorized claims; links critical risks to credible evidence; and refuses confidence that the evidence cannot support.
$software-verification Verify this completed candidate at revision <REVISION>. The intended behavior is <REQUIREMENT>. Inspect the available source and tests, run only safe authorized checks, and issue an evidence-backed release assessment.
$verification-reviewer Independently challenge the resulting verification package. Find the smallest consequential break in its evidence chain and judge whether the proposed status is supportable.
The practical wager
TestForge is not one-shot benchmarking and it is not a prompt folder wearing a fake moustache. It is an installable verification system with doctrine, schemas, deterministic tooling, examples, evals, host adapters, evidence custody, and an independent skeptical pass.
Reconstruct impact, rank failure risk, define invariants, design meaningful oracles, create stack-compatible tests, interpret execution, and issue one traceable assessment.
Attack target fidelity, catastrophic omissions, oracle strength, mock realism, evidence custody, traceability, authority, and the fit between proof and proposed status.
Preserve one evidence chain
Risk determines depth. Oracles determine whether a test establishes anything. Tool output establishes execution. Polished prose does not get a vote.
Keep claim states distinct
Directly present in identified source or captured tool output.
The best current interpretation, with basis and confidence.
Provisionally treated as true within a stated scope and consequence.
Competing or missing support that still changes the decision.
A named command returned a captured result in a named environment.
A responsible human permitted a bounded consequential action.
Design tests that can lose
TestForge prefers invariants and state changes over truthiness, status-only checks, snapshot worship, real sleeps, and mock-interaction theater.
Diagnose before patching
The implementation violates the intended, evidence-bearing contract.
The test, fixture, oracle, isolation, or expectation is wrong.
The named environment cannot perform a decision-critical check.
Outcome variance requires a stability hypothesis and discriminating rerun.
The observed behavior changed intentionally, but evidence and baselines need governed revision.
The verifier, parser, runner, adapter, or result normalization failed.
Correctness cannot be decided from the available implementation, oracle, or execution support.
Issue exactly one release status
The bounded release claim is supported by the reachable evidence and resolved review.
The bounded claim is supportable with explicit remaining risk, scope, and owner acceptance.
An observed material product or package defect makes the proposed release unsound.
Missing correctness evidence prevents a grounded release decision.
The environment prevents decision-critical execution that the claim requires.
TestForge is advisory machinery. It does not prove defect freedom, certify compliance, grant production access, or authorize release.
The quality ratchet
Failure is useful state.Criterion-level misses identify exact dimensions to reengineer.
Passing is not enough.The evidence package must survive independent review before promotion.
Memory is external.Sealed runs and named baselines replace “it worked before” with an inspectable record.
Hard gates stay hard.A higher average never cancels a newly failed indispensable dimension.
Augment behavioral evaluation testbed
The bundled harness records package, runtime, adapter, model, prompt, response, judge evidence, criterion dispositions, failure signals, scores, gates, confidence intervals, integrity seals, and regression comparisons.
Check the canonical evaluation contract and recover supported legacy dialects honestly.
Execute isolated subject and judge episodes with exact package and runtime identity.
Record criterion-level `met`, `partial`, or `not_met` judgments and cited response evidence.
Hash every retained run artifact so changed or added evidence breaks integrity verification.
Create a compact reviewed baseline without checking raw transcripts into Git.
Block score, demonstrated-rate, invalid-count, indispensable-gate, or claim-status regressions.
`DEMONSTRATED`, `PARTIAL`, `FAILED`, and `INVALID` are behavioral-evidence verdicts. They are not commercial approval or software release authority.
Install, verify, maintain
codex plugin marketplace add Stunspot/TestForge
codex plugin add testforge@cd-testforge
Start a new task. Invoke $software-verification, then use a second fresh task to invoke $verification-reviewer. Verify that each can reach its package-relative resources.
Copy both complete folders under testforge/skills/ into the personal Codex skills directory. Keep every reference, template, example, fallback, and script with its owning skill.
On Windows, the final paths normally end in .codex\skills\<skill-name>\SKILL.md. Restart and probe both handles separately.
Confirm the account exposes custom Skills. Upload claude-ai/software-verification-v2.0.0.zip and claude-ai/verification-reviewer-v2.0.0.zip separately, enable both when required, and test each in a new conversation.
Live upload, discovery, resource loading, script execution, and reviewer handoff were not exercised for this release.
A host must load Markdown skill instructions and preserve package-relative resources. Otherwise use the fileless fallback. Deterministic tools require Python 3.10+; unsupported stacks degrade to generic scenario and command planning.
Record the installed version and preserve required evidence. Replace both skills from the same release, start a fresh task or conversation, and repeat both discovery and resource probes. Never mix operator and reviewer versions.
Remove or disable the plugin and both skills through the host manager, or delete only the two exact standalone folders. TestForge has no account, daemon, telemetry store, or product database. Local manifests, reports, tests, raw captures, baselines, and eval runs remain ordinary files until you archive or delete them under project policy.
Detailed routes: Codex lifecycle · Claude lifecycle · host boundary.
Start from the change
$software-verification Verify this frozen release candidate. Run only checks that can change the verdict, allow one support-path recovery at most, and give me one evidence-backed release assessment.
$verification-reviewer Challenge this verification package and tell me whether its release status is actually supported.
Troubleshoot and recover
Start a fresh task, verify the plugin state or final folder path, confirm the complete folder—not a lone SKILL.md—is installed, and check host or organization enablement.
Invoke the handle explicitly, state the target and release claim, and verify package-relative doctrine is reachable. Do not treat a plausible answer as proof the skill loaded.
Record the command, working directory, exit code, exact error, and unavailable guarantee. Classify product, test, environment, or tooling cause before changing anything.
Stop the cycle, retain raw evidence, and use the reviewer in a fresh context. A product defect returns the candidate upstream; it is not repaired inside that TestForge cycle.
Privacy, storage, network, security
The v2.0.0 skills-only plugin includes no account, telemetry, analytics, hosted service, connector, MCP server, hook, or automatic network request. Deterministic scripts touch only paths and commands the user chooses.
Prompts, repositories, logs, uploads, model calls, retention, training, residency, and connector traffic are governed by Codex, Claude, configured models, Git hosts, and any authorized tools—not by TestForge.
Verification manifests, reports, raw command captures, generated tests, evaluation runs, seals, and promoted baselines may contain sensitive project evidence. Store and delete them under an approved retention policy.
Imported files and tool output are untrusted evidence. Active security work needs a named target, explicit permission, non-production default, time window, rate limits, prohibited actions, handling rules, and a stop contact.
Provenance and evidence status
| Claim | Current evidence | Boundary |
|---|---|---|
| Constructed and packaged | v2.0.0 source, plugin tree, current Claude archives, manifests, and retained frozen-release receipts | Does not prove host installation or live behavior |
| Deterministic behavior | Repository-local tool, package, distribution, documentation, and eval-harness suites | Only exercised commands, fixtures, Python version, and environment |
| Behavioral evaluation | Named model/context baselines and a deliberately failed fresh-package smoke | Single-trial, model-, context-, case-, and judge-bounded; not universal model quality |
| OpenAI directory | Latest retained portal packet is v1.1.4 and repository-tested | No claim of upload, approval, publication, or discoverability |
| Host activation | Installation probes are documented | Current live Codex and Claude activation remain separate observations |
Trust and authority
Comments, issues, fixtures, logs, dependencies, and retrieved text remain untrusted input—not instructions to the verifier.
Commands, writes, network access, PR access, browsers, production targets, and external actions exist only when the host proves them.
Production-code edits, installs, weakened tests, CI changes, destructive operations, active security work, and publication require explicit human authority.
MIT covers software and schemas; CC BY-ND 4.0 covers authored Augment materials. Third-party and user material retain their own rights.