Open source / v0.1.0

Verify what your AI agent is actually running.

RuntimeTruth is an evidence-backed runtime verification CLI for AI agents. It separates declared configuration, resolved effective state and live evidence, then detects semantic drift against a caller-controlled baseline.

Python 3.12+ · MIT · Local-first · No account · No RuntimeTruth telemetry

Evidence before inference Local-first Deterministic core

Source code is not the whole runtime.

Effective model routing, project instructions, MCP capabilities, sandbox policy and runtime versions can change independently. Looking only at the repository answers what was written, not necessarily what the agent actually resolved and can use.

01

Model and provider

Workspace configuration can leave values unset while the runtime resolves a concrete effective model and provider when work starts.

02

Instructions

Project instruction files can change without application code changing. Their source paths and fingerprints are part of the effective runtime identity.

03

Capabilities and permissions

Sandbox/network policy and the MCP capability surface determine what an agent can actually reach and do.

Declared, resolved and live are different kinds of truth.

D

Declared

What configuration and deployment inputs say should be used.

R

Resolved

What the inspected runtime's own resolution path materializes from that configuration.

L

Live

What RuntimeTruth can observe directly from the environment at collection time.

Design rule Unknown is better than confidently inferred.
  • every evidence record keeps its provenance
  • missing values stay different from JSON null
  • provenance changes are rejected when they cannot be compared faithfully

The effective runtime appeared only when the runtime resolved it.

In real validation, canonical workspace configuration left model and provider fields unset. Creating a new ephemeral Codex thread materialized the effective model/provider plus approval, sandbox and instruction-source state.

Config read
  • workspace identity
  • canonical configuration
  • some effective fields still unresolved
Ephemeral thread
  • effective model/provider
  • approval and sandbox state
  • instruction-source paths

RuntimeTruth uses Codex's own app-server interfaces for that resolution instead of reimplementing Codex precedence rules. The probe creates no model turn.

Strict when you need it. Selective when reality is noisy.

A baseline can gate the whole supported snapshot, or protect only high-value invariants such as effective model, sandbox policy or instruction fingerprints.

01

Capture

Collect an evidence-backed snapshot for a real workspace or runtime target.

02

Compare

Use one semantic diff engine for human-readable changes, strict baseline verification and live Codex verification.

03

Gate

Return stable exit codes: 0 PASS, 2 DRIFT, 1 ERROR. Add JSON output when CI needs structured changes.

protect codex.thread.model
protect codex.thread.sandbox
protect codex.instructions

Verify the agent consuming the dependency, not the dependency contract itself.

RuntimeTruth records thread-scoped MCP server status and deterministic tool-catalog fingerprints as part of agent runtime identity. It does not classify breaking schema changes, tool-description rug pulls or protocol compliance.

RuntimeTruth
  • Which MCP capability surface can this agent see?
  • Did that agent-visible surface drift?
  • What else changed in the same effective runtime?
Dedicated MCP testing
  • schema/description compatibility
  • protocol and behavioural checks
  • server-focused security heuristics

Tools such as MCPWard fit naturally beside RuntimeTruth when a pipeline needs both layers.

Built from a real runtime outward.

01

Real Codex runtime

Config resolution, ephemeral-thread state, instruction fingerprints and MCP status were validated against a real local Codex environment.

02

Real drift

Changing an instruction file produced a fingerprint-only semantic drift without storing the instruction plaintext.

03

Public release

The v0.1.0 tag builds into a clean wheel, installs from the public repository and reports the expected version.

PASS is intentionally a narrow claim.

PASS means the selected evidence matched the selected baseline under RuntimeTruth's comparison rules. It is not a security certification, compliance assessment or behavioural guarantee.

Current project model Free, local-first open source.
  • no account or RuntimeTruth telemetry
  • no hosted control plane or SLA
  • caller controls and protects the baseline

Inspect the project yourself.

For the engineering decisions behind the project, see Artur Panek's RuntimeTruth case study. For the underlying evidence model as a standalone technical argument, read A config file does not tell you what an AI agent is actually running.