Skip to content

Softprobe Agent Evaluation

Run existing eval suites in controlled environments. Capture complete native evidence. Compare and gate with one workflow.

Softprobe Agent Evaluation is an open workflow/control plane for AI evaluation frameworks.

System at a glance

What Softprobe does

  • Packages and pins framework-native definitions
  • Controls execution environment (network, mounts, secrets, limits)
  • Captures complete native outputs (results, logs, traces, cost/usage)
  • Normalizes outer lifecycle (requested → validated → running → terminal)
  • Compares runs and drives gates across local CI and managed execution

What Softprobe does not do

  • Replace Promptfoo/DeepEval DSLs
  • Promise full assertion-type parity in a new schema
  • Require users to author a new eval language

Product areas on this site

AreaYou use it when...
TestingJava record/replay regression with JVM agent
PlatformIstio/SESSIFY observability
Agent EvaluationWorkflow + environment control + evidence + gates for framework suites

Start here

PersonaStart page
Framework user (Promptfoo/DeepEval)Quick start
Migration leadPromptfoo integration
Agent environment / record-replayRecord and replay an agent environment
Platform operatorArchitecture
CI / AI coding agentFor AI agents

Zero code changes · Full-context visibility · Cost optimization