GDPval Actions

Agent evaluation infrastructure

Measure whether AI agents can do the work, not merely describe it.

GDPval Actions extends deliverable benchmarking into executed digital labor. Agents operate inside simulated enterprise systems, and every state change, approval, recovery and cost is measured against expert human baselines.

L1
Deliverable

Produces an expert-quality artifact.

L2
Directed Execution

Follows a defined multi-step workflow across applications.

L3
Situated Execution

Acquires context, adapts, and recovers from exceptions.

L4
Sustained Digital Labor

Manages work over time across stakeholders and deadlines.

Quality score Q

Q = 0.40·Outcome + 0.20·Execution + 0.15·Verification + 0.15·Governance + 0.10·Efficiency, with hard gates that zero a run.

  • O · Outcome Correctness
  • X · Execution Quality
  • V · Verification & Recovery
  • G · Governance
  • E · Efficiency

Four evaluation stages

Deterministic state tests, artifact review, trajectory review and blinded expert comparison against human work.

  • A · State assertions
  • B · Artifact rubric
  • C · Trajectory rubric
  • D · Blinded preference

Certification ladder

Reliability is a lower confidence bound, not an average. Autonomy is earned tier by tier.

  • Observed Assistant
  • Supervised Contributor
  • Bounded Operator
  • Autonomous Digital Worker
  • Resilient Digital Worker