Agent evaluation infrastructure
GDPval Actions extends deliverable benchmarking into executed digital labor. Agents operate inside simulated enterprise systems, and every state change, approval, recovery and cost is measured against expert human baselines.
Produces an expert-quality artifact.
Follows a defined multi-step workflow across applications.
Acquires context, adapts, and recovers from exceptions.
Manages work over time across stakeholders and deadlines.
Q = 0.40·Outcome + 0.20·Execution + 0.15·Verification + 0.15·Governance + 0.10·Efficiency, with hard gates that zero a run.
Deterministic state tests, artifact review, trajectory review and blinded expert comparison against human work.
Reliability is a lower confidence bound, not an average. Autonomy is earned tier by tier.