Real-session compositional workflows
200 executable tasks are reconstructed from anonymized, privacy-screened multi-turn sessions, preserving interaction history and workspace state across 8 broad scenarios and 17 fine-grained task types.

Research benchmark · DuMate
AI Agent Benchmark for Real-World Work
Evaluate agents on realistic tasks and environments.
DuMateBench evaluates whether autonomous agents can coordinate tools, produce artifacts, and recover from failures across realistic end-to-end workflows.
Why DuMateBench
The benchmark is designed around the gap between clean, isolated demos and the messy, multi-step work agents must complete in practice.
200 executable tasks are reconstructed from anonymized, privacy-screened multi-turn sessions, preserving interaction history and workspace state across 8 broad scenarios and 17 fine-grained task types.
Isolated Docker environments model insufficient tools or resources, unstable network and tool behavior, and noisy workspaces with distracting data.
Five agent frameworks and four base models form 20 configurations, measuring end-to-end performance, workspace-noise robustness, efficiency, and failure modes.
Environment & capability coverage
Tasks combine multiple capabilities and run under controlled environmental constraints.
These conditions are instantiated in isolated Docker containers so reliability can be evaluated under controlled and repeatable constraints.