Coding agents handle substantial portions of professional software workflows, clinical decision-support agents triage patients, and conversational agents mediate human learning and relationships. The community's evaluation methods, built largely for static benchmark performance on chat completion, have not adapted at the same pace.
A recent review of agentic-AI evaluation found that 83% of evaluations are dominated by technical metrics, while human-centered (30%), safety (53%), and economic dimensions (30%) remain peripheral (Jafari Meimandi et al., 2025). HAIC 2026 builds a methodological foundation for the empirical evaluation of human-agent teams. The central question is how to evaluate and govern human-agent systems rigorously as they coevolve with the people who use them, both in general deployment and in domains where the gap between benchmark and reality has the highest stakes?
This is the second edition of the HAIC series, following the inaugural
ICLR 2025 workshop on human-AI coevolution. Where the first workshop mapped coevolution broadly across five themes, this one commits to a single focused operationalization: rigorous empirical evaluation of human-agent teams, with each theme anchored in recent peer-reviewed evidence.
The program grounds its discussion in case studies from high-stakes domains where the organizing team has direct research access: healthcare, mental health, aviation, and finance. The format weights discussion over talks, with breakouts feeding an open-problems registry and a community position paper.