HAIC NeurIPS 2026
Second HAIC workshop · NeurIPS 2026

Human-AI Coevolution

Measuring human-agent teams in the agentic era.

  • Atlanta
  • December 12–13, 2026
  • In person, one day

Papers due August 29, 2026 (AoE) Decisions September 29, 2026

About the workshop

Evaluation has not kept pace with deployment

Coding agents handle substantial portions of professional software workflows, clinical decision-support agents triage patients, and conversational agents mediate human learning and relationships. The community's evaluation methods, built largely for static benchmark performance on chat completion, have not adapted at the same pace.

A recent review of agentic-AI evaluation found that 83% of evaluations are dominated by technical metrics, while human-centered (30%), safety (53%), and economic dimensions (30%) remain peripheral (Jafari Meimandi et al., 2025). HAIC 2026 builds a methodological foundation for the empirical evaluation of human-agent teams. The central question is how to evaluate and govern human-agent systems rigorously as they coevolve with the people who use them, both in general deployment and in domains where the gap between benchmark and reality has the highest stakes?

This is the second edition of the HAIC series, following the inaugural ICLR 2025 workshop on human-AI coevolution. Where the first workshop mapped coevolution broadly across five themes, this one commits to a single focused operationalization: rigorous empirical evaluation of human-agent teams, with each theme anchored in recent peer-reviewed evidence.

The program grounds its discussion in case studies from high-stakes domains where the organizing team has direct research access: healthcare, mental health, aviation, and finance. The format weights discussion over talks, with breakouts feeding an open-problems registry and a community position paper.

Themes

How the themes relate

The themes are not independent. Deployment moves the validity target (Theme 1), the human feedback meant to correct course is itself contested (Theme 2), and evaluation must adapt as systems and users coevolve (Theme 3). Full descriptions are in the call for papers.

Theme 1

Validity of evaluation in deployed contexts

Benchmarks assume a static target. In deployment the target moves, and the constructs being measured, such as productivity, helpfulness, and safety, are themselves contested across domains. How do validity frameworks adapt?

Theme 2

Expert disagreement and the limits of human feedback

RLHF assumes aggregated feedback approximates a coherent target. A growing line of work treats disagreement as signal rather than noise. When does it mark evaluation invalidity, and when does it reflect domain pluralism that deployed systems should preserve?

Theme 3

Adaptive testing and continual evaluation

Deployed teams coevolve: skills reallocate, populations shift, distributions drift. A fixed test set can lose validity with no visible signal. We seek methods that track and respond to drift.

The three themes form a cycle: Theme 1 feeds Theme 2, Theme 2 feeds Theme 3, and Theme 3 returns to Theme 1 as drift moves the evaluation target again.

Key dates

Timeline

Dates follow the NeurIPS 2026 recommended workshop timeline. All deadlines are Anywhere on Earth.

Submission and notification schedule
MilestoneDateNotes
Submission deadline August 29, 2026 (AoE) Double-blind via OpenReview
Acceptance notification September 29, 2026 (AoE) Hard deadline
Camera-ready To be announced Set with authors after decisions
Workshop, Atlanta December 12–13, 2026 One day within that window, confirmed by NeurIPS
Invited speakers

Speakers and panel

To be announced. The invited speakers are being confirmed and will be listed here, together with the four panelists for the panel on evaluation in practice, which covers clinical practice, AI safety research, applied industry evaluation, and policy.

Invited speakers will be announced on this page. Enquiries may be directed to the organizing committee.

Organizing committee

Organizers

Ahmad Rushdi

Ahmad Rushdi

Sponsorship & inclusion

Stanford University (HAI)

Call for papers

Scope of submissions

We invite new evaluation methods and benchmarks, datasets of human-agent interaction, reproducible benchmark critiques, case studies from high-stakes domains, and position papers. The workshop is non-archival, and each submission receives three double-blind reviews via OpenReview.