HAIC NeurIPS 2026
Call for papers

Measuring human-agent teams in deployment

Submissions are due August 29, 2026 (AoE). The workshop is non-archival, with a 9-page limit and three double-blind reviews per submission via OpenReview. We particularly welcome methods the community can adopt in operational settings.

Overview

Aims and scope

Agentic AI has moved from research demos to broad deployment. Coding agents handle substantial portions of professional software workflows (Anthropic, 2026), clinical decision-support agents triage patients (Taylor et al., 2025), and conversational agents mediate human learning and relationships (Shi et al., 2025; Zhang et al., 2025). The community's evaluation methods, built largely for static benchmark performance, have not kept pace.

Recent work finds that 83% of agentic-AI evaluations are dominated by technical metrics, while human-centered, safety, and economic dimensions remain peripheral (Jafari Meimandi et al., 2025). HAIC 2026 builds a methodological foundation for the empirical evaluation of human-agent teams. The central question is how to evaluate and govern human-agent systems rigorously as they coevolve with the people who use them, both in general deployment and in high-stakes domains such as healthcare, mental health, aviation, and finance.

This is the second edition of the HAIC series, following the inaugural ICLR 2025 workshop. Where the first workshop mapped coevolution broadly across five themes, this one commits to one focused operationalization, with each theme anchored in recent peer-reviewed evidence.

Themes

How the themes relate

The themes are not independent: deployment moves the validity target (Theme 1), the feedback meant to correct course is itself contested (Theme 2), and evaluation must adapt as systems and users coevolve (Theme 3). Works cited below are listed under references.

Theme 1 · Validity of evaluation in deployed contexts

Traditional benchmarks assume a static evaluation target. In deployed human-agent systems the target moves, and the constructs being measured, such as productivity, helpfulness, and safety, are themselves contested across domains. Recent work documents an evaluation blind spot in deployed healthcare agentic AI, where systems with strong technical metrics fail in clinical practice and usability, safety, and clinical-outcome evaluation are underrepresented (Njei et al., 2026; Aránguiz Dias et al., 2026). The causes are both conceptual and structural, from contested definitions of "agentic" to mismatch between benchmark assumptions and workflow reality. This theme asks how validity frameworks can be adapted to deployed agentic contexts, including the five-pillar validity model (Salaudeen et al., 2025), BetterBench (Reuel et al., 2024), and the Agentic Benchmark Checklist (Kapoor et al., 2024), and invites case studies from clinical care, aviation, finance, education, and general-purpose deployment.

Theme 2 · Expert disagreement and the limits of human feedback

Reinforcement learning from human feedback rests on the assumption that aggregated human feedback approximates a coherent target. A growing line of work treats disagreement as signal rather than noise: high-disagreement annotation items often reflect genuine ambiguity rather than error (Aroyo and Welty, 2015; Pavlick and Kwiatkowski, 2019), the DICES dataset encodes rater votes as distributions across demographic groups rather than a single label (Aroyo et al., 2023), and even foundational RLHF work reports only moderate expert-crowd agreement on safety (Bai et al., 2022). A mixed-methods study of three psychiatrists evaluating 360 LLM-generated mental-health responses finds systematic disagreement driven by incompatible clinical frameworks, not measurement noise (Jafari et al., 2026). This theme asks when inter-rater disagreement signals evaluation invalidity versus genuine domain pluralism that deployed systems should preserve, how to build feedback pipelines robust to it, and what costs annotators bear at scale (Casper et al., 2023) as degraded feedback trains the next generation of models.

Theme 3 · Adaptive testing and continual evaluation for coevolving systems

Once human-agent teams are deployed, they coevolve: skill reallocation emerges over months, user populations shift, and the training distribution drifts as humans adapt to the system's outputs. The effect on evaluation is documented: benchmarks of temporal distribution shift report average performance drops on the order of twenty percent from in-distribution to deployed data, and existing methods do not close that gap (Yao et al., 2022). Fixed benchmarks break down here, and a static test set can lose validity with no visible signal that it has done so. This theme focuses on methods that track and respond to drift: adaptive testing, automated red-teaming (Hardy et al., 2024), continual evaluation infrastructure (Saxena et al., 2025), longitudinal study design for deployed systems (Silacci et al., 2026), rare-event simulation for emergent failure modes (Lee et al., 2020), and detection of when a benchmark has lost validity to coevolutionary drift.

Submission topics

Topics of interest

  • New evaluation methods and benchmarks for human-agent systems.
  • Datasets capturing human-agent interaction in deployed or longitudinal settings.
  • Studies where strong benchmark performance fails to transfer to practice.
  • Reproducible critiques and replications of existing agentic benchmarks.
  • Case studies from high-stakes domains: healthcare, mental health, aviation, finance.
  • Position papers advancing validity-centered or coevolutionary evaluation.

The workshop is non-archival. Accepted contributions are posted on OpenReview, but the workshop is not a publication of record, so authors remain free to submit the same work to a conference or journal afterwards. Length, format and policy details are in how to submit below.

Key dates

Timeline

Dates follow the NeurIPS 2026 recommended workshop timeline. All deadlines are Anywhere on Earth.

Submission and notification schedule
MilestoneDateNotes
Submission deadline August 29, 2026 (AoE) Double-blind via OpenReview
Acceptance notification September 29, 2026 (AoE) Hard deadline
Camera-ready To be announced Set with authors after decisions
Workshop, Atlanta December 12–13, 2026 One day within that window, confirmed by NeurIPS
Submitting

How to submit

Submission instructions

The paper submission deadline is August 29, 2026, 11:59pm Anywhere on Earth. Papers are submitted through the HAIC 2026 OpenReview portal. Supplementary material is due at the same time as the main paper.

Papers may be revised as many times as needed up to the submission deadline. Revisions are not permitted while a paper is under review.

Paper length

The main text must take up at most 9 pages. This limit is enforced strictly: a submission with main text on the 10th page will be desk rejected. The limit applies to the initial submission and to the final camera-ready version alike.

References do not count toward the limit, and there is no cap on pages used for the bibliography. Authors may add as many pages of appendices after the bibliography as they wish, but reviewers are not required to read the appendix. The full nine pages are worth using only where larger figures or additional detail genuinely call for them.

Style files and templates

Submissions must use the NeurIPS 2026 LaTeX style files: Formatting Instructions for NeurIPS 2026 (ZIP, 20 KB). The archive contains neurips_2026.sty, the template neurips_2026.tex, and checklist.tex.

This workshop reviews double-blind, so load the style file with the dblblindworkshop option and nothing else. Omitting final and preprint is what anonymizes the paper and adds the line numbers reviewers refer to:

\usepackage[dblblindworkshop]{neurips_2026}
\title{Your paper title}
\workshoptitle{Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era}

Workshop papers need \workshoptitle{} in addition to \title{}; it sets the footnote identifying the workshop. Accepted papers add final for the camera-ready version, \usepackage[dblblindworkshop, final]{neurips_2026}, which reveals author names and removes the line numbers. Do not use sglblindworkshop, which de-anonymizes the paper, and do not alter the margins, font sizes, or line spacing the style file sets.

Submissions must be anonymized. Author names, affiliations and acknowledgements are omitted, and references to the authors' own prior work are written in the third person.

Policies

Submission policies

Double-blind reviewing

Reviewing is double blind. Reviewers cannot see author names while conducting reviews, and authors cannot see reviewer names. Reviewer comments remain anonymous to authors.

Reviews are visible to the paper's authors and to the organizing committee, and are not published. Rejected submissions are not made public and are never deanonymized. Accepted contributions are posted publicly on OpenReview after the notification, and the program includes spotlight talks and posters of accepted papers.

Non-archival status and dual submission

The workshop is non-archival. Acceptance is not a publication of record, so authors may submit the same work to a conference or journal afterwards without prejudice.

Posting a submission to a preprint server such as arXiv is permitted, both before and during the review period. Work that has appeared on a non-peer-reviewed website, or that has been presented at a venue without published proceedings, may be submitted. Papers citing the authors' own earlier related work may be submitted, subject to the anonymization requirement above.

Use of large language models

Large language models may be used as a general-purpose assistive tool. Authors and reviewers take full responsibility for everything submitted under their names, including model-generated content that could be construed as plagiarism or scientific misconduct, such as fabricated results or references. Language models are not eligible for authorship.

Withdrawal

Authors may withdraw a submission at any time before the notification. Withdrawn submissions are removed from consideration and are not made public.

Code of conduct and ethics

All participants, authors included, are required to follow the NeurIPS Code of Conduct. Authors should also read the NeurIPS Ethics Guidelines, which apply to submission, reviewing and discussion alike.

A workshop code of conduct, setting out expectations for a respectful and inclusive environment, will be published here before the workshop.

Reviewing

Review process

Each submission receives three double-blind reviews via OpenReview. Reviews follow NeurIPS conflict-of-interest policy strictly. Organizers will not submit contributions to the workshop, nor will their immediate students or postdocs, and reviewers will not review papers from their own institutions or from organizations defined as conflicting under the NeurIPS policy.

The program committee comprises approximately 30 researchers, with domain coverage targeting one third agentic-AI evaluation and machine-learning methodology, one third human-robot interaction and human-agent teaming, and one third applied researchers with direct deployment experience in healthcare, aviation, finance, or related high-stakes settings. Each reviewer is capped at three papers.

The full program committee will be announced here once recruitment closes.

Outcomes

Intended outputs

  • All accepted contributions posted on OpenReview.
  • A community position paper from the organizers within six months of the workshop.
  • An open-problems registry on GitHub scoped to human-agent teaming evaluation, assembled from the breakout sessions.
  • A quarterly virtual continuation series.
References

Works cited

Deadline August 29, 2026 (AoE)

Submissions

Reviewing is double blind and the workshop is non-archival, with a 9-page limit on the main text. Full requirements are set out in how to submit.